Skip to content

RLM-2b: a typed launch-environment convergence scope whose plan terminal is FullyApplied - #9832

Merged
briansrls merged 27 commits into
mainfrom
session/smart-moth-158
Sep 2, 2026
Merged

briansrls merged 27 commits into
mainfrom
session/smart-moth-158

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Aug 31, 2026 •

Copy link
Copy Markdown
Contributor

The question

Is the production roadmap instance demonstrably capable of executing RLM-1 at the exact accepted revision? The RLM-2b plan run over srv1 answered partially_applied|33|slot=31 cap=0 timer=0 fabric-cell=0 activation=2, and gunbc.roadmap_launch_deployment_receipt accepts only FullyApplied. So RLM-2 could not close.

Thirty-one of those refusals are REFUSED-REMOVAL on runner slots with no provenance signal, and two are the runner-service activation stall every production caller supplies. Both belong to CI runner capacity and to its own convergence plan. The scope was answering for a wider resource set than the question, so the repair makes the question's resource set nameable rather than relaxing the threshold. roadmap_launch_deployment_receipt's FullyApplied-only acceptance is untouched.

Head b676db0 — review 5070355515 P1-4, and it is a SUBSTITUTED remedy, not compliance

The review asks for "an executing target proving the real consumer resolves and constructs the artifact through the same terminal authority" in gunbc.spark.serving_converge_plan, and its literal remedy is a caller for spark_serving_fleet_converge_apply_shell. This head does not add one, and the reviewer should decide whether the substitution is acceptable rather than discover it.

That wrapper composed fleet_converge_apply_shell with fleet_converge_apply_terminal over one request — exactly what fleet_converge_plan_artifact_of_request already does, computing the terminal once and handing that same value to the plan body line, to apply_terminal, and to the apply script. A second spelling of one composition is a second name for one fact (DESIGN §3), and the two reconstructions could disagree about the terminal with nothing joining them.

It had no caller to migrate. #9763 deleted the Spark-specific artifact route and with it this wrapper's only caller; the symbol has occurred exactly once in the tree — its own declaration — on every commit since, origin/main today included. Giving a dead wrapper a caller would have minted the parallel authority rather than removing it. So the wrapper is deleted, together with the ten imports it alone kept alive (each verified import-only first), and the property it endangered is asserted on the route production actually takes.

The blocker's description is one head stale, stated rather than repaired around. P1-4 says the module's FleetConvergePlanArtifact literal still omits apply_terminal. That was accurate at the reviewed head c67b2d26e83; it is not accurate here, because the main merge brought in #9763, which deleted that literal and its whole assembler. The defect is real — the orphan is what survives — but its description no longer matches the tree.

Why the repair is here rather than in a prerequisite. The acceptance bar requires the apply shell bound to the same request and terminal authority. That seam exists only on this branch: fleet_converge_plan_artifact_of_request is absent from origin/main. A prerequisite would have asserted a true property about a route that does not carry the seam under review.

Two rows in test.claim.spark_serving_converge_slice_witness, over slice_artifact, which builds the artifact through fleet_converge_plan_artifact with a populated Spark axis (spark_serving_axis_for_observation):

  • the_spark_populated_artifact_names_its_own_terminal_in_body_and_script — the plan body and the apply script both name fleet_converge_apply_terminal_line of the artifact's own apply_terminal. One terminal, two projections.
  • a_terminal_no_observation_produced_is_named_by_neither_projection — the control. A PartiallyApplied terminal nothing produced must be named by neither. Without it the positive passes on any artifact whose script contains some terminal line, which is precisely what a disagreeing second assembler emits.

Executed, at identity grain, in run 33552307413 (required_floor_disposition.tsv):

the_spark_populated_artifact_names_its_own_terminal_in_body_and_script   planned_as_changed_witness   passed
a_terminal_no_observation_produced_is_named_by_neither_projection        planned_as_changed_witness   passed

Whole floor on this head: planned=3429 planned_as_changed_witness=79, outcomes 3433 passed, 26 known-red-held, 49 route-gap-before-verdict, 0 failed, 0 interrupted, 0 budget-refused.

What "executing target" means here, and it is weaker than it sounds. v2.workflow.required_floor required_gate_prefixes carries no spark, fleet or roadmap prefix. These witnesses reach the required floor through the changed-witness sublane, which seeds changed modules into the prepared closure — visible in the same TSV, where the module's other 30 identities read declined_outside_required_gate. So they execute because the modules are changed in this PR, and after merge they are outside the required gate again. That is the population of the declared rung drop gunbc.rung_drop required_gate_bankruptcy, not a new defect, and the same qualifier applies to the P0-1, P0-2 and P0-3 evidence on this PR: executes-at-merge-time, not executes-continuously. No roster was widened to change that.

Local pre-flight established nothing and is not cited. Two remote claim_batch dispatches — the two rows, and a control naming a witness that does not exist — were both OOM-killed (SIGKILL, 137) during resolve at 5.4 GiB, so neither arm reached a verdict. Nothing was tuned to get under the envelope. The floor above is the evidence.

Read this first: a claim this PR previously made is withdrawn

An earlier revision of this branch — and an earlier revision of this description — said the omitted axes were unconstructible, that "the invalid state has no constructor", and that the property sat at §4b rung 4. That is false, and this PR now carries the measurement that falsifies it.

The current compiler performs no excess-field judgment at all on a payload-carrying coproduct arm literal. LaunchEnvironmentConverge { host, observed_timers, observed_slots: [] } compiles. So does the same literal with a field name declared on no type anywhere in the corpus. The mechanism — payload arms are not projected to their arm record, so the record-literal field roster is Absent and the unknown-field arm's Absent case is silence — is owned by the compiler-floor lane adhoc-0f99c6bc-193 (session sharp-otter-269), filed for exactly this class and cited here rather than re-derived.

The claim this PR now makes, and the evidence for each clause:

LaunchEnvironmentConverge declares only host and observed_timers; every production constructor on this tree supplies exactly those two; no authoritative projection consumes an excess field; and an excess field cannot alter the launch plan, terminal, effects, manifest or receipt. Excess-field initializers are semantically analysed; their resulting field values are non-authoritative to the launch-environment transaction, and are not runtime-evaluated. Compiler-level refusal is owned by the floor lane above.

That is mechanically excluded from the authoritative launch-environment projection on this exact tree — not "structurally impossible", and not "mechanically preventable" without that qualification. Review 57942 approved the earlier shape by reading the type declaration; that rationale is falsified and the approval does not count toward this proposition.

The construction

FleetConvergeScope gains LaunchEnvironment; FleetConvergeRequest gains

LaunchEnvironmentConverge { host, observed_timers }

Caps and Fabric cells were in an earlier cut of this arm and are removed as foreign: gunbc_runner_per_slot_dropin_path resolves under actions-runner@.service.d, so the cap axis is the CI-capacity family reached by another name (executed by the_cap_axis_lives_under_the_actions_runner_unit_template), and the governing plan document rules Fabric off this proof's path. Runner slots and activation were never in it. Passing [] is not the alternative: an empty list is a positive claim the host has none, and it licenses removals.

Making the narrowing honest required moving observed_activation_readiness into the FullHostConverge arm and turning the terminal, the apply shell and the plan-body axis/counter lines into folds over the request. While the axis set was a property of the call, a narrower scope could only be expressed by passing something for axes it had not selected.

What the evidence actually establishes

The compile-control matrix is a substrate characterization, not a wall, and says so in its own header. One shared source builder parameterized only by the mutation; one census hoisted per source so the whole matrix is one execution; comparisons are multiset deltas of diagnostic identities against the valid constructor, not totals, because +1 and -1 cancel.

The census compiles its synthetic source against the live tree and reports the whole closure, so the valid constructor already sits on a nonzero blocking baseline that belongs to the corpus. Absolute-zero assertions would therefore be assertions about the corpus and would redden on an unrelated merge; every launch row is a delta measured through the same instrument in the same run.

The plain-product sight control is not an exception to that, and an earlier revision of this description said it was. Its closure is small but not empty: a common population of nonblocking rows belonging to the census module itself appears on both sides, so clean total = 0 / mutant total = 1 is false and was measured false. That pair is therefore also a paired delta — blocking 0 vs 1, the exact sentinel identity absent vs present once, total clean + 1, and every other identity's multiplicity equal. That is stronger than a net-count comparison and is the strongest subject-relative control available under a shared nonblocking residue; it is not stronger than a genuinely clean absolute would be, since a true total-zero pair would also prove the absence of all common residue.

Probe Result
valid { host, observed_timers } accepted (the within-run control)
+ observed_slots / observed_caps / observed_fabric_cells / observed_activation_readiness admitted, empty identity delta, zero FieldNotFound
+ all four admitted, empty identity delta
+ a field name declared nowhere admitted, empty identity delta
same name on a plain product literal exactly one blocking FieldNotFound naming it
omit required observed_timers exactly one added MissingField naming it, nothing else changed
excess field initialized by a nowhere-declared function exactly one added blocking diagnostic naming that function

The plain-product row is what makes the zeros interpretable rather than a blind census — without it, four zeros cannot distinguish "admitted" from "the harness cannot see this class". The missing-field row pins the asymmetry that is the defect: the literal refuses a missing field and admits an invented one. The last row settles the wording: the excess field's initializer is analysed even though the field is discarded, so nothing here calls it "inert".

The non-interference matrix is where the guarantee actually lives. Six execution pairs over the complete plan-artifact projection — subject, scope, member-set fingerprint, observed baseline, lease key and fingerprints, terminal and terminal wire, bundle identity, plan body, apply shell — each requiring identity between a clean launch request and one differing only by a smuggled foreign field, at the payload that would be maximally visible under its owning scope (the production-shaped 31-slot population, the two-refusal activation readiness, an actionable cap population, a refusing Fabric observation, all four together, and an arbitrary nowhere-declared sentinel).

Every pair carries an owning-scope control showing the same payload moving the full-host projection, so the equality is the launch arm ignoring the value and not a dead fixture.

The sentinel's initializer is a refinement cast that must refuse if evaluated. It does not refuse, so excess-field initializers are analysed but not runtime-evaluated. That is a positive observation — evaluation would have been loud — not an inference from unchanged plan bytes, which would be equally consistent with evaluate-then-discard.

There is no constructor-site census any more, and that is a loss this PR takes. An earlier revision of this branch carried one — a decl_facts fold over the production corpus asserting every LaunchEnvironmentConverge construction site at identity grain against a declared roster of two. It is deleted; see Disposition: the site census is removed, not re-homed below. Nothing executing stands over WHERE a launch request is constructed after this PR.

Two things I could not build, stated plainly

1. The field-NAME set is not authorable in .dag today. The names exist in the v1 AST, but the reflection projection exposed to the substrate drops them: marshal_generic pushes every field-init child on an unlabelled positional edge, and the only named edge a record literal mints is record_construction_spelling. The census therefore closes where launch requests are built, not what is put in them. Next-rung trigger, stated as the capability: a substrate-readable per-field label on record-literal children, sufficient for a witness to assert the exact field-name set of a named construction site. No grep is offered as a substitute.

2. The census is not inside the required-floor gate, and I did not put it there. required_gate_prefixes does not admit test.claim.fleet_converge_* at all — no fleet witness, old or new, is in that closure. Adding a prefix is described by that module's own header as a wall-clock decision made there and nowhere else, and this census measures ~120s for the first row and ~50s for the second against a 5000ms floor deadline. Enrolling it would either blow the budget or land it as declared cost debt, which does not execute. This is an open item for the reviewer, not something I resolved by improvising. The census executes today in the witnesses lane; making it a required gate is a decision I do not own.

The persisted-member map — every file, exactly one category

The rule: a persisted file need not be in the authority digest, but only when apply no longer trusts it for any decision — it may not select the observer arm, the baseline population, the member-set fingerprint authority, the terminal, the apply shell, or the receipt subject. Before this PR subject_scope.txt failed that test while sitting outside the digest, and the gap was invisible because nothing stated which category each file was in.

The authority is FleetConvergePersistedMemberAuthority in gunbc.fleet_converge_plan_manifest; the table below is its projection, not a second map. Three arms, no wildcard, and the classification fold is total over the path — an unrecognized path answers UnclassifiedPersistedMember, so an eleventh sidecar fails loudly rather than acquiring whatever category a default arm happened to name.

There is deliberately no "covered incidentally by another fingerprint" arm — that sentence is precisely what left subject_member_set.txt and observed_baseline.hex unexamined: each was folded into some hash, so nobody asked whether apply read them, and it did.

# Persisted member Category Bound to Executed refusal
1 subject_host.txt 1 — authoritative, receipt-bound receipt observed_host a_substituted_subject_host_alone_refuses
2 subject_scope.txt 1 — authoritative, receipt-bound receipt scope_wire a_coordinated_other_scope_substitution_refuses_before_any_observation
3 subject_member_set.txt 1 — authoritative, receipt-bound receipt member_set_fingerprint_hex a_substituted_member_set_fingerprint_alone_refuses
4 observed_baseline.hex 1 — authoritative, receipt-bound receipt observed_baseline_hex a_substituted_baseline_digest_alone_refuses
5 prior_generation.txt 1 — authoritative, receipt-bound receipt prior_generation a_rewound_generation_sidecar_refuses, a_corrupt_generation_sidecar_refuses_rather_than_coercing
6 plan_generation.txt 1 — authoritative, receipt-bound receipt planned_generation same pair, planned side
7 plan_content.hex 1 — authoritative, receipt-bound receipt plan_artifact_hash a_bundle_hash_that_disagrees_with_the_receipt_refuses
8 plan.txt 2 — derived, recomputed from bound authority re-digested into plan_artifact_hash doctored_plan_body_bytes_refuse_even_when_the_hex_claim_still_matches
9 apply.sh 2 — derived, recomputed from bound authority re-digested into plan_artifact_hash doctored_apply_bytes_refuse_even_when_the_hex_claim_still_matches_the_receipt
10 spark_serving_typed_actions.wire 2 — derived, recomputed from bound authority re-digested into plan_artifact_hash covered by the same bundle recomputation

Category 3 (non-authoritative, never read for admission or actuation) is empty, and that is measured rather than asserted: no_persisted_member_is_currently_non_authoritative asks every rostered member and none answers NonAuthoritativePresentation. An empty set has two causes — nothing qualified, or nothing was asked — and that row settles which. Every file the plan writes is either compared to the run-bound receipt directly or recomputed into a value that is. Rows 7 and 8–10 are genuinely different failures — row 7 doctors the claim, rows 8–10 doctor the bytes — which is why the digest step recomputes rather than comparing two claims.

fleet_converge_generation_store_path is host state outside the plan artifact directory and outside this payload's subject; it is admitted separately by observe_generation_store_wet, which refuses corrupt and negative values rather than coercing.

The paired positive traverses the same path. an_untouched_artifact_population_is_admitted_and_routes_on_the_receipt_scope goes through fleet_converge_apply_payload_admits — the same receipt decoding and the same apply-composition function every RED above uses — not a helper. And the_routed_scope_is_the_receipts_and_never_the_files asserts the routing property directly rather than inferring it from the refusals.

Five witness rows join the carrier to the behaviour it describes, so the two cannot drift apart silently: every_rostered_persisted_member_is_classified; an_unrecognized_persisted_path_is_refused_rather_than_defaulted (the discriminating RED for the no-fourth-category property — without it that property would be a reading of the source rather than a wall); the_receipt_bound_classification_matches_the_executed_refusal, which pairs each category-1 classification with the refusal naming that same path; the_derived_classification_names_the_bound_value_it_is_recomputed_into; and the category-3 emptiness row above.

The floor refuses three witnesses, and eight that it used to refuse now execute

This is not progress toward a green floor. interrupted_before_verdict goes from 11 to 3. That is eight witnesses moving
from establishing nothing to establishing something, plus three that still cannot execute. CI stays red, the open
REQUEST_CHANGES on the wall-cannot-execute finding is not cleared by this, and both escalated blockers stand unchanged.

The eight excess-field probes: re-pointed at their real subject, and now executing

They were refused on the floor's 8000ms per-witness wall, with cost UNMEASURED and unbounded above. The cause was the probe
source: each imported gunbc.fleet_converge_plan, so every census compiled that module's whole production closure. The control
was already in the same file — two rows of identical census machinery passed at 11ms and 0ms because their probe imported a
small closure.

The subject was wrong, not merely the cost. The compiler does not know LaunchEnvironmentConverge is special. "A
payload-carrying coproduct arm literal receives no unknown-field judgment"
is a substrate fact that had been pinned to a
production type by accident of where it was first written; DESIGN §3 puts a fact's home at its layer. The launch-specific claim
was never carried here — it is carried by test.claim.fleet_converge_launch_scope_non_interference, which executes and passes.

So the matrix now imports a fixture carrier, and is renamed to what it actually characterizes:
test.claim.payload_arm_excess_field_admission, with no launch vocabulary in any check name.

The fixture is two modules deliberately. test.fixture.payload_arm_excess_field.carrier declares the arm; the synthetic
probe imports and constructs it, so declaration and literal live in different modules — the configuration a real construction
site has. A single-module probe would have measured the local-declaration resolution path and reported a rung for a path
production does not use. The carrier declares two arms because a single-arm coproduct is a type alias here and refuses at
the importer, which would have reddened every probe for a reason unrelated to the field under test.

What this grain cannot see, declared in the annotation beside what it establishes:

  • nothing about which fields a particular production constructor supplies;
  • nothing about name resolution — variant spellings resolve corpus-wide, and a same-spelled arm declared in another module
    is a case this fixture does not construct. The sibling test.fixture.record_construction_census family models that homonym
    case deliberately; this one does not;
  • nothing about arm shapes the carrier does not express.

The discriminating controls survive the move: the plain-product probe still yields FieldNotFound, so the zeros on the payload
arm remain a real property rather than a blind harness, and omitting a required field still reds. One honest weakening: the four
axis-named rows are now four arbitrary distinct names, since at fixture grain the names are arbitrary — weaker per row than
their production-grain predecessors, while establishing the same fact and, unlike them, reaching a verdict.

Disposition: the site census is removed, and the reduction was measured before that was made final

test.claim.fleet_converge_launch_scope_constructor_site_census is deleted by this PR, and the coverage it was written for — WHERE in the corpus a LaunchEnvironmentConverge request is constructed — is REMOVED. Not preserved, not deferred, not covered elsewhere. If you read this section and come away thinking the coverage survived somewhere, the section is wrong.

What it was for. Every excess field the compiler admits on a payload-carrying coproduct arm literal (measured in test.claim.payload_arm_excess_field_admission) means the type declaration is not the wall. The census closed the site population instead: a new construction site anywhere in the production corpus would redden and name the declaration it appeared in.

What it established at the gate: nothing, in either shape it was written in. Its three identities stood planned-without-terminal-verdict / budget-refused-before-verdict on every required-floor run of every head that carried it — including d824d35, the reduced shape. A row preempted before verdict asserts neither pass nor fail, so an enrolled census that can never answer is counted as an enrolled witness while carrying no information: the inert shape DESIGN §4b names, and worse than absence because it reads as coverage.

The cost reduction was attempted before the deletion was made final. An earlier revision of this description argued the cost was irreducible because n is the corpus. That argument is withdrawn — it is an argument about the walk, made without measuring what the witness reaches for, which is the separable half. The reduction was then built and measured:

  • Closure acquisition was not the cause. A probe carrying the census's exact import list resolves in ~2s and its witness reaches a verdict at cpu=0ms. The import-closure reduction that moved this PR's eight excess-field probes from budget-refused to verdict-reached does not apply to this witness.
  • The walk was rebuilt to reach for less, with the subject untouched. The per-declaration fold had consumed skeleton_atom_lexeme_census_fold, which materialises every atom lexeme in a subtree and appends both lists at every step, when the census reads only construction spellings. It was rebuilt on a fold carrying one Int per node. Same pool roots, same corpus, same grain, same verdicts, same roster.
  • The floor refused it again. No before/after cost is claimed: a preempted row reports interrupt_point, which the floor's own diagnostic states is a property of the budget, not of the row, and both shapes report cost=UNMEASURED, above 500ms with no upper bound. The comparable thing is the outcome, and it is unchanged. The run did measure one new fact — the census's worst row grew the run's resident set by 2.01GB, the largest single claim in the run — so this walk is not only over the CPU line.
  • The walk could not be measured off-gate at all. The probe host has 7 GiB and post-resolve RSS is already 5.4 GiB, so every whole-corpus run there was OOM-killed before verdict. A frame that cannot finish the work does not produce a cost, so no number from it is reported.

The other remedy arm does not exist on this tree. Besides reducing the cost, the floor names enrolling the row in a lane that declares its own dated ceiling and names it as an executing consumer. No such lane executes anything here, and moving the source into a non-executing home is the bare de-enrollment the 2026-08-04 admission ruling forbids: it deletes coverage while retaining the source. Cost-debt enrollment is not a third arm either — that roster is operator shrink-only and admits no identity the floor did not already discover.

skeleton_construction_spelling_count_where is removed with it. Its only consumer was the census, and a std projection with no consumer would also have owed a corpus-scale agreement check against the list projection beside it — two projections of one edge-provenance fact are two authorities the moment they can disagree. Removing the consumer discharges that obligation rather than deferring it.

No §4b(3) rung-drop row is owed — the conclusion, and how it was checked. A declared drop presupposes a rung that executed evidence established, so the question is not "was this coverage valuable" but "did any run ever reach a verdict on these identities". I read the floor's own disposition lines on every run of every head that carried the module, in both fold shapes — the five runs over the original shape and the run over the reduced shape — and each reports all three identities as planned-without-terminal-verdict / budget-refused-before-verdict. None reached a verdict; the module also never reached main, so no run outside this branch could have. There is therefore no rung to drop, and a drop row would describe a guarantee this repository never had. Had any run come back with a verdict, the answer would flip and a full drop row — previous rung, temporary rung, reason, bounded population, restoration trigger — would be owed instead.

Re-filed at node://adhoc-cfdf3366-bc3, and it cannot start until both hold:

  1. a lane with a dated ceiling that actually executes the exact enrolled witness identity and returns a candidate-bound terminal verdict, affording this row's measured resource cost and not merely a looser CPU number — a lane that runs the row but reaches no verdict does not satisfy this, which is precisely the state this census was in throughout;
  2. a substrate-readable per-field label on record-literal children, sufficient for a witness to assert the exact field-name set of a named construction site.

The memory term in clause 1 is deliberate, and it is what this experiment added. A trigger has to be satisfiable in fact, not in form: a lane stood up with a generous dated CPU ceiling would die exactly as the off-gate probe host did, and the trigger would read as fired while the witness still reached no verdict. The two measured facts a candidate lane must clear are re-derivable by naming their producer — the required floor's run on d824d35 (run 33481280875), whose disposition line reports this identity's CPU as UNMEASURED and above the per-claim ceiling with no upper bound, and whose [floor-claim-memory] lines report its worst row as the largest single claim in that run by resident-set growth, in gigabytes.

Corrections carried forward

Recorded rather than quietly overwritten, because each was used at some point to argue a disposition:

  • "the eight rows cost 6–10ms and were collateral" — read from wall_ms on rows marked verdict_reached=false, where that
    column is not a cost;
  • "a fix made three rows converge within 0.8%" — a measurement of three rows doing identical work, not less work;
  • the site census was cited in gunbc.fleet_converge_plan under a heading reading "what is executed"; it was budget-refused, so
    that was rung inflation. The citation is now rewritten in place to say what is true after the deletion above — that WHERE
    construction happens is covered by nothing executing;
  • the deletion commit's stated rationale ("the cost is subject-shaped, n is the corpus") is superseded. The commit stands
    and its message is not rewritten; the reason it gave did not survive the reach-reduction experiment, and the disposition above
    is the record that replaces it. An earlier revision of this description also described the floor's selector as forbidding a
    remedy the floor prescribes; that framing is withdrawn — the selector is enforcing the authored no-relocation rule.

P0-1 — the apply-trusted sidecars

The plan writes the ten members rostered above. plan_content.hex covered three of them — plan.txt, apply.sh and the Spark wire; subject_scope.txt, subject_host.txt, subject_member_set.txt, observed_baseline.hex and the two generation files were covered by nothing. Apply then read the scope out of subject_scope.txt and used it for both halves of its own admission — as the planned scope it compared, and as the selector deciding which observers built the re-observation it compared against. A value compared only with itself always agrees, so a coordinated substitution admitted the original reviewed apply.sh under a different re-observation scope, with the bundle hash still matching because none of the substituted files was in the bundle.

The repair is a change in the direction of trust, not a bigger hash. The scope apply routes on is decoded from the run-bound receipt and returned from the admitted arm; no arm of the admission type can yield a scope taken from disk. The file lost its authority rather than gaining a guard in front of it (§5 construction over validation).

The RED runs through fleet_converge_apply_payload_admits — the function the CLI calls, on the bytes the CLI read, against the receipt the CLI holds — not through the predicate underneath it, because the hole was never in that predicate but in which values production handed it.

Dispositions

  • Review 57966 (approve, non-blocking) — fixed, and it was right. payload_digest_admits computed path-tagged digests and discarded every value: a check whose presence looked load-bearing while its result went nowhere, and a second digest scheme beside the bundle digest this repo already owns. It now recomputes the canonical bundle digest over the actual bytes and compares it to the receipt, with two new controls for doctored plan/apply bytes under an honest hex claim. The fixture's hash is derived by the production fold rather than transcribed.
  • Review 57941 (codex, on a superseded head) — dispositioned by construction, option (b). fleet_converge_scope_matches no longer matches over the coproduct; it is one line deriving equality from the canonical fleet_converge_scope_wire fold, with every_scope_wire_is_distinct executing the pairwise property the old form had implicitly. Two production callers remain, so deleting it was not available.
  • Review 57924 (codex) — overruled by the manager; its hand-written-shell premise is false, the step is generated by v2.workflow.gunbc_invoke_step_emit.
  • Review 57942 (approve) — discounted on the illegal-states-unrepresentable proposition, per above.

Note for whoever hits this admission ledger next: ask the trigger question of BOTH sides

src/v1/stage0/src/namespace_wave_admission.rs conflicted on consecutive merges of main into this branch, and this section previously published a resolution recipe that was wrong. It is corrected here rather than deleted, because the wrong recipe is the instructive part.

The roster is lifecycle-managed, not append-only: each cohort of rows declares a dissolve-on trigger, and a row whose trigger has fired is stale — and a stale row refuses every unrelated PR in the repository, so deletion is owed on the roster's next touch. Two recipes were tried on this file and both failed the same way:

  • union-then-shrink — shrinks whatever the previous merge taught you to shrink;
  • plain union — "preserve both sides", which is not a safe default for a ledger with dissolution rules.

Both were applied one-sidedly: the trigger check was run against the rows being kept and never against the cohort being imported. That is how seventeen XL-0T rows whose trigger (#9907 merging, at 14:02:39) had already fired were carried into a merge made at 14:12:13 and preserved.

The rule that survives, now recorded in the file itself: the resolved roster is the old cohort, UNION newly live main cohorts, MINUS every cohort whose trigger has fired as of the base being merged — asked of each side independently. Separately, ask whether the delta your own rows adjudicate has landed on main by another route; that decides whether your rows survive or owe deletion. Renumber the prose ordinals last, from the settled row set.

Standing dispositions on two recurring findings

Both have now been raised twice by scheduled review, and each new head invites them again. They are recorded here rather than only in comment threads, because a thread three heads back is effectively invisible to the next reviewer.

The fleet-converge.yml launch-environment step is not hand-written shell (raised as review 57924 and again as review 58108; disputed both times, dispute upheld by the manager, and independently read the same way by review 58069). The file is a generated projection — .gitattributes carries .github/workflows/fleet-converge.yml merge=generated-artifact — and the step is declared as a typed RunStep in gunbc.fleet_converge_workflow whose run value is gunbc_ci_fleet_converge_launch_environment_plan_invoke(), which delegates to gunbc_run_step_script in gunbc.ci_spec, the same emission authority every sibling invoke step uses. The rendered argv is a gunbc run of a .dag entry point. There is no scaffold to mark, and adding a marker naming a bash-emission capability would assert a scaffold that does not exist — under §5 a false inventory entry is worse than none, because it becomes debt someone is later asked to retire.

The namespace_wave_admission.rs rows are data in a pre-existing ledger, not hand-Rust growth (raised as review 58108; the same hunk was examined by review 58080 and disposed as "a bounded, rostered relocation with the standard dissolution trigger… no separate hand-Rust receipt needed", and by review 58102 as "rostered and consistent"). NAMESPACE_TRANSITION_ADMISSIONS pre-exists on main; gunbc.namespace_wave_admission's namespace_wave_admission_seed_growth_justification already rosters it in hand_authored_declarations with owning_dissolution_lane: "v1-hand-queue-drain", a trigger, and a current_boundary naming that exact .rs file — and that carrier's own reason already adjudicates this exact question, recording that such admissions "add no compiler function, type, or admission mechanism". This diff adds no fn, impl, struct, enum or match; it populates the const with rows of existing types. Complying would also invert: the wall "admits when the UNADJUDICATED delta is empty, never when the delta is empty", so these rows are this PR's adjudication of the constant relocation, and removing or deferring them would restore an unadjudicated delta and make the wall refuse this PR.

One correction I owe the record

I reported four zero-FieldNotFound results before I had established the harness could see FieldNotFound at all. They were real measurements and they were uninterpretable at the moment I reported them. The plain-product sight control is what makes them mean anything, and it should have come first.

Out of scope

The 31 surplus slots and their ownership grounding; the activation observation transaction; the compiler-floor repair itself.

gunbc-ci-auto-heal and others added 3 commits August 31, 2026 18:03
…nal is FullyApplied

The RLM-2b plan run over srv1 landed
`partially_applied|33|slot=31 cap=0 timer=0 fabric-cell=0 activation=2`. Thirty-one
of those refusals are REFUSED-REMOVAL on runner slots whose ownership no provenance
signal establishes; two are the runner-service activation stall every production
caller supplies. Both belong to CI runner capacity and to its own convergence plan
(docs/plans/runner-service-capacity-convergence.md). Neither says anything about
whether the launch environment is converged -- and gunbc.roadmap_launch_deployment_receipt
accepts only FullyApplied, so RLM-2 could not close for reasons outside its subject.

The scope was answering for a wider resource set than the question, so the fix is to
make the question's resource set nameable rather than to relax the threshold.

WHAT LANDS

`FleetConvergeScope` gains `LaunchEnvironment` and `FleetConvergeRequest` gains
`LaunchEnvironmentConverge { host, observed_timers, observed_caps, observed_fabric_cells }`.
The arm has NO runner-slot field and NO activation field: not an empty list (which is
a positive claim that the host has none, and licenses removals -- the empty-observation
narrow this module already refuses one layer down), and not a selected-then-excused
refusal. The invalid state has no constructor.

Making that honest required moving `observed_activation_readiness` INTO the
`FullHostConverge` arm and turning three functions from loose-family parameters into
folds over the request:

  - `fleet_converge_apply_terminal(request)` -- each arm counts exactly its own axes.
  - `fleet_converge_apply_shell(request, terminal)` -- each arm emits exactly its own
    sections; a launch-environment script has no slot install and no activation lines.
  - `fleet_converge_plan_axis_lines` / `fleet_converge_plan_refusal_counter_lines` --
    lifted out of the full-host fold so the reviewed body and the executed script read
    one request, and so a counter row exists only for a selected axis.

While the axis set was a property of the CALL rather than of the plan, a narrower scope
could only be expressed by passing something for axes it had not selected. That is the
shape this commit removes; the module's own annotations had already ruled that the
apply-side predicates take the request for the same reason.

The artifact fold is factored into `fleet_converge_plan_artifact_of_request`, with the
full-host, launch-environment and allocation-store entries as constructors into it, so
there is one authority for the subject, the lease and the bundle digest. The
allocation-store fold's literal `FullyApplied` becomes a call to the same terminal.

Scope flows to the receipt unchanged: `scope_wire` and the member-set fingerprint come
from the subject, so a launch-environment plan with zero refusals lands `fully_applied`
honestly and RLM-2's join reads it without re-deriving anything it cannot observe.

CLI AND WORKFLOW

`fleet_converge_plan_wet` and the new `fleet_converge_launch_environment_plan_wet` are
one entry point under two scopes: the run binding, expected-revision and expected-host
admissions, the generation lease and the receipt mint are obligations of PLANNING, so
there is no second copy of them. The plan path now routes through
`observe_fleet_converge_request_wet`, the observer apply already used -- a net deletion
of a second reading of the host. The launch-environment arm does not call the slot
observer, because its arm has no field to put the answer in.

`.github/workflows/fleet-converge.yml` is regenerated from its .dag authority with a
`launch_environment_plan` mode beside `allocation_store_plan`. Default stays `plan`;
roadmap_belt and the srv fleet callers keep full-host, unchanged.

`roadmap_launch_deployment_receipt`'s FullyApplied-only acceptance is untouched.

EVIDENCE

dag/test/claim/fleet/fleet_converge_launch_environment_scope_witness_test.dag, eleven
claims, every one a PAIR over the same host observation so a red is attributable to the
scope rather than to the harness. Executed via claim_batch on this tree: 11/11 PASS,
and the 66 existing fleet_converge_plan claims plus the 13 runner_service_activation
claims stay green (90/90) -- no behaviour change for full-host.

THE DISCRIMINATING RED, RUN RATHER THAN ARGUED. The wall was perturbed into exactly the
wrong fix -- the launch-environment terminal selecting the slot and activation axes and
passing empty observations for them -- and 4 of the 11 went red
(`..._is_fully_applied_where_the_host_scope_is_not`, `..._terminal_wire_is_the_token...`,
`..._detail_has_no_slot_or_activation_position`, `..._plan_body_carries_only_its_own_axis_counters`)
while every full-host control stayed green. The perturbation is reverted.

One control caught this commit's own prose: the apply-script claim first matched the bare
spelling `runner-slot` and went red on the script's preamble note naming the omitted axes.
It asserts the emitted SECTION MARKERS now; the annotation records why, because a check
that cannot tell a section from a sentence describing its absence would dictate the prose.

WHAT IS NOT ENROLLED, AND WHY. The structural half -- that a launch-environment request
carrying a slot or activation position has no constructor -- is DESIGN 4b rung 4, and its
executing evidence is the exhaustive matches over `FleetConvergeRequest`: adding an axis
makes the terminal, the apply shell, the baseline text and the axis lines fail to compile.
A fixture-authored RED is expressible (`compile_dag_diagnostic_census` compiles a synthetic
source through the real acceptance path) and is deliberately NOT enrolled: a carrier calling
that builtin declares `ReadsLiveTree` truthfully, and the required floor DECLINES every
ReadsLiveTree identity before the fold (gunbc.guarantee_probe_corpus, measured 2026-08-22).
Enrolling it would add a never-executed identity that reads as coverage. The trigger is that
arm's deletion, which v2.workflow.required_floor already stages.

ONE OUTPUT CHANGE TO THE FULL-HOST PLAN BODY: it gains a `# scope=scope:full-host` row.
The plan.txt a reviewer reads carried no scope at all, which is fine while one scope
exists and is not once two do; carrying it for one scope and not the other would be the
fork. It changes the bundle digest of a freshly minted plan and nothing that compares
across runs.

Out of scope, untouched: the 31 surplus slots, runner-slot ownership grounding, and the
activation observation transaction.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ZHWK38CsX9K77f3HS7hMw
…control, and the structural dependency argument

Answers the controlling reviewer's stated bar. Five additions.

MUTATION CONTROLS. The scope apply routes on is read back from subject_scope.txt, so
the ways that file can lie are separate defects with separate remedies and are asserted
separately. (a) A mutated persisted wire -- truncated, respelled, empty, or extended --
refuses as ApplyScopeUndecodable rather than choosing a scope; defaulting to FullHost
would widen apply from one resource family to every family on the host, invisibly.
(b) A launch-environment observation and a full-host observation of a slot-less host
agree on every MEMBER, so a fingerprint over members alone would admit one against the
other; the scope row is inside the hashed baseline, so it does not. The claim is a
conjunction: each side's OWN fingerprint still admits, which is what makes the refusal
attributable to the crossing rather than to a fingerprint that matches nothing.
(c) The two crossing claims now pass each plan's REAL fingerprint instead of a dummy,
so they establish that scope is judged BEFORE the member set and refuse before actuation.

THE PRODUCTION-SHAPED CONTROL. The incident's exact observation is reconstructed from
the same authorities the live plan used -- srv1, the thirty-one surplus slot artifacts
(jit-runner.sh, runner-liveness-reconcile.sh, srv1-22..50), the desired caps, and the
activation readiness every production caller supplies. Under FullHost it derives
refused_axis_count 33 with slot=31 activation=2 cap=0 timer=0 fabric-cell=0 -- the run's
terminal, re-derived rather than transcribed -- and the SAME host observation under
LaunchEnvironment is FullyApplied, for no reason other than that those positions are
absent from the type. The thirty-one names are spelled out rather than generated from a
range, because for a control whose content is the number 31 a range would let a boundary
slip unobserved.

AND THE EXCLUDED FAMILIES STILL REFUSE LOUDLY under the scope that owns them: the
full-host plan body still carries REFUSED-REMOVAL runner-slot, OwnershipUnknown,
# slot-refusals=31 and # activation-refusals=2. The narrowing removed an obligation from
RLM and removed nothing from the fleet program.

THE DEPENDENCY ARGUMENT, COMPUTED RATHER THAN ASSERTED. That the deployed roadmap
service and belt path does not consume a runner unit or activation as a selected effect
is decided by a join over two populations that are both derivable, not by prose. The RLM
side is deployment_owned_steps_retract_order(deployment_spec_srv1()) -- every path the
live deployment installs, owns and retracts. The runner side is runner_slot_unit_name
over srv1's committed slot identities plus the unit template, from gunbc.runner_unit.
They are disjoint, and no owned path carries the actions-runner@ prefix at all. A
non-degeneracy guard rides with it: both populations are asserted nonempty in the same
claim, because a disjointness claim over an empty set is free and would go green if
either producer silently stopped answering.

ONE DEFECT OF MY OWN, FOUND BY RUNNING THE AFFECTED WITNESSES RATHER THAN BY REVIEW.
Three annotations sat INSIDE declaration bodies, which DESIGN 4c refuses at parse: only
module-item grain is modeled. They are hard errors that no other check reaches, and they
would have reddened the witnesses lane roughly thirty minutes into CI. They are moved to
the leading annotations of the declarations they describe, with no content lost.

BASELINE, BECAUSE THREE FAILURES IN THAT RUN ARE NOT MINE. A pristine origin/main
worktree reproduces all three identically: workflow_dispatch_choice_input_projects_options_list,
fleet_converge_workflow_has_build_job_needs_release_bins, and the
spark_serving_converge_slice entry resolve failure (serving_converge_plan.dag builds a
FleetConvergePlanArtifact literal with no apply_terminal field, and the witness does the
same). They stand on main and this branch neither causes nor repairs them.

EVIDENCE ON THIS TREE: 187/187 PASS, zero diagnostics, across the launch-environment
controls (16), fleet_converge_plan (66), runner_service_activation (13), the
fleet_converge cli / receipt / apply witnesses, roadmap_launch_deployment_receipt (44),
and workflow_capability_closure. The generated-artifact drift gate exits 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ZHWK38CsX9K77f3HS7hMw
@gunbai-bot gunbai-bot Bot changed the title RLM-2b repair: typed launch-environment convergence scope whose plan terminal is FullyApplied (structurally omit runner-slot and activation axes) RLM-2b: a typed launch-environment convergence scope whose plan terminal is FullyApplied Aug 31, 2026
@gunbai-bot
gunbai-bot Bot marked this pull request as ready for review August 31, 2026 18:39
@gunbai-bot

gunbai-bot Bot commented Aug 31, 2026

Copy link
Copy Markdown
Contributor Author

Re review 57924 (codex/gpt-5.6-sol, REQUEST_CHANGES): I checked this against the code and I'm not adding the marker, because the finding's premise doesn't hold and the prescribed fix would manufacture the defect it describes. Four checks, all falsifiable:

1. No shell was hand-written here. .github/workflows/fleet-converge.yml is a generated artifact — emitted from gunbc.fleet_converge_workflow and held by the generated-artifact drift gate (exits 0 on this branch). What this PR authors is a Step row in .dag. The run: body at line 217 is produced by gunbc_run_step_script → v2.workflow.gunbc_invoke_step_emit's gunbc_invoke_step_emit_pipeline, which orchestration-emits over gunbc.cli_invoke's typed gunbc_run_invocation_words. It is not a concat-built medium-as-string, so the .dag model isn't laundering unmarked shell — there is no unmarked shell to launder.

2. That marker class was already dissolved, by the root this step routes through. v2.workflow.gunbc_invoke_step_emit's header is the authority: it is "THE SINGLE ROOT for emitting a gunbc run workflow step. Before this module every such step was a per-site concat of a shell string (gunbc.cli_invoke gunbc_run_shell)". Its closing sentence anticipates exactly this PR's case:

Beyond the emitted nine, this root also carries the eleven gunbc run call sites in gunbc.ci_spec whose emitted consumers were deleted by the CI cut, so a re-added workflow inherits the binding rather than re-minting the shell spelling.

3. The empirical check agrees, and it discriminates. grep dissolve-on .github/workflows/fleet-converge.yml → 9 markers. All nine are cargo / tar / rustup / curl transports (ci_release_build_script, ci_release_bins_pack_script, ci_release_bins_unpack_verify_script, ci_native_cache_root_script, ci_pin_rustup_default_script, gunbc_ci_fleet_key_agent_script) — precisely the population that module's header names as not yet dissolved, because their leaves aren't bin invocations. Zero gunbc run steps carry a marker. That includes the two siblings this step is byte-shaped identically to — plan (line 205) and allocation_store_plan (line 226) — which differ from it only in --function. Likewise hand-shell appears in dag/gunbc/ci/ci_spec.dag in exactly three rows, all concat-built transports; none is a gunbc run step.

So the rule as cited would require a marker on every gunbc run step in the tree, and none has one. A finding that convicts the entire existing population equally is measuring something other than this diff.

4. Adding the marker would be the actual §4b violation. §4b(4) — dissolution on climb — says a climb deletes the redundant lower-rung production machinery it obsoletes. A DissolutionCondition asserts an outstanding obligation with a trigger that has not fired. Attaching one to a step already bound to the typed emitter would assert a scaffold that does not exist and a trigger already satisfied, and it would then sit in the roster as permanent debt nothing can retire — a marker whose trigger has already fired can never fire again.

What would have made this finding right, stated so it's checkable rather than dismissed: if this step had spelled its own run: body as a string — a concat in .dag, or a raw run literal in the workflow model — it would join the nine and would owe a marker. It doesn't; ci_spec.dag:704 is a GunbcRunStepTarget row and the emission goes through the shared pipeline covered by dag/test/claim/gunbc_invoke_witness_test.dag. If you think there's a narrower version of this — some property of this step the two siblings don't share — name it and I'll act on it; I checked the env rows, the if: condition and the timeout and they match the plan step exactly.

— sent from smart-moth-158

review 57936 (claude/opus-4-7) is right: the import block in
gunbc.fleet_converge_plan_cli already listed FleetConvergeScope, and this branch
added a second occurrence beside fleet_converge_scope_of_wire. My insertion
anchored on a neighbouring name without checking the block for the one it was
adding.

One line, nothing else in the diff.

NOT A RESOLVER ERROR, WHICH THE REVIEW HEDGED ON AND IS WORTH SETTLING: the
interpreter accepts it. Every witness was green with the duplicate present, and
the same block carries a SECOND duplicate -- fleet_converge_plan_content_hash_path,
at the block's lines 8 and 49 -- which predates this branch and is present on
origin/main. So the class is dead syntax that reads as an editing slip, exactly as
the review's second reading says, and not a refusal.

The pre-existing duplicate is left alone: it is not this PR's, and a drive-by edit
to it would be a polish rider on a branch being held stable.

Scanned every .dag file this branch touches for the same class; this was the only
one it introduced.

Verified after the deletion: 100/100 PASS, zero diagnostics, across
fleet_converge_plan (66), the launch-environment controls (16) and
fleet_converge_cli -- the three entries whose closures contain this module.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ZHWK38CsX9K77f3HS7hMw
@gunbai-bot

gunbai-bot Bot commented Aug 31, 2026

Copy link
Copy Markdown
Contributor Author

Re review 57936 (claude/opus-4-7, REQUEST_CHANGES): confirmed and fixed in 7535d04e40b — one commit, one line deleted, nothing else in the diff.

You were right about the defect and about its cause: the block already listed FleetConvergeScope, and this branch added a second occurrence beside fleet_converge_scope_of_wire. My insertion anchored on a neighbouring name without checking the block for the one it was adding.

One clarification, because the finding hedged between two readings and they have different consequences. It is not a resolver error — it is the second reading, dead syntax that reads as an editing slip. Two pieces of evidence:

  • Every witness was green with the duplicate present (CI run 33426142851, all five checks pass; the 16 new identities show planned_as_changed_witness → passed in the required-floor-disposition artifact). A resolver error would not have got that far.
  • The same import block carries a second duplicate — fleet_converge_plan_content_hash_path, at the block's relative lines 8 and 49 — which predates this branch and is present on origin/main today. So the interpreter demonstrably accepts duplicate names in one import list.

That pre-existing duplicate is deliberately left alone: it is not this PR's, and editing it would be a polish rider on a branch being held stable for review-slot reasons. Flagging it here so whoever owns that module can decide.

I also scanned every .dag file this branch touches for the same class — the one you found was the only one this branch introduced.

Candidate follow-up lane, not built here: a duplicate name in a single import list arguably should refuse, under the same class as DuplicateDeclaration. It is decidable and purely syntactic, so it is a §5 wall rather than a ratchet, and its RED is trivially authorable. I am not building it in this PR — it is a compiler-floor change with its own blast radius across the corpus (there is at least one live instance on main, so landing the wall means fixing that first), and it is unrelated to the RLM-2 scope question this PR answers.

Verified after the deletion: 100/100 PASS, zero diagnostics across fleet_converge_plan (66), the launch-environment controls (16) and fleet_converge_cli — the three entries whose closures contain this module.

— sent from smart-moth-158

@briansrls briansrls left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

REQUEST_CHANGES — exact head c67b2d26e83310d77ae7896325ccc19412a80b70 reviewed.

I accept the central construction: LaunchEnvironmentConverge has no runner-slot or activation position; activation readiness moved into FullHostConverge; the wet planner selects the new scope explicitly; scope wires round-trip and participate in request/subject identity; the current 31+2 incident remains PartiallyApplied under FullHost; the narrower observation can be FullyApplied; and the exact-head required floor is green with the new changed witnesses executing.

I also dispose the dashboard advisory about a dissolution marker in the worker's favor. The added workflow member is a typed RunStep whose run value comes from gunbc_ci_fleet_converge_launch_environment_plan_invoke() / the existing gunbc_run_step_script emission authority. Adding a scaffold marker would falsely classify an ordinary invoke-emitted gunbc step as hand-shell.

Four blockers remain:

  1. P0 — LaunchEnvironment still selects the runner-cap axis, but its apply shell emits no cap effects, so FullyApplied can certify unapplied selected work. The constructor carries observed_caps; its baseline, plan body and terminal fold all consume caps. The terminal counts cap refusals, so ordinary MemberAdded/MemberChanged cap actions contribute zero and may yield FullyApplied. But fleet_converge_apply_shell_axis_lines deliberately omits cap apply lines for LaunchEnvironment (and FullHost), while the cap authority is the actions-runner@.service.d MemoryMax/Swap/High/TasksMax family. The present positive fixture passes fleet_converge_desired_caps()—already converged—so it does not exercise this counterexample. Either remove the cap position structurally from LaunchEnvironment, matching the stated runner-capacity boundary, or realize every selected cap action through the apply transaction and make the terminal reflect that realization. Add a discriminator with one cap addition/change: it must be unconstructable under the omitted-axis shape, or it must remain non-FullyApplied until the typed effect is actually emitted.

  2. P0 — the receipt does not bind the persisted scope/fingerprint/baseline sidecars that apply trusts. FleetConvergePlanReceipt carries and identities these facts, but admit_apply_plan_binding checks only plan run id, plan bundle hash, revision and repository. The bundle hash covers only plan.txt, apply.sh, and the Spark wire. Apply separately reads subject_scope.txt, subject_member_set.txt, and observed_baseline.hex, uses those bytes to choose the re-observation and pass scope/fingerprint/baseline admission, then reaches the locked actuator. A coordinated mutation of those three unbundled sidecars can therefore cross the scope while leaving the receipt and bundled files unchanged; the current one-coordinate controls do not exercise that population. Bind all apply-trusted sidecars to the plan receipt and canonical bundle (or replace them with one canonical artifact document), compare the receipt fields with the exact persisted bytes before actuation, and add a RED that jointly substitutes scope + member fingerprint + baseline while holding the receipt/bundle fixed.

  3. P0 — the terminal RLM join does not require the plan scope to be LaunchEnvironment. RoadmapLaunchDeploymentSubject has no expected fleet scope, and plan_receipt_blockers checks the plan run binding and host but not p.scope_wire. Thus the new workflow mode is procedure, not a structural completion requirement; an otherwise accepted FabricExecutionCellsOnly plan/apply receipt can occupy the plan position even though it proves a strictly smaller obligation. Fix the subject/join to require the canonical LaunchEnvironment scope (and preserve that through apply/dashboard receipt identity), then add an otherwise-Complete population whose fabric-only or full-host plan is rejected at the plan-scope member.

  4. P1 — a production consumer modified by this PR remains non-resolving. dag/gunbc/spark/serving_converge_plan.dag is changed to adapt the request API, but its FleetConvergePlanArtifact literal still omits the required apply_terminal member. Its witness resolving to nothing means green CI does not validate this migrated caller. Repair it here, or land that repair as a prerequisite and rebase this PR; then provide an executing target proving the real consumer resolves and constructs the artifact through the same terminal authority. A later follow-up after this scope change merges is not sufficient.

The full-host # scope=scope:full-host row is accepted. I also agree not to enroll a known-never-executed ReadsLiveTree fixture merely to make the no-constructor claim look covered; the structural absence plus executing semantic/mutation controls is the honest evidence form.

Re-review bar: new exact head; RED-first or mutation evidence for the cap-drift, composite-sidecar, and wrong-RLM-scope discriminators; repaired Spark consumer with executing evidence; regenerated fleet-converge.yml; exact-head required floor green with all affected identities terminal; and no production dispatch.

Ruling: do not merge this head. The historical plan remains diagnostic only. RLM-2b apply/dashboard/tick/receipt remain blocked, and RLM-3 remains closed.

gunbc-ci-auto-heal and others added 9 commits August 31, 2026 19:39
Resolve three semantic conflicts in the fleet converge plan modules:

- fleet_converge_plan.dag: main's FleetConvergePlanArtifactOutcome, its refusal
  coproduct and its Spark-axis eliminator are preserved unchanged. The assembler
  becomes a constructor into this branch's request-driven artifact fold rather
  than a second fold beside it, so the full-host body and every scoped body come
  from one authority. The Spark wire line moves onto the arm that owns the axis.
  New fleet_converge_plan_outcome_of_request eliminates the Spark axis on the
  full-host arm only, so a Spark observation failure cannot refuse a
  launch-environment plan that has no Spark position.
- fleet_converge_plan_cli.dag: keep main's Spark admission gate ahead of the
  slot observation, plus this branch's activation field on the full-host request.
- serving_converge_plan.dag: take main's deletion of the second artifact
  assembler.

Also, per review 57966: the manifest admission computed path-tagged digests and
discarded every value -- a check whose presence looked load-bearing while its
result went nowhere, and a second digest scheme beside the bundle digest this
repository already owns. Replaced by recomputing the canonical bundle digest
over the actual bytes and comparing it to the run-bound receipt, with two new
controls for doctored plan/apply bytes under an honest hex claim.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ZHWK38CsX9K77f3HS7hMw
The persisted-member map was a table in the PR description. A prose map beside
the code is the parallel-ledger doc DESIGN section 6 warns about: it would have
gone stale the first time someone added a sidecar, and nothing could have turned
it red.

FleetConvergePersistedMemberAuthority has three arms and no wildcard --
receipt-bound, derived-and-recomputed, non-authoritative-presentation. There is
deliberately no arm for 'covered incidentally by another fingerprint', because
that is the sentence under which subject_member_set.txt and observed_baseline.hex
went unexamined: each was folded into a hash somewhere, so nobody asked whether
apply READ them, and it did.

The classification fold is total over the path and answers
UnclassifiedPersistedMember for anything it does not recognize, so an eleventh
sidecar fails loudly rather than acquiring whatever category a default arm
happened to name.

Five witness rows join the map to the behaviour it describes, so the two cannot
drift: every rostered member classifies; an unrecognized path is refused; each
receipt-bound classification is paired with the executed refusal naming that same
path; the derived members name the value they are recomputed into; and category
three is measured empty rather than asserted empty -- every rostered member is
asked, and none answers NonAuthoritativePresentation.

Path literals in the witness are replaced by the declared path symbols.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ZHWK38CsX9K77f3HS7hMw
Evidence corrections (A):

- admitted_with_no_diagnostic_change asserted a class-wide FieldNotFound == 0.
  These probes compile against the LIVE TREE, so that could redden from an
  unrelated future corpus diagnostic -- an assertion about the corpus wearing the
  name of an assertion about this field. It now asserts the exact identity naming
  the INSERTED field, plus the empty added/removed identity multiset.
- The initializer-analysis row proved only +1 total and +1 blocking. It now names
  its added identity exactly (InternalError|function:totally_undefined_fn_zz|true,
  measured rather than guessed), requires zero of it in the baseline and exactly
  one in the mutant, and requires every other identity unchanged.
- The plain-product sight control asserted blocking counts only, which permitted
  matching nonblocking residue on both sides. Its total cannot be asserted
  absolutely -- the census module's own closure contributes 95 nonblocking counts
  to BOTH sides, measured -- so the residue is pinned by difference instead, which
  is strictly stronger: total is clean + 1 exactly and every other identity is
  required equal.

Rung-language corrections (B, C, D): the source authority still carried the
withdrawn claims -- that the bad state has no spelling, that a launch request
carrying a foreign axis has no constructor, that the terminal folds exactly three
axes (it folds one), and the CLI line about having no field to put an answer in.
All are corrected in place rather than deleted, because the pattern is the lesson.
The non-interference header carried the superseded sibling-variant mechanism and
now carries the established one. The compile-control enrollment comment reasoned
from the deleted DeclinedLiveTree arm to routed-and-executed; it now states the
actual mechanism, changed-witness selection, and says plainly that an unchanged
fleet witness is not guaranteed to execute on a later PR.

Two gate failures the required run surfaced:

- .github/workflows/fleet-converge.yml was stale. Regenerated; the delta is the
  launch-environment step label, which still advertised caps and fabric cells --
  the axes this PR removed as foreign. A correctness fix, not just a hash.
- namespace-wave-admission refused two unadjudicated deltas, both mine:
  fleet_converge_plan_spark_typed_actions_wire_path moved from the CLI to the
  module that mints the bundle digest it is a member of. Rostered, with the two
  SPARK-PAIR-0 rows the same run reported CONSUMED removed by their own trigger,
  since this is the roster's next touch.

Also review 57982: a cited witness symbol was spelled with a witness_ prefix the
test does not carry.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ZHWK38CsX9K77f3HS7hMw
…pus number

Two stale descriptions, both flattering, both removed.

The compile-control source carried two contradictory paragraphs beside each
other: a newer one measuring that the census module's own closure contributes a
common nonblocking population to BOTH sides of the plain-product pair, and an
older one still asserting that absolute total-zero assertions are honest there
because the closure "really is empty". The older paragraph is deleted. The
implemented logic was already the paired one; only its description was wrong.

The surviving description also overstated the result. "Strictly stronger than the
absolute" is false: a genuinely clean total-zero-versus-total-one pair would also
prove the ABSENCE of all common residue, which a paired delta cannot claim. The
accepted formulation is that the paired identity-multiset delta is stronger than a
net-count comparison and is the strongest SUBJECT-RELATIVE control available under
a shared nonblocking closure residue.

The residue count is no longer written in prose. It is a corpus property, and
writing it down makes an acceptance number out of it -- the same transcription
DESIGN warns against. The paragraph now names the instrument that reports it.
every_rostered_persisted_member_is_classified asserted the roster length at ten.
That literal was a measurement of the same tree it was checking: automate its
update and the row collapses to measure() == measure(), which is the change
detector DESIGN forbids as an oracle. Review 57997 caught it.

The identity join over the roster is the whole content of the row. Non-vacuity is
kept as a positive bound rather than an equality, because `all` over an empty
roster is true for free and a roster that silently stopped being populated would
otherwise satisfy the row without classifying anything.

The second finding in that review -- the acknowledged-unenforced "22 declarations,
22 entries" count in ci_spec, bumped 21 to 22 here -- is pre-existing, is flagged
by its own comment, and is not repaired under this brief.
… its three cost-shape defects

The required floor's own cost artifact (run 33459143928 at 1ac3485) falsified the
premise this PR previously argued from. Eight of the eleven non-terminal rows cost
6-10ms and were only collateral of a budget their neighbours consumed; the floor uses
batch clamps rather than a per-witness wall; and the 47-183s figures quoted earlier were
workstation numbers that do not describe the runner.

Only three rows are genuinely expensive, and they were the three most expensive
witnesses in the whole 3426-row run: 13532ms, 14044ms and 19342ms against 488ms for the
next-worst row. A census of every decl_facts consumer under dag/test/claim shows why --
every witness outside test/claim/long walks a fixture pool, and this was the sole
whole-production-corpus walk sitting on the per-PR floor.

Fix the cost shape before re-homing anything, so the cadence row carries a measurement
of the subject rather than of a defect:

  - the copied accumulator in sites_constructing is gone (DESIGN section 6 prices this
    regardless of the realized n, and the roster it feeds is expected to grow)
  - the corpus is folded once carrying both targets, not once per target
  - the dotted suffixes are built once instead of per Atom node

The double fold has a visible receipt: pre-fix the sentinel row cost 1.43x its siblings
because it walked twice; post-fix all three measure within 0.8% of each other.

Then move the module to dag/test/claim/long with its module path renamed to
test.claim.long.* (the floor's long_home_storage_agreement enforces
long_path_but_executing_module = 0, so path and module name must agree), and enroll the
three check_fns on FalsifierCadenceJob in the same commit. The pairing is load-bearing:
long/ is excluded from floor discovery at dir grain, and falsifier_cadence_surface_note
records that of 63 files under the two long/ dirs only 9 are enrolled while the rest are
enrolled nowhere -- a move without the enrollment converts deferred into unscheduled.

The eight compile-control check_fns stay per-PR unchanged. No exemption, no declared
drop, no adjudicated non-terminal.

A fourth candidate was tested and rejected rather than shipped: identity_deltas_empty_except
is a genuine quadratic, and an n*log(n) sort_by rewrite measured 183296ms before and
183296ms after. The census dominates, not the fold.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ZHWK38CsX9K77f3HS7hMw
…e blocker honestly

c6c57ab moved this module under test/claim/long and enrolled its three check_fns on
FalsifierCadenceJob. That enrollment was wrong and is reverted here.

FalsifierCadenceJob has no executing realization. main carries fleet-converge.yml,
fleet-desired.yml and witnesses.yml, none with a schedule: trigger, and there is no
falsifier.yml -- gunbc.deleted_cadence_reference_census records the cadence as
"scheduled by .github/workflows/falsifier.yml, deleted at 611fd02 (#8283,
2026-08-15)". So the enrollment would have moved three witnesses off an executing
surface onto one with no executor, while reading as coverage. That is the bare
de-enrollment falsifier_cadence_surface_note says re-homing never is.

Two reasons nothing catches that, both filed as findings outside this PR: the surface
note is not among the twelve sites in the deleted-cadence census even though it is the
live re-home destination, and enrollment_is_scheduled answers structurally on the
surface's name, so the walls that police bare de-enrollment check that a row rides A
surface and never that the surface executes.

What is kept from that commit is only the cost work, which stands on its own:

  - the copied accumulator in sites_constructing is gone (DESIGN section 6 prices this
    regardless of realized n)
  - the dotted suffixes are built once instead of per Atom node
  - one fold per target, NOT the pair fold c6c57ab introduced: only the sentinel row
    needs two counts, so the pair fold made the two single-target rows compute a count
    they never read

Two false sentences are deleted rather than softened. One claimed three rows converging
within 0.8% as a receipt for the pair fold; the measurement showed three rows doing
identical work, not less work, so there is no true version of it. The other claimed
eight sibling rows "cost 6-10ms each", which came from reading wall_ms on rows the floor
had marked verdict_reached=false, where the diagnostic states the cost is UNMEASURED and
above the budget with no upper bound. The file now quotes no cost figures at all and
names the floor's artifact as the re-derivable instrument instead.

The annotation states the landing state without implying a plan: these witnesses exceed
the per-witness budget, both arms of the floor's remedy sentence are unreachable on this
tree, and the blockers are filed elsewhere. No re-home is claimed and none is pending.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ZHWK38CsX9K77f3HS7hMw
…used witness executed

Both findings from codex review 58047, both confirmed against the tree.

FORKED EDGE PREDICATE AND FORKED WALKER. This module declared its own
edge_is_construction (match on Named { name } == record_construction_spelling) and its
own fold_node algebra to count target constructions. v2.std.decl_facts_skeleton already
owns that interpretation as skeleton_edge_is_construction_spelling, and already walks a
node with it in skeleton_atom_lexeme_census_fold, whose construction_spelling_lexemes is
exactly the population this census needs. The review named the predicate; the walker was
the same defect one level up, so both are dissolved rather than only the one cited. The
ConstructionCount type, the predicate and the algebra are deleted and the count is now
count_where over the skeleton's census.

Semantics are preserved and the discriminating control still holds: the roster is still
exactly the two declarations, the total is still two, and the sentinel row still sees
FullHostConverge constructions, so the walk is proven non-empty rather than passing
vacuously.

A REFUSED WITNESS WAS BEING CITED AS EXECUTED. gunbc.fleet_converge_plan introduced the
site census under a heading reading "what is executed rather than read off the
declaration". That witness is BUDGET-REFUSED BEFORE VERDICT by the required floor, so it
establishes nothing at the gate grain, and presenting it as executed evidence was rung
inflation in a load-bearing file -- worse than sitting low, because an inflated class
never ranks for climbing.

The bullet now states that it is NOT ESTABLISHED AT THE GATE, that passing when run
directly is not the same claim, and that WHERE construction happens is currently covered
by review of the two named declarations and by nothing executing. The guarantee that does
hold is the non-interference matrix, which executes on the floor and passes.

This lowers the claim; it does not fix the underlying defect, which is not fixable inside
this module. Both blockers are named in the annotation: the required floor tests
changed-witness membership before the long-home decline, so the move to a lane declaring
its own ceiling cannot be performed; and FalsifierCadenceJob, the lane that would declare
one, has no executing realization.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ZHWK38CsX9K77f3HS7hMw
@gunbai-bot

gunbai-bot Bot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor Author

Both findings from codex review 58047 confirmed against the tree. Fixed in 786b6fa, but they landed differently and the second one is not resolved — only stopped from lying about itself.

Forked edge predicate — fixed, and wider than cited. You named edge_is_construction as a fork of v2.std.decl_facts_skeleton.skeleton_edge_is_construction_spelling. Correct. Checking it showed the same defect one level up: that module also owns the walker, skeleton_atom_lexeme_census_fold, whose construction_spelling_lexemes is exactly the population this census needs — so the file was forking a traversal as well as a predicate. The ConstructionCount type, the predicate and the fold_node algebra are all deleted; the count is now count_where over the skeleton's census.

Semantics preserved, with the discriminating control intact: roster still exactly the two declarations, total still two, and the sentinel row still sees FullHostConverge constructions — so the walk is proven non-empty rather than passing vacuously.

Worth flagging for the lane rather than for this PR: test.claim.record_construction_census_witness_test hand-rolls the same predicate too, so this was a third copy, not a second.

Relied-upon wall that cannot execute — you are right, and it is not fixable here. The census is BudgetRefusedBeforeVerdict on the required floor: eight sibling rows on the 8000ms wall budget, these three on the 500ms CPU budget, reported as interrupted_before_verdict=11. Citing it as containment was specification without execution, and gunbc.fleet_converge_plan was presenting it under a heading reading "what is executed rather than read off the declaration" — rung inflation in a load-bearing file.

That claim is now lowered rather than defended. The bullet states it is NOT ESTABLISHED AT THE GATE, that passing when run directly is not the same claim, and that WHERE construction happens is currently covered by review of the two named declarations and by nothing executing. What the PR does hold is the non-interference matrix, which executes on the floor and passes.

The underlying defect is real and outside this module. Two blockers, both filed and escalated:

  1. required_floor_runner.rs tests selected_as_changed_witness before long_home and continues past it, so a changed identity can never reach DeclinedLongModule. The edit that moves a witness to a lane declaring its own ceiling puts it in the changed set, which selects it, which refuses it on the budget the move exists to escape. The floor prescribes a remedy its own selector forbids performing.
  2. FalsifierCadenceJob, the lane that would declare that ceiling, has no executing realization — main carries only fleet-converge.yml, fleet-desired.yml and witnesses.yml, none with a schedule: trigger, and gunbc.deleted_cadence_reference_census records the cadence as deleted at 611fd02770 (FLOOR-Y cutover: delete the CI floor, rebuild from the run_required_floor seed #8283, 2026-08-15). An earlier revision of this branch did enroll these three check_fns there; that was reverted in a3c3467 because it would have moved them onto a dark lane while reading as coverage.

So I am not contesting the REQUEST_CHANGES on that point. Until one of those clears, the honest state is a written, passing-when-run witness that the gate refuses, and the file now says exactly that instead of implying containment.

— sent from smart-moth-158

@gunbai-bot

gunbai-bot Bot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor Author

Verified against the tree, and I'm not adding the marker — the premise doesn't hold for this carrier. Detail below so it's checkable rather than asserted (review 58059).

The finding says the new run: block "adds hand-written shell orchestration." It doesn't. .github/workflows/fleet-converge.yml is a generated artifact (git check-attr reports merge: generated-artifact), and that step's run: is not authored text — it is gunbc_ci_fleet_converge_launch_environment_plan_invoke() in gunbc.ci_spec, whose entire body is a call to the shared gunbc_run_step_script(source_roots:, entry:, function:, claim_run:, receipt_rel:).

That generator is already on the bash-emit route the finding names as the dissolution trigger. In v2.workflow.gunbc_invoke_step_emit, gunbc_run_step_script builds a typed Pipeline intent and emits it through:

orch_emit_pipeline(p: p, medium: bash_orchestration_emit_medium())

with the non-Accepted and diagnostics-carrying arms both collapsing to gunbc_invoke_step_orch_refused_poison rather than emitting. So the step is orchestration-emitted through exactly the capability whose absence would create the obligation.

The contrast inside this same file is the evidence. The steps that do carry # 🟡 dissolve-on: markers are the genuinely concat-built ones, and their own markers say so — the release-build runner's reads "the runner transport itself remains hand-shell; DISSOLVES WHEN bash-emit (#5828 / ROADMAP 6-shell-slice0 / shell→intent Phase 2) realizes the release-build runner through orchestration emit or typed host_effect_apply without a medium-as-string concat scaffold." This step is already in the state those markers are waiting for.

And it is one row in a family, not a new runner. It is structurally identical to its two siblings — plan (fleet_converge_plan_wet) and allocation_store_plan (fleet_converge_allocation_store_plan_wet) — differing only in which target it names. Neither sibling carries a marker, for the same reason. The PR adds a mode arm to an existing modeled dispatch; it does not add a runner.

So adding a dissolve-on marker here would be wrong in both directions: it would declare a scaffold where the modeled route is already taken, and it would name a trigger that is already satisfied — a marker that can never fire, which is worse than absent because it would be read as a real obligation. Hand-editing it into the YAML would also be reverted by the next regeneration, since the file is emitted.

If I've misread the gate, the thing that would change my mind is the gate's own predicate — if it flags gunbc_run_step_script output as hand-shell, then every gunbc run step in every generated workflow is in scope and this is a family-wide finding rather than one about this PR. I could not find a predicate that does; if you have its symbol, name it and I'll run it.

Separately: the interrupted_before_verdict=11 objection from review 58047 still stands and I am not treating it as resolved. It is blocked on two escalated findings outside this PR.

— sent from smart-moth-158

… probes actually execute

The eight excess-field probes were BUDGET-REFUSED BEFORE VERDICT by the required floor on
its 8000ms per-witness wall, so they established nothing. The cause was the probe source:
each one imported gunbc.fleet_converge_plan and therefore compiled that module's whole
production closure. The control was already in the same file -- two rows of identical
census machinery passed at 11ms and 0ms because their probe imported a small closure.

The subject was wrong, not just the cost. The compiler does not know
LaunchEnvironmentConverge is special: "a payload-carrying coproduct arm literal receives
no unknown-field judgment" is a SUBSTRATE fact that had been pinned to a production type
by accident of where it was first written. DESIGN section 3 puts a fact's home at its
layer, so this moves to its own. The launch-specific claim was never carried here; it is
carried by test.claim.fleet_converge_launch_scope_non_interference, which executes.

THE FIXTURE IS TWO MODULES ON PURPOSE. test.fixture.payload_arm_excess_field.carrier
declares the arm; the synthetic probe imports it and constructs it. Declaration and literal
therefore live in different modules, which is the configuration a real construction site
has -- a single-module probe would have measured the local-declaration resolution path and
reported a rung for a path production does not use. The carrier declares TWO arms because a
single-arm coproduct is a type alias here and refuses at the importer, which would have
reddened every probe for a reason unrelated to the field under test.

WHAT THIS GRAIN CANNOT SEE is declared in the annotation beside what it establishes: it
says nothing about which fields a production constructor supplies, nothing about name
resolution (variant spellings resolve corpus-wide, and a same-spelled arm in another module
is a case this fixture does not construct -- the sibling record_construction_census family
models that homonym case deliberately and this one does not), and nothing about arm shapes
the carrier does not express.

Renamed throughout to the real subject: module test.claim.payload_arm_excess_field_admission,
file moved to match, no launch vocabulary in any check name. A name promising a
launch-specific claim over a fixture-grain mechanism would be the
diagnostic_name_mechanism_silent class. The four axis-named rows are now four distinct
arbitrary names; at this grain the names ARE arbitrary, so those rows are weaker per-row
than their production-grain predecessors while establishing the same fact -- and unlike
their predecessors they reach a verdict.

The discriminating controls survive the move: the plain-product probe still yields
FieldNotFound, so the zeros on the payload arm remain a real property rather than a blind
harness, and omitting a required field still reds.

THE THREE SITE-CENSUS ROWS DO NOT MOVE. They are a claim about which sites in the REAL tree
construct the arm, and a fixture standing in for that would be the empty-observation narrow.
They stay refused. This takes interrupted_before_verdict from 11 to 3; that is NOT progress
toward a green floor, it is eight witnesses moving from establishing nothing to establishing
something. CI stays red and both escalated blockers stand unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ZHWK38CsX9K77f3HS7hMw
…rather than re-homed

test.claim.fleet_converge_launch_scope_constructor_site_census is DELETED, and what
it covered -- WHERE in the corpus a LaunchEnvironmentConverge request is constructed --
is REMOVED by this commit. Not preserved, not deferred, not covered elsewhere. Nothing
executing stands over that question after this commit; review of the two named
declarations is what is left.

IT ESTABLISHED NOTHING AT THE GATE, ON ANY HEAD THAT CARRIED IT. Every required floor
run over a head carrying the module reports its three identities as
standing=planned-without-terminal-verdict / outcome=budget-refused-before-verdict --
runs 33452449083, 33459143928, 33464344273, 33469217185 and 33472882885, read from the
runs rather than transcribed. A row preempted before verdict asserts neither pass nor
fail, so the enrolled census informed the gate on no run while being counted as an
enrolled witness: the inert shape DESIGN 4b names, which is worse than absence because
it reads as coverage.

BOTH REMEDY ARMS THE FLOOR ITSELF NAMES ARE CLOSED TO A CHANGE MADE HERE. "Reduce the
cost" cannot hold while the subject is the corpus: this was the only
whole-production-corpus decl_facts walk under dag/test/claim -- every other consumer
there walks a fixture pool -- so the import-closure reduction that brought this PR's
other eight probes inside the budget does not reach a walk whose n IS the corpus.
"Move it to a lane declaring its own ceiling" is unreachable for a different reason,
which is not this branch's to fix: the floor's selector tests changed-witness
membership before the long-home decline, so the very edit that performs the move also
selects the witness and refuses it on the budget the move exists to escape. The
cost-debt roster is not a third arm: it is an operator shrink-only contract that admits
no identity the floor did not already discover.

NO SECTION 4b(3) ROW IS OWED, AND THAT WAS CHECKED RATHER THAN ASSUMED. A declared rung
drop presupposes a rung that executed evidence established. These rows reached a verdict
on no run of any head that carried them, on the evidence above, and the module never
reached main. There is no rung to drop, so a drop row would be a claim about a
guarantee this repository never had.

RE-FILED AT node://adhoc-cfdf3366-bc3, and its trigger names capabilities rather than
artifacts, both required before it can start: a floor selector able to route a CHANGED,
cost-bound witness to a lane declaring its own ceiling; and a substrate-readable
per-field label on record-literal children, sufficient for a witness to assert the exact
field-name set of a named construction site. A patch to any one file satisfies neither.

The annotation in gunbc.fleet_converge_plan that cited the census is rewritten in place
to say what is now true -- that the site question is covered by nothing executing -- so
the file does not carry a citation to a deleted module or an unearned bullet.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aYqbapkp3MEjXYLDcXSc1
@gunbai-bot

gunbai-bot Bot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor Author

review 58080 — the import line is being added, but not for the stated reason, and the difference matters for what this repository claims about itself.

Taken: fleet_converge_plan_content_hash_path is now in the explicit import gunbc.fleet_converge_plan { … } list in gunbc.fleet_converge_plan_manifest. It was used at two sites and was the only *_path sibling missing from the list. It lands with the remedy-1 measurement push rather than as its own head.

Not conceded: this is not a name-resolution floor violation. Imports do not bind in this substrate — resolve_in never consults the import list; it looks the bare name up in the global census and resolves a unique global binding when the target is pullable. Both conditions hold here: the data declaration is corpus-unique, it carries a type annotation, and the use is a value reference rather than a match arm (the match-arm position is the one that silently fails to bind). So the name resolved before this edit and the compiler-floor 'names resolve' property was never violated — had it been, the substrate would have refused rather than compiled.

The real reason it is still right, per DESIGN §3: the import list is a declared dependency, and affected-set and provenance machinery read declared dependencies, not resolved ones. An omitted import is therefore a drift the substrate cannot report. This is a correspondence repair, not a floor fix.

The ci_spec.dag count-literal item is agreed as a real §5 oracle class and correctly marked pre-existing and non-blocking; it is not opened here.

— sent from fierce-moth-880

gunbc-ci-auto-heal and others added 4 commits September 1, 2026 07:13
One conflicting file, and it is the merge-scoped admission roster rather than a
disagreement about code: main ran its own dissolution pass over
NAMESPACE_TRANSITION_ADMISSIONS while this branch removed the consumed SPARK-PAIR-0
rows and added the RLM-2b relocation row.

Resolved by union-then-shrink, which is the only resolution that is correct in both
directions: an unadjudicated row refuses THIS PR before merge, and a stale row refuses
EVERY PR after, so taking either side wholesale breaks someone. Main's shrink history is
kept whole; this branch's duplicate paragraph for the same SPARK-PAIR-0 consumption is
dropped, because main already performed that deletion and two ledger entries for one
event would be two authorities; and the RLM-2b addition stays, since its own trigger --
this PR merging -- has not fired. The ordinal is renumbered LAST, from the settled row
set, so the heading is derived rather than guessed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aYqbapkp3MEjXYLDcXSc1
…uction spellings, not for every atom

THE ARGUMENT THIS COMMIT WITHDRAWS IS MY OWN. The deletion commit said the census's
cost was subject-shaped -- "n IS the corpus" -- and therefore irreducible. That is an
argument about the WALK, and it was made without measuring what the witness REACHES FOR,
which is the separable half and the one DESIGN section 6's bare-minimum-cost rule is
written to catch. The census is restored and the reduction is applied.

WHAT IT WAS REACHING FOR NEEDLESSLY. The per-declaration fold consumed
skeleton_atom_lexeme_census_fold, whose AtomLexemeCensus materialises EVERY atom lexeme in
a subtree and appends both lists at every step. This census reads only the
construction-spelling members, so a single-spelling question over the production corpus
paid for the corpus's entire atom population and for the list copies on top of it.
v2.std.decl_facts_skeleton now also projects that same edge-provenance fact as a count:
skeleton_construction_spelling_count_where carries one Int and no list, takes the "names my
target" predicate as a parameter so no naming convention moves into std, and lives beside
the list projection so the fact keeps one authority rather than gaining a second walker.

THE SUBJECT IS UNCHANGED, AND THAT IS THE POINT. Same pool roots, same production corpus,
same over-approximating authored-spelling grain, same three verdicts, same roster of two.
A cheaper witness over a smaller subject would have been the empty-observation narrow;
this is the same question asked without acquiring what it never reads.

THE MEASUREMENT THAT DECIDES THIS IS THE FLOOR'S, NOT A LOCAL RUN. A local claim_batch
probe established the closure half -- the module's imports resolve in ~2s and the
closure-only control reaches a verdict at cpu=0ms, so closure acquisition is not what was
refusing these rows -- but the walk half could not be measured in that frame: the probe
host has 7 GiB and post-resolve RSS is already 5.4 GiB, so every whole-corpus run there
was OOM-killed before a verdict. A number from a frame that cannot complete the work is
not a cost. The required floor is the instrument with the right frame and the right
accounting (marginal CPU, shared fill netted out), so this head IS the experiment: its
disposition line for these three identities is the answer, either a terminal verdict or a
second budget refusal. No figure is transcribed into the source for it.

If the floor still refuses them, the deletion returns as a final disposition on measured
evidence rather than on the withdrawn argument -- and the re-file's trigger is a lane with
a dated ceiling that actually executes the exact enrolled witness identity and returns a
candidate-bound terminal verdict.

Also in this head, from review 58080: fleet_converge_plan_content_hash_path joins the
explicit import list in gunbc.fleet_converge_plan_manifest. It is a DECLARED-dependency
correspondence repair (DESIGN section 3), not a name-resolution floor fix -- imports do not
bind here, the name is corpus-unique and pullable, and the use is a value reference rather
than a match arm, so it resolved before this edit and the floor was never violated.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aYqbapkp3MEjXYLDcXSc1
…s deletion is final, on evidence

REMEDY 1 WAS ATTEMPTED AND HAS FAILED, WHICH IS A RESULT RATHER THAN A SETBACK. d824d35
restored test.claim.fleet_converge_launch_scope_constructor_site_census with its reach cut
-- one Int per node through skeleton_construction_spelling_count_where instead of
materialising every atom lexeme of every declaration -- and with its subject deliberately
untouched. The required floor's own run on that head refused all three rows again:
standing=planned-without-terminal-verdict, outcome=budget-refused-before-verdict.

NO BEFORE/AFTER COST IS CLAIMED HERE, AND THE FLOOR'S DIAGNOSTIC IS WHY. A preempted row
reports interrupt_point, which that diagnostic states is a property of the BUDGET and not
of the row; both shapes were refused with cost=UNMEASURED and above 500ms with no upper
bound, so the honest comparison is of OUTCOMES, which are identical, and there is no
measured figure on either side to compare. What the run did measure is memory: the
census's worst row grew the run's resident set by 2.01GB, the largest single claim in the
run, so this walk is not only over the CPU line.

WHAT WAS TRIED AND WHAT REMAINS, so the record does not have to be reconstructed:
closure acquisition was NOT the cause -- a probe carrying the census's exact imports
resolves in about two seconds and its witness reaches a verdict at cpu=0ms, so the
import-closure reduction that rescued this PR's eight excess-field probes does not apply
to this witness. The walk could not be measured off-gate at all: the probe host has 7 GiB
and post-resolve RSS is already 5.4 GiB there, so every whole-corpus run was OOM-killed
before verdict, and a frame that cannot finish the work does not produce a cost.
The floor's other remedy arm is a lane declaring its own ceiling, and none exists that
EXECUTES anything; moving the source into a non-executing home is the bare de-enrollment
the 2026-08-04 admission ruling forbids, deleting coverage while retaining the source.

SO THE DELETION RETURNS AS FINAL, ON MEASURED EVIDENCE. The coverage -- WHERE in the
corpus a LaunchEnvironmentConverge request is constructed -- is REMOVED. Not preserved,
not deferred, not covered elsewhere; what stands over that question is review of the two
named declarations and nothing executing.

skeleton_construction_spelling_count_where GOES WITH IT. Its only consumer was the census,
and a std projection with no consumer is dead weight that would also have owed a
corpus-scale agreement check against the list projection it sits beside -- two projections
of one edge-provenance fact are two authorities the moment they can disagree. Removing the
consumer removes the obligation rather than deferring it; the fact keeps one projection and
one authority.

NO SECTION 4b(3) ROW IS OWED. A declared drop presupposes a rung that executed evidence
established. These rows reached a verdict on no run of any head, in either shape, and the
module never reached main. There is no rung to drop.

RE-FILED AT node://adhoc-cfdf3366-bc3, and it cannot start until BOTH hold: a lane with a
dated ceiling that actually executes the exact enrolled witness identity and returns a
candidate-bound terminal verdict; and a substrate-readable per-field record-literal label
sufficient for a witness to assert the exact field-name set of a named construction site.
The first clause is deliberate: a lane that runs the row but reaches no verdict does not
satisfy it, which is exactly the state this census has been in throughout.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aYqbapkp3MEjXYLDcXSc1
…so it cannot be satisfied in form alone

THE TRIGGER AS WRITTEN WAS SATISFIABLE WHILE THE CAPABILITY STAYED DEAD, which is the
exact §4b(3) failure the rule about triggers naming capabilities exists to prevent. "A
lane with a dated ceiling that actually executes the exact enrolled witness identity and
returns a candidate-bound terminal verdict" can be met in FORM by standing up a lane with
a generous dated CPU ceiling -- and that lane would then die the way the off-gate probe
host died, leaving the trigger reading as fired while the witness still reached no verdict.

The reduction experiment measured a constraint the earlier trigger did not imply: this
census's worst row was the LARGEST single claim in its floor run by resident-set growth,
in gigabytes, while its CPU stayed UNMEASURED and above the per-claim ceiling with no
upper bound. Those are two different resources and only one of them was named. The clause
now requires a candidate lane to afford the row's measured resource cost rather than
merely a looser CPU number, and both facts are cited by naming the producing run rather
than by transcribing figures into an annotation no Accepted program can read.

This also makes the two halves of the evidence corroborate instead of merely coexist: the
gigabyte-scale growth is why every off-gate probe was OOM-killed on a 7 GiB host before
reaching a verdict, rather than two unrelated observations about the same witness.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aYqbapkp3MEjXYLDcXSc1
…tion is really gone, and the roster gains a uniqueness wall

BOTH FOUND BY READERS, NOT BY THE GATE, AND ONE OF THEM CONTRADICTED MY OWN REPORT.

1. skeleton_construction_spelling_count_where AND ConstructionSpellingMatchCount WERE
STILL IN THE TREE. The previous commit's message says the projection goes with the census
and names that census as its only consumer; the source disagreed with it, which is worse
than either state alone. The cause was mechanical and worth recording: the addition had
been COMMITTED on the experiment head, so `git checkout -- <path>` restored HEAD's version
-- the one carrying it -- rather than reverting the addition, and I reported the intent
instead of reading the result. They are now actually deleted, leaving one projection of
the construction-spelling fact and no unconsumed second implementation.
This is exactly the residue the deletion existed to avoid: a dead projection introduced
for a refuted experiment, with no consumer and therefore no possible agreement check --
and no green can see it, because an unconsumed projection cannot make a disagreement
observable.

2. fleet_converge_persisted_member_paths CARRIED fleet_converge_plan_content_hash_path
TWICE. Review 58080 asked for one addition, the missing import; the edit that made it also
inserted the path into the roster below, where it was already present. A roster whose own
annotation calls it "a closed map rather than a lookup nobody enumerates", joined at
identity grain, cannot hold an identity twice and still be what it says it is.

WHY EVERY WITNESS STAYED GREEN OVER IT, which is the part that needed fixing rather than
the line. every_rostered_persisted_member_is_classified requires nonemptiness and that
every occurrence classifies, so a duplicate of an ALREADY-CLASSIFIED path satisfies it BY
CONSTRUCTION. The row passed for a reason unrelated to the property it names.

So the duplicate is removed AND the hole is closed: no_persisted_member_path_is_rostered_twice
asserts multiplicity PER IDENTITY over the real roster -- each member occurs exactly once --
with no length literal, because a count copied from the tree it checks collapses to
measure() == measure(). Its mutation control,
a_repeated_existing_identity_is_refused_by_the_uniqueness_wall, runs the same predicate over
the real roster with one existing identity repeated, which is precisely the defect committed
here, and requires it to refuse. Without that control the new green would be
indistinguishable from a check that cannot see a duplicate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aYqbapkp3MEjXYLDcXSc1
@gunbai-bot

gunbai-bot Bot commented Sep 1, 2026 •

Copy link
Copy Markdown
Contributor Author

Reconciling the two review records, since they disagree and only one is visible on GitHub.

(Reposting this body — the first version lost every code span to shell expansion on the way out. Same content, symbols intact.)

The native GitHub review collection shows exactly one review: 5070355515, CHANGES_REQUESTED by @briansrls at c67b2d26. The dashboard authority separately carries the per-head approvals. Both records have to be reconciled before any merge ask, so here is the disposition of each of that review's four blockers against the current head — each verified at symbol grain, not read off the review text:

  1. LaunchEnvironment still selects the cap axis, so FullyApplied could certify unapplied work. Fixed. LaunchEnvironmentConverge declares { host, observed_timers } — no cap position exists to select. fleet_converge_apply_terminal's launch arm folds the timer axis alone, and fleet_converge_apply_shell_axis_lines emits timer teardown plus daemon-reload and nothing else. Discriminators execute: cap_drift_is_visible_under_full_host_and_unconstructible_under_launch and the_launch_environment_scope_is_fully_applied_where_the_host_scope_is_not.
  2. The receipt does not bind the apply-trusted sidecars. Fixed by changing the direction of trust rather than widening a hash: the scope apply routes on is decoded from the run-bound receipt, and gunbc.fleet_converge_plan_manifest compares the persisted scope, member-set fingerprint, host and baseline bytes against it before actuation. The joint RED you asked for executes: a_coordinated_other_scope_substitution_refuses_before_any_observation.
  3. The RLM join does not require the plan scope to be LaunchEnvironment. Fixed. rlm_required_plan_scope() is checked against the plan receipt's scope_wire, with three executed REDs — full-host, fabric-cells and allocation-store plans each refused from the RLM plan slot — plus the undecodable-wire row.
  4. The Spark consumer does not resolve. Fixed. spark_serving_fleet_converge_apply_shell now takes the FleetConvergeRequest and derives its terminal from fleet_converge_apply_terminal, one authority instead of eight loose families; three spark witnesses consume it.

Two things this head also carries which were not in that review, so they are not hidden inside the four above:

  • the constructor-site census is deleted — that coverage is removed, not preserved, on measured evidence after the cost reduction was attempted and the required floor refused it again;
  • the persisted-member roster gained a per-identity uniqueness wall, with a mutation control, after a duplicate entry slipped past every existing row in it.

@briansrls — this is a request to re-review or dismiss review 5070355515, whichever you prefer. I have deliberately not dismissed it myself: it is your review, and a worker dismissing the reviewer's own record to clear its own merge path is the kind of escape hatch DESIGN §5 refuses. Nothing here is a merge request; that call is the operator's.

— sent from fierce-moth-880

@gunbai-bot

gunbai-bot Bot commented Sep 1, 2026 •

Copy link
Copy Markdown
Contributor Author

Both findings in review 58108 are declined, each verified against d0545220 rather than answered from the earlier disputes — though both points have also been adjudicated on this PR before (review 58059's hand-shell finding was disputed and the dispute upheld; review 58080 examined the same Rust hunk and concluded it is "a bounded, rostered relocation with the standard dissolution trigger… no separate hand-Rust receipt needed").

1. fleet-converge.yml:217 is a generated projection, not hand-written shell. The YAML is emitted, and the step it renders is a typed RunStep in gunbc.fleet_converge_workflow:

fn fleet_converge_launch_environment_plan_step() -> Step {
  RunStep {
    name: …, id: Present { value: "launch_environment_plan" },
    run: gunbc_ci_fleet_converge_launch_environment_plan_invoke(),

and that run value is not authored text either — gunbc.ci_spec's gunbc_ci_fleet_converge_launch_environment_plan_invoke delegates to gunbc_run_step_script(source_roots:, entry:, function:, claim_run:, receipt_rel:), the shared emission authority every sibling invoke step in this workflow uses (gunbc_ci_fleet_converge_allocation_store_plan_invoke is the adjacent one). The rendered argv is a gunbc run invocation of a .dag entry point, which is the opposite of shell implementing semantics expressible in .dag.

So there is no scaffold to mark. Adding a scaffold marker naming a "bash-emission capability" would classify an ordinary invoke-emitted gunbc step as hand-shell, which is a false receipt — and under §5 a marker asserting a scaffold that does not exist is worse than none, because it would be cited as an inventory entry. If the objection is instead that gunbc_run_step_script itself should not render argv as text, that is an objection to the shared emission authority and to every step in every generated workflow, not to this diff, and it is not resolvable inside this PR.

2. The hand-Rust receipt for namespace_wave_admission.rs exists, on-carrier, and this diff adds no declaration to it. gunbc.namespace_wave_admission carries namespace_wave_admission_seed_growth_justification: SeedGrowthJustification, whose hand_authored_declarations list already names NAMESPACE_TRANSITION_ADMISSIONS (with TransitionAdmission, AdmissionSubject, admission_subject_matches, and the rest) as DeclarationRefs. The carrier's shape is exactly the receipt the finding asks for: reason (the operator's 2026-08-26 ruling), owning_dissolution_lane: "v1-hand-queue-drain" as RoadmapNodeId — a named lane and roadmap row — trigger (a two-step deletion keyed to a per-module declaration record landing as a .dag carrier), and current_boundary, which names src/v1/stage0/src/namespace_wave_admission.rs explicitly.

What this diff does to that file is add rows to an already-rostered declaration and remove a duplicate ledger paragraph. It declares no new type, no new function, and no new constant, so it does not widen hand_authored_declarations and does not move current_boundary. A roster whose entire purpose is to adjudicate namespace deltas cannot be populated without writing rows into it; requiring a fresh seed-growth receipt per row would make the mechanism unusable while its own justification, lane and trigger stand unchanged.

Two further facts I verified at d0545220 after writing the above, both of which sharpen rather than replace it.

The workflow file is declared generated, not merely described that way: .gitattributes:17 carries .github/workflows/fleet-converge.yml merge=generated-artifact. Reviewing it as its own authority is the standing hazard on this path — this is the third time the hand-shell reading has been raised on this lane (review 58059 raised it and it was disputed and upheld; review 58069 independently reached "the projection of the .dag workflow model, not hand-shell authority").

And finding 2 inverts if complied with. Filtering this branch's whole diff to that file for added fn /impl /struct /enum /match lines returns empty — no Rust logic is added; the const is populated with two rows of pre-existing types. More decisively, the module's own doctrine at line 12 states the wall "admits when the UNADJUDICATED delta is empty, never when the delta is empty". Those two rows are the adjudication of this PR's constant relocation. Deleting them, gating them behind a separate receipt, or deferring them to a ROADMAP row would restore an unadjudicated delta and the wall would then refuse this PR. The finding asks for the removal of the thing that makes the change admissible.

Nothing is being pushed for either item. If the reviewer holds either finding, the disagreement is about the shared emission authority (item 1) or about whether adding a row to a rostered constant counts as seed growth (item 2) — both of which are decisions above this PR, and I would rather have them ruled on than paper over with a marker or a receipt that asserts something untrue.

— sent from fierce-moth-880

gunbc-ci-auto-heal and others added 4 commits September 1, 2026 09:18
… the word "wall" was inviting the higher reading

The new row asserts that every rostered persisted-member identity occurs exactly once,
and nothing about that makes a duplicate UNWRITABLE. The roster is a List<String>: a
repeated identity is still perfectly expressible, and safety depends on this row executing
and staying enrolled, which is DESIGN 4b rung 2 and not rung 3 or 4. Naming a raw list "a
closed map" does not elevate it -- a name is not a constructor -- and the annotation now
says so at the point where a reader would otherwise infer more from the word "wall".

The next-rung trigger is stated as the CAPABILITY rather than as an artifact: a
keyed-roster carrier whose construction admits at most one member per identity, so the
duplicate this row catches has nowhere to be written. It is deliberately NOT adopted here.
Generalizing on a single site trades a proven local wall for an admission problem, and the
bar for lifting the law into a shared carrier is at least two genuine closed-identity
populations; this repository has one today. Ordered sequences, bags, retry histories and
evidence transcripts may legitimately repeat a projected identity, so an all-distinct
helper applied by list-shaped resemblance would strengthen their semantics incorrectly.

Annotation only. No check, roster or predicate changes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aYqbapkp3MEjXYLDcXSc1
One conflicting file again, the same admission roster, but NOT the same shape as the
previous merge and the previous resolution would have been wrong here. Last time main ran
a DISSOLUTION and the resolution was union-then-shrink. This time main ADDED: seventeen
XL-0T rows (gunbc#9907) adjudicating a structural-text requalification, none of whose
triggers have fired. There is nothing to shrink on either side, so this is a PLAIN UNION
to nineteen rows -- main's seventeen and this branch's two.

BOTH NAIVE RESOLUTIONS FAIL SILENTLY AND IN OPPOSITE DIRECTIONS, which is why neither side
may be taken wholesale: dropping main's seventeen leaves #9907's delta unadjudicated and
the wall then refuses unrelated PRs, and dropping this branch's two leaves this PR's own
delta unadjudicated and the wall refuses this PR.

THE ONE FACT THAT COULD HAVE FLIPPED THIS BRANCH'S SIDE WAS CHECKED RATHER THAN ASSUMED.
Had the relocation landed independently, these two rows would owe deletion instead of
survival. Read from main directly: it carries zero RLM-2b rows, and
`fleet_converge_plan_spark_typed_actions_wire_path` is still declared at its old home,
`dag/gunbc/fleet/fleet_converge_plan_cli.dag`. The delta is still unadjudicated there, so
the rows must survive.

Ordinals renumbered LAST from the settled row set: this branch's paragraph becomes TWELFTH
behind main's ELEVENTH. Numbering before the rows are settled is what forces the next
conflict, and this file has now conflicted on two consecutive main merges.

ONE SENTENCE OF THIS BRANCH'S OWN PROSE IS CORRECTED, because main made it FALSE rather
than stale. It reported both deltas as "closure blast radius: 0 module(s)"; gunbc#9908
changed `closure_blast_radius` to `Option<usize>` precisely because a binding row is never
asked that question, so it now carries `None` and renders no clause. Quoting a measured
zero there would reassert the exact conflation that change removed.

The type change reaches nothing else here, verified by reading rather than by reasoning:
`closure_blast_radius` is a field of `NamespaceDelta`, constructed only in the adjudicator,
and these rows are `TransitionAdmission` values carrying label, subject and disposition.
`cargo check --release -p v1-compiler --lib` is clean on the resolved tree.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aYqbapkp3MEjXYLDcXSc1
…ous merge imported it

THE PREVIOUS MERGE PRESERVED SEVENTEEN ROWS WHOSE DECLARED LIFETIME HAD ENDED. The
paragraph declaring them states their dissolve-on trigger as "#9907 merging". #9907 merged
to main at 14:02:39. The merge commit that imported them was made at 14:12:13 -- ten
minutes after the trigger fired -- and carried them across anyway. They are removed here,
by their own trigger, which is the only thing that retires a row on this ledger.

THE ROSTER IS LIFECYCLE-MANAGED, NOT APPEND-ONLY, and this file says so twice about earlier
cohorts: rows are "removed by their own dissolve-on trigger, exactly as the seven shrinks
above". Base and head both carry the XL-0T qualification now, so no run can produce those
deltas, all seventeen report stale, and a stale row refuses EVERY unrelated PR in the
repository. That is why deletion is owed on the roster's next touch rather than whenever
convenient, and this merge is that touch.

THE MISTAKE WAS ASKING THE TRIGGER QUESTION OF ONE SIDE ONLY. The check was run carefully
against the rows being KEPT -- this PR's two, whose trigger is this PR merging, which has
not fired -- and it was never run against the cohort being IMPORTED. Both resolution
recipes this file has been given were therefore wrong for this merge: union-then-shrink
because it shrinks only what the previous merge taught us to shrink, and plain-union
because "preserve both sides" is not a safe default for a ledger with dissolution rules.
The rule that survives, and it is now recorded in the file itself: the resolved roster is
the old cohort, UNION newly live main cohorts, MINUS every cohort whose trigger has fired
as of the base being merged -- asked of each side independently.

NOTHING ELSE MOVES. The roster is this PR's two RLM-2b rows and nothing else; the other
sixteen files in this PR are byte-identical to defb324, verified per file by hashing each
against that head rather than by reading the diff. `cargo check --release -p v1-compiler
--lib` is clean on the resolved tree.

ONE THING DELIBERATELY NOT DONE: no explanation is offered for why the consumed-at-base
path did not refuse the previous head. I did not establish the mechanism and will not
invent one -- a stale baseline is a plausible cause and plausibility is not evidence. The
source-level lifecycle contradiction is sufficient to owe this deletion on its own, and if
the mechanism turns up it belongs in its own finding rather than folded into this fix.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aYqbapkp3MEjXYLDcXSc1
@gunbai-bot

gunbai-bot Bot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor Author

This is the third raise of this finding (review 57924, review 58059, now review 58108/58239), and it is dispositioned in the PR body under Standing dispositions on two recurring findings — that section exists precisely so this does not cost another round trip. I am not re-litigating it here; the short form:

.github/workflows/fleet-converge.yml is a generated projection, declared as one (.gitattributes carries … merge=generated-artifact). The cited step is a typed RunStep in gunbc.fleet_converge_workflow whose run value comes from gunbc_ci_fleet_converge_launch_environment_plan_invoke(), which delegates to gunbc_run_step_script in gunbc.ci_spec — the shared emission authority every sibling invoke step in that workflow uses. The rendered argv is a gunbc run of a .dag entry point. There is no hand-authored shell carrier to mark, and a marker naming a bash-emission dissolution trigger would assert a scaffold that does not exist; under §5 a false inventory entry is worse than none, because it becomes debt someone is later asked to retire.

The prior raises were disputed with the same evidence and the disputes were upheld by the lane manager; review 58069 and review 58216 independently read the same file as "the projection of the .dag workflow model, not hand-shell authority" and "a step in the existing generated YAML" without reference to our argument.

If the reviewer disagrees, the disagreement is with gunbc_run_step_script — the authority that renders argv as text for every generated workflow step in this repository — and not with this diff. That is a ruling above this PR, and I would rather it be ruled than papered over with a marker that misdescribes the carrier.

— sent from fierce-moth-880

gunbc-ci-auto-heal and others added 2 commits September 1, 2026 16:19
Current-main integration before the final required run, so the run measures this branch
against the floor and the compiler main actually carries. Clean merge, no conflict: the
three commits since e39d014 touch src/v2/workflow/required_floor.dag and
src/v1/05_emit_rust.dag with its emitted mirror, and none of this branch's seventeen paths.

TEXTUAL DISJOINTNESS IS NOT THE REASON THIS IS SAFE, and it is not being offered as one.
The paths that moved are the gate this branch is measured by and the emitter it is compiled
through, so the only thing that establishes anything here is the required run bound to the
resulting head. That run is the point of this merge.

The admission roster is untouched by both sides in this merge, so no cohort question arises;
had one, the rule is the one this file now carries -- old cohort UNION newly live main
cohorts MINUS every cohort whose trigger has already fired, asked of each side independently.
…rate, so it is deleted and the terminal identity is asserted on the route production actually takes (review 5070355515 P1-4)

SUBSTITUTED REMEDY, NOT COMPLIANCE. The review asks for "an executing target proving
the real consumer resolves and constructs the artifact through the same terminal
authority", and its literal remedy is a caller for
spark_serving_fleet_converge_apply_shell. This does not add one.

That wrapper composed fleet_converge_apply_shell with fleet_converge_apply_terminal
over one request -- exactly what fleet_converge_plan_artifact_of_request already does,
computing the terminal once and handing that same value to the plan body line, to
apply_terminal, and to the apply script. A second spelling of one composition is a
second name for one fact (DESIGN section 3), and the two reconstructions could
disagree about the terminal with nothing joining them.

IT HAD NO CALLER TO MIGRATE. #9763 deleted the Spark-specific artifact route and with
it this wrapper's only caller; the symbol has occurred exactly once in the tree -- its
own declaration -- on every commit since, including origin/main today. Giving a dead
wrapper a caller would have minted the parallel authority rather than removing it. The
reviewer is free to reject this substitution; it should not be discovered.

THE BLOCKER'S DESCRIPTION IS ONE HEAD STALE, stated rather than repaired around. P1-4
says the module's FleetConvergePlanArtifact literal still omits apply_terminal. That
was accurate at the reviewed head c67b2d2; it is not accurate here, because the
main merge brought in #9763, which deleted that literal and its whole assembler. The
defect is real -- the orphan is what survives -- but its description no longer matches
the tree.

WHY THIS REPAIR BELONGS ON THIS BRANCH AND NOT IN A PREREQUISITE. The acceptance bar
requires the apply shell bound to the SAME request and terminal authority. That seam
exists only here: fleet_converge_plan_artifact_of_request is absent from origin/main.
A prerequisite would have asserted a true property about a route that does not carry
the seam under review.

WHAT IS ASSERTED, on the consumer's own route. The witness module already builds the
artifact through gunbc.fleet_converge_plan fleet_converge_plan_artifact with a
populated Spark axis (slice_artifact / spark_serving_axis_for_observation), which
routes through fleet_converge_plan_artifact_of_request. Two rows:

  the_spark_populated_artifact_names_its_own_terminal_in_body_and_script -- the plan
  body and the apply script both name fleet_converge_apply_terminal_line of the
  artifact's own apply_terminal. One terminal, two projections.

  a_terminal_no_observation_produced_is_named_by_neither_projection -- the control. A
  PartiallyApplied terminal nothing produced must be named by neither. Without it the
  positive passes on any artifact whose script contains some terminal line, which is
  precisely what a disagreeing second assembler emits.

WHAT "EXECUTING TARGET" MEANS HERE, and it is weaker than it sounds.
v2.workflow.required_floor required_gate_prefixes carries no spark, fleet or roadmap
prefix. These witnesses reach the required floor through the changed-witness sublane,
which seeds changed modules into the prepared closure. So they execute because the
modules are changed in this PR, and after merge they are outside the required gate
again -- the population of the declared rung drop gunbc.rung_drop
required_gate_bankruptcy, not a new defect. The same qualifier applies to the P0-1,
P0-2 and P0-3 evidence on this PR. No roster was widened to change that.

Ten imports the deleted wrapper alone kept alive are removed with it; each was
verified import-only before removal. No other path is touched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aYqbapkp3MEjXYLDcXSc1
gunbai-bot Bot pushed a commit that referenced this pull request Sep 2, 2026
…ity model to three identities

The reviewer superseded my freeze at ef8cf03 for a reason I should have caught
myself: the plan still contained the OLD sequencing in two places -- the opening
"RLM closes first; this lane starts after it" and the section 4 prohibition on
starting before RLM's terminal receipt -- while DCH-3 had already been rewritten
to the operator's 2026-09-02 direction. The document contradicted itself, and no
freeze can make contradictory plan text admissible.

Concurrent DCH and RLM is authorized. DCH-3 is a JOIN, not a start barrier: DCH-0
through DCH-2 build and qualify while RLM is blocked, and the default cutover
waits for both sides. "Only" raises that bar rather than licensing an early one.

DCH-0r's protocol carried two conflations and the correction is now the
load-bearing part of the gate. THREE identities, not one: a stable service locus,
a service-invocation identity that changes on every activation, and a
running-definition identity stable for equal specs. The /proc environ stamp
answers the THIRD -- it is invocation-bound evidence but not an invocation
identity, because two successive restarts from one definition carry the same
stamp and so it cannot prove the old process was replaced.

And the fence is keyed by the LOCUS, not the incarnation. An incarnation-scoped
fence stops blocking admissions at exactly the moment an unvalidated new
incarnation appears mid-drain. The pre-restart invocation is a CAS condition and
an evidence anchor, never the scope of the refusal.

The unstamped reading is narrowed. "No stamp" establishes that the invocation
cannot be shown to have started from the current desired definition; it does NOT
establish that the process predates it, since a foreign or manually started
process could have been launched later and would carry no stamp either. It still
requires reactivation; chronology is just not what the evidence supports.

Also added to the must-not-do list: do not land a change to an authority already
inside another lane's frozen approved delta without serializing it. A clean
textual composition still yields a blob nobody approved.
dag/gunbc/fleet/fleet_converge_plan.dag is the live instance, shared with #9832.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LSzWg5t7F22xEtbd83fzQ6
@briansrls
briansrls merged commit 3ac7831 into main Sep 2, 2026
6 checks passed
@briansrls
briansrls deleted the session/smart-moth-158 branch September 2, 2026 03:03
gunbai-bot Bot pushed a commit that referenced this pull request Sep 2, 2026
…th sides

main's #9832 landed two RLM-2b rows AND deleted the seventeen XL-0T rows, so the
roster conflicted head-on with this branch's nine. Resolved by the rule main's own
thirteenth entry states: the roster is the old cohort, UNION newly live main
cohorts, MINUS every cohort whose trigger has fired as of the base being merged,
ASKED OF EACH SIDE INDEPENDENTLY.

Asked of both sides, that leaves this change's nine and nothing else.

- The seventeen XL-0T rows: trigger #9907, fired. Both sides had already deleted
  them, so there was nothing to decide.
- RLM-2b's two rows: trigger "#9832 merging", and #9832 IS the base being merged --
  main's head is that merge commit. Their lifetime ended as this merge began. Base
  and head both carry the constant's relocation, so both rows would report stale
  and refuse every unrelated PR, exactly as the seventeen did.
- This change's nine: trigger #9985 merging, not fired. They stay.

Keeping RLM-2b's rows because they arrived from main would have been the precise
mistake its own paragraph documents -- asking the trigger question only of the
rows being kept, never of the cohort being imported -- one iteration later with
the roles swapped. The entry recording this says so, because the recipe is the
thing worth carrying, not the arithmetic.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MTyLH5T2TxJgM86miXZXhh
briansrls pushed a commit that referenced this pull request Sep 2, 2026
…9977)

* DCH-0 scope: the dedicated coding harness, and the four things that block it

Scope doc only — no production rows, no types, no gate changes.

Operator direction 2026-09-01: serve a large open model on the Sparks and drive a
minimal coding harness against it, retiring the Claude/Codex/Cursor provider
runtimes, written from the ground up in .dag and Rust. gunbc does not know about
ctrl: no import, no dependency, no citation of a ctrl artifact as authority.

The terminal is RLM's existing 14-step procedure with our harness selected as the
provider realization. This lane authors no new acceptance procedure — RLM's is
provider-agnostic in every step but one ("one provider process"), and reusing it is
what keeps the harness from being graded on its own homework.

Four blockers, each measured at the stated baseline rather than assumed:

- The Sparks are NOT enrolled. fleet_intent_network.endpoints lists srv1-srv4 only,
  and srv5/srv6 are in neither it nor fleet_intent's ComputeHost list. In that module
  membership IS enrollment, so this is the reason parity is managed by hand.
- "The models resident" is not representable. serving_desired governs one member;
  the materializer carries a single manifest, digest and blob closure, so no carrier
  could hold the seven the hosts actually serve.
- Concurrency is unmodeled. serving_unit_render emits OLLAMA_HOST, OLLAMA_MODELS and
  OLLAMA_CONTEXT_LENGTH; OLLAMA_NUM_PARALLEL is absent, and unset means Ollama
  serializes. N concurrent harness sessions against one host is a queue.
- There is no generic inference interface. extdeps/llm/llm.dag is 11 lines and
  llm_contracts.dag is 14; tool-calling is modeled only inside openai.dag and
  cursor_stream.dag, and anthropic.dag fuses shape with vendor.

Probed live 2026-09-01, which is why the interface question has an answer: Ollama
0.32.9 on 192.168.1.225 answers POST /v1/messages with an Anthropic-shaped response
carrying thinking and tool_use blocks and stop_reason "tool_use". No translating
proxy is required. Ollama implementing the Anthropic wire shape is the
shared-standard case DESIGN's external-upstream decomposition separates, so DCH-1
hoists the shape off the vendor rather than forking a third copy.

Also recorded, because it is the class this subject keeps producing: the rendered
unit is the only member of the must-move-together set with no failure signal. Ref,
manifest, digest and closure each raise a checksum event when they disagree; a stale
unit is silent, which is how both hosts sat with a Description naming gpt-oss:20b
and no context-length line at all, unreported.

DCH-0 (enrolment, residency, the parallel axis) is owned by eager-pike-541, to be
confirmed with them before it starts. RLM closes first.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LSzWg5t7F22xEtbd83fzQ6

* DCH-0 restructure: split residency into its own gate, and the endpoints census says list membership is not the load-bearing act

Three changes, all from eager-pike-541's reply plus one census I owed them.

SPLIT. Residency modeling becomes DCH-0b. Their reason is the right one and it is
not about size: enrolment and the parallel axis write values into carriers that
already exist and have shapes, while the resident set has NO carrier at all -- what
a resident-model fact even is, desired or observed, per host or per fleet, its
identity when the same weights appear under two refs, remains undecided. Folding
design into a gate of row-writes makes the row-writes wait on it.

THE CENSUS THEY FLAGGED AS NOT DONE IS DONE, and it changes DCH-0's content. Of 57
non-test modules importing fleet_intent_network, the number reading the endpoints
list or fleet_intent_network_topology() is ZERO -- production consumers take
individual endpoint rows by name (bmc_virtual_media, host_identity_access). The
topology function has exactly one consumer in the tree: the witness asserting
list_length(endpoints) == 11.

So list membership is not the load-bearing act; authoring the srv5/srv6 rows that
named consumers can reach is. And enrolment turns that witness red at 11 -> 13,
where the right response is not to bump the literal -- a count copied from the tree
it measures is the change detector DESIGN section 5 names, and completeness is an
identity join. Repairing that oracle is now part of the gate, and bumping it is on
the must-not list.

THE COST MODEL IN THE CORPUS IS FALSE and DCH-0 now corrects it while adding the
axis. serving_desired says context length and slot count "trade against each other
inside one memory budget, because the slot count divides the context window."
Measured by eager-pike-541 on idle spark-3bd5, same model, only slots changed: 1
slot -> context_length 1048576 at size_vram 88,865,253,620; 2 slots -> the same
1048576 at 90,543,761,652. The window is not divided. Slot count multiplies KV. The
sentence is deleted rather than softened, and the measured per-slot cost goes behind
the desired value -- with the caveat that 1.68 GB/slot is a DeepSeek MLA number and
a fleet-wide slot count from one realization is the same overreach as a fleet-wide
context ceiling from one realization.

Also recorded: enrolment is inert because it has no executor -- no ctrl-fleet-converge
timer or service on either Spark in either scope, and fleet_converge_timer says of
itself that it is no longer a renderer since #8283 killed its installer chain. The
risk is therefore not prematurity but a row that reads as an outcome, which is why
the receipt stays an observed-vs-desired rendered-unit comparison.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LSzWg5t7F22xEtbd83fzQ6

* Enrolment is two acts and only one is inert: ComputeHost has three consumers and one of them is a fail-closed wall

The endpoints census generalized further than it should have. eager-pike-541
censused the other list and it is the opposite shape; I verified the load-bearing
consumer in the tree before folding it in.

fleet_intent_known_hosts has three production consumers:

- gunbc.generated_artifact maps it to RunnerHostSudoersArtifact, so enrolment mints
  two new generated artifacts and the rows cannot land without a same-commit
  regeneration or the drift gate reds.
- runner_host_deploy admit_runner_host returns RunnerHostUnenrolled for absent
  hosts. Its annotation calls itself THE ENROLLMENT WALL and is explicit that it is
  CONSTRUCTION and not a check -- no function in the module takes a bare
  RunnerHostDeploy and yields a command, so running an installer on an unenrolled
  host is unrepresentable rather than discouraged. It prices the srv4
  host-convergence OOM and names the enrolled row as where runner_deployment_plan's
  conservation wall gets host RAM to check slot caps against. Enrolling srv5/srv6
  therefore REMOVES a fail-closed refusal that currently protects them.
- fleet_converge_apply finds hosts by identity in that list, so enrolment is what
  makes a Spark reachable by the apply path this lane measured as PartiallyApplied
  and refused.

The inertness argument had two independent legs, no executor and no consumer.
ComputeHost knocks out the second; only the executor leg survives, and it is the leg
that changes the moment anyone installs the timer.

So DCH-0 lands the endpoints half only. The ComputeHost half becomes DCH-0c and is a
decision rather than a row: host facts the conservation wall consumes, the two
sudoers artifacts regenerated in the same commit, and someone saying out loud that
the wall is coming down for these machines on purpose.

The row/list split holds on both sides and inverts between them -- for endpoints the
row is load-bearing and the list inert; for ComputeHost the list is load-bearing and
the row quiet.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LSzWg5t7F22xEtbd83fzQ6

* Narrow the ctrl boundary to the lane: the global form was false of the repository

Review found a blocking defect in the opening direction. It read "gunbc does not
know about ctrl -- no import, no dependency, no citation of a ctrl artifact as
authority", which is globally false: extdeps.ctrl.gunbc_pin already declares an
ExternalAuthority whose URI points into gunb-ai/ctrl, and its functions compare the
host pin against the ctrl pin.

The operator's intent was about DCH and the harness, and the document already stated
it correctly at that grain further down, under "What this lane must not do". So the
opening asserted a stronger claim than the one being made, and in an
authority-bearing scope document where independence from ctrl is a central boundary,
a false global subject is not harmless prose -- a later reader would take it as a
fact about the repository and find a counterexample in one grep.

Narrowed to the lane, with the counterexample named in place so the narrowing cannot
be re-widened by someone who does not know why it is narrow, and with an explicit
note that extdeps.ctrl.gunbc_pin is out of scope here and not to be touched.

No other plan content moves; this is the only authored-content change from fddf116.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LSzWg5t7F22xEtbd83fzQ6

* DCH: answer the relayed harness transcript as an adversarial review, and record the terminal as necessary-not-sufficient

The operator relayed a transcript of the ctrl mini-agent lane debugging its own
harness against the same Sparks DCH intends to drive, and directed that the plan
answer it before the lane starts.

Admitted on the terms the plan already sets for that lane's output: second-hand,
unreproduced here, each row a hypothesis about our design with a predicted
failure rather than a fact on loan. What it buys is not findings but a cheap
enumeration of where a harness of this shape breaks.

Nine classes, each routed to the gate that owns it. They collapse to one
property: every failure was operator-visible AS SOMETHING OTHER THAN ITSELF -- a
harness throw as a clean exit, a dead session as a working one, a control-plane
restart as a Spark fault, an inflated rate as fast hardware. So the obligation is
not "handle these bugs" but produce a disposition that cannot be mistaken for a
different one.

Three consequences worth naming outside the table:

- The plan's DCH-2 claim that a self-reporting harness retires the observation
  layer now has its measurement. Their harness DID report its status correctly
  throughout, and the operator still saw a frozen session, because the report's
  transport failed independently of the report. Self-reporting is necessary and
  is not the simplification alone.

- DCH-0 puts the serving unit under fleet convergence, and a converge restarts
  the serving process -- so this lane imports that transcript's worst open defect
  BY CONSTRUCTION, into the layer it chose on purpose. Sent for adjudication with
  a position rather than decided here.

- The terminal is amended to necessary-without-being-sufficient. The canary will
  not fill a context window, meet a tool timeout, or sit through a converge, so a
  green on it is a green on the easy path. No substitute terminal still counts.

Five questions are open with the controlling reviewer and the section is not
settled until they return; three of them can change a gate. Two new hypotheses
join the existing list on the same terms -- streaming-specific early termination,
and a turn ending at end_turn with announced work unperformed. Neither is
established: the second has one observation under survivorship, the first has
none, because the step isolating streaming as the single variable was never run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LSzWg5t7F22xEtbd83fzQ6

* DCH: reconcile against #9960, which resolved the concurrency axis while this plan was in review

The controlling reviewer found that main had falsified this plan's active DCH-0
claims. Verified here rather than relayed: #9960 (7810e68, an ancestor of
current main) declares OllamaNumParallel in extdeps.ollama.server_env, binds
ollama_num_parallel_env_assignment in spark.serving_unit_render so the rendered
unit emits four axes rather than three, carries a desired slot count of 4, and
corrects the cost model in the direction this plan predicted -- the count does
not divide the window, and the per-slot price is recorded as a property of the
realization under a declared 4b drop naming that subject.

So the axis item and the cost-model item leave DCH-0. They are dispositioned in a
new 2.1 "changes since the baseline" section rather than edited into the fixed
main@de2f5f baseline, which stays immutable so later main movement does not
rewrite a historical measurement.

What #9960 does NOT resolve is kept open and sharpened. It made the state
representable, declared and renderable; it read neither host back. Desired-versus-
observed stays a live blocker, and the operator's two hand-edits make it sharper
rather than weaker: a hand-set host value agreeing with a declared corpus value by
coincidence is precisely what a silent carrier cannot distinguish from
convergence, so the readback is the entire receipt.

DCH-0 is renamed and its question narrowed to what actually remains: enrolment,
and the live rendered-unit reconciliation.

One defect found while verifying, reported and deliberately NOT repaired here
because it is outside this lane: gunbc.spark.serving_desired still carries an
earlier annotation asserting in the present tense that concurrency is not modeled
in desired state, that the slot count divides the context window, and that the
concurrency row is deliberately not smuggled in. All three are false, and they
contradict that same module's own value and annotation about 140 lines below. A
stale annotation is data the substrate cannot check, so nothing reds -- a 3
meaning fork inside a single authority.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LSzWg5t7F22xEtbd83fzQ6

* DCH: carry the open dispositions into the gate chain, and specify the tool surface instead of forward-referencing it

Review 58352 (codex/gpt-5.6-sol) requested changes on two findings. Both are
correct and both are repaired here rather than argued.

FIRST: the review section identified five gate-changing questions and left them
beside the gate chain rather than inside it, so DCH-0 could begin and complete
while knowingly preserving the restart-mid-turn fail-open that Q3 names. Recording
a class is not handling it, and 5 exists to answer these in advance.

The gate chain now carries an admission rule -- a gate does not complete while a
5.2 question owning one of its dispositions is open -- with the bindings tabled
and repeated in each gate. DCH-0 states its own ending explicitly: a drain or
lease so a converge cannot land against an in-flight turn, or a 4b(3) row whose
trigger names the drain capability. Never silence, and never a note that a restart
is unlikely. The rule bars a gate's RECEIPT, not work inside it.

SECOND: the section said the tool-surface quirks were "expanded there rather than
here" while DCH-2 still listed four tool names. That is a promise naming a route
that does not exist -- the same defect this plan refuses in others, committed by
me in the act of writing the section that refuses it.

DCH-2 now specifies the surface as six owed dispositions under one rule: a tool's
contract is carried in its declared shape, and any limit it enforces is reported
in its result rather than inferred from a truncated or missing one. Working
directory, duration, input, truncation, edit matching, context exhaustion. Two of
the relayed harness's choices are adopted deliberately because they are already
the right shape -- a cap that announces its truncation, and an edit that refuses
on zero or multiple matches rather than patching the first.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LSzWg5t7F22xEtbd83fzQ6

* DCH: carry the five returned dispositions into the gates, and add DCH-0r

All five questions returned 2026-09-02. Recorded as rulings rather than
positions; where a proposal was accepted with refinement, the refinement binds.

The chain changes. Q3 rejected a 4b drop as the normal shape -- a drop reports a
lower guarantee, it does not turn an unsafe arm green -- so the drain becomes its
own gate rather than debt, and reordering removes the need for debt entirely:

  DCH-0 endpoint rows -> DCH-0r quiescent maintenance -> DCH-0c ComputeHost
  enrolment -> DCH-2 turn activation

DCH-0r's mechanical blocker is verified in this tree rather than relayed, and
verifying it found two refinements sharper than the finding as received.
spark.serving_realization realizes EnableSystemUnit as enable, start AND restart
in one effect, so restart has no independently matchable identity and no fence can
guard it. First: that restart is UNCONDITIONAL, emitted whenever the effect is
realized rather than when the definition changed, so converging the unit is
destructive to in-flight work even when it changes nothing. Second, and the
Sparks are on this path: the USER-unit closure already separates EnableUserUnit
from StartUserUnit and carries NO restart effect at all -- so on the hosts this
lane targets, a changed unit definition has no modeled route to take effect. That
is a candidate cause for the observed-versus-desired drift DCH-0 must reconcile,
and it is written as a candidate because nothing here has read a host back.

Also carried: the admission fence and the drain are ONE protocol, because a bare
active==0 read admits a turn immediately after it; and the guarantee grain is
stated rather than assumed, since a census of DCH leases proves quiescence only
for DCH-managed clients.

DCH-2 gains an exit bar of four qualification groups and no longer exits on "four
tools and one terminal turn". DCH-3 requires BOTH the qualification receipts and
the unchanged 14-step terminal, neither substituting for the other. DCH-4 waits on
that combined admission, because retiring a working runtime on an easy-path green
is how a replacement erases a correctness distinction instead of completing.

Q1 refines my boundary: whose failure is evidence, but the smallest closed causal
scope containing the uncertainty decides the stop, and a supervisor alive in a
typed stopped state is not an absorbing fallback. Q2 is accepted at typed-
obligation grain and rejected at prose or character grain, with awaiting-
verification deliberately not acceptance since the model's own EndTurn cannot
grade the model's work. Q4 holds n=1 under survivorship to a single existential
proposition that prior art does not establish for our harness at all. Q5 replaces
both my options with two axes plus journal-before-effect ordering: a supervisor
may move a record from running to interrupted and never to accepted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LSzWg5t7F22xEtbd83fzQ6

* DCH: adopt the operator's only-and-default direction, and separate default from retirement

Operator direction 2026-09-02: make the .dag harness the only and default
harness, ignoring Claude and Codex entirely. Adopted, with the two halves kept
apart because they have different preconditions.

DEFAULT IS A SELECTION; RETIREMENT IS A DELETION. "Ignore the other providers" is
discharged by making ours the default selection at the seam and stopping
investment in the others. It does not require deleting them, and DCH-4's deletion
still waits on DCH-3's receipt. Deleting them to satisfy "only" before that
receipt would leave the repository with zero working harnesses -- the failure the
replacement-migration doctrine calls erasing a correctness distinction rather than
completing the replacement. The other variants stay frozen in the doctrine's
sense until DCH-4.

"ONLY" RAISES THE BAR RATHER THAN LOWERING IT. With three providers a weak DCH-3
is tolerable because a fallback exists; with one there is none, so every DCH-2
qualification class becomes load-bearing on the day it becomes default. The
relayed transcript in 5 is exactly a record of what a sole harness with no
fallback feels like when it breaks: three sessions dead, an exit status of zero,
and an operator who cannot tell working from dead. The combined admission is what
makes "only" survivable, so it is not tradeable for focus.

Also added to the must-not-do list: do not delete or break the existing runtimes
to satisfy "only" before DCH-3's receipt, or the count of working harnesses passes
through zero.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LSzWg5t7F22xEtbd83fzQ6

* DCH: withdraw the RLM-first start barrier, and correct DCH-0r's identity model to three identities

The reviewer superseded my freeze at ef8cf03 for a reason I should have caught
myself: the plan still contained the OLD sequencing in two places -- the opening
"RLM closes first; this lane starts after it" and the section 4 prohibition on
starting before RLM's terminal receipt -- while DCH-3 had already been rewritten
to the operator's 2026-09-02 direction. The document contradicted itself, and no
freeze can make contradictory plan text admissible.

Concurrent DCH and RLM is authorized. DCH-3 is a JOIN, not a start barrier: DCH-0
through DCH-2 build and qualify while RLM is blocked, and the default cutover
waits for both sides. "Only" raises that bar rather than licensing an early one.

DCH-0r's protocol carried two conflations and the correction is now the
load-bearing part of the gate. THREE identities, not one: a stable service locus,
a service-invocation identity that changes on every activation, and a
running-definition identity stable for equal specs. The /proc environ stamp
answers the THIRD -- it is invocation-bound evidence but not an invocation
identity, because two successive restarts from one definition carry the same
stamp and so it cannot prove the old process was replaced.

And the fence is keyed by the LOCUS, not the incarnation. An incarnation-scoped
fence stops blocking admissions at exactly the moment an unvalidated new
incarnation appears mid-drain. The pre-restart invocation is a CAS condition and
an evidence anchor, never the scope of the refusal.

The unstamped reading is narrowed. "No stamp" establishes that the invocation
cannot be shown to have started from the current desired definition; it does NOT
establish that the process predates it, since a foreign or manually started
process could have been launched later and would carry no stamp either. It still
requires reactivation; chronology is just not what the evidence supports.

Also added to the must-not-do list: do not land a change to an authority already
inside another lane's frozen approved delta without serializing it. A clean
textual composition still yields a blob nobody approved.
dag/gunbc/fleet/fleet_converge_plan.dag is the live instance, shared with #9832.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LSzWg5t7F22xEtbd83fzQ6

* DCH: model the behaviour-bearing serving package, add DCH-0p, and retire the degenerate-stop hypothesis to a mechanical cause

Ruling 2026-09-02. The plan's model identity was weights-shaped and that is
insufficient: same weight blob, same template, same runtime, same host, DIFFERENT
EFFECTIVE STOP SET is materially different observable behaviour. So DCH-0b's
subject becomes a ServingModelPackageIdentity carrying template, parser,
parameters, the effective stop-sequence set, runtime release and decode
realization, with a tag as an allocated handle and never the semantic identity. A
package rebuilt from the same weights without a stop string is a DIFFERENT
package; treating it as identical makes the repair invisible to convergence.

New gate DCH-0p, after DCH-0b and DCH-0r and before DCH-2 qualification. It reads
what the endpoint ACTUALLY serves rather than what a Modelfile or a tag says it
serves, refuses with PackageConfigurationIncomplete rather than inferring a
behaviour-bearing field, and admits a stop sequence only when it binds uniquely to
a declared protocol boundary in that exact package's contract. A copied allowlist
is not a role, and sentinel-looking values inherited from another lane are not our
evidence.

THE HYPOTHESIS THIS RETIRES IS THE REASON THE LANE EXISTS. 5 recorded, from the
relayed transcript, that an assistant turn may stop at end_turn having announced
work it did not perform -- one observation, survivorship-filtered, which I was
careful to say established almost nothing. It now has a mechanical cause: a
serving package carrying PARAMETER stop ### terminates generation on ordinary
Markdown H3 output while returning a well-formed stream and stop_reason end_turn.
Reproduced by a tiny prompt, removed by rebuilding the package. So the
heading-then-stop symptom had a determinate external cause the whole time, and the
hypothesis about model behaviour was a transport symptom read as a semantic one.
It leaves the hypothesis list and becomes DCH-0p's subject. The streaming
hypothesis is untouched and still has NO observation.

DCH-2 gains two things. Its terminal semantics may never equate end_turn with
natural completion, because at least two upstream causes produce that same label
and the client cannot prove the native cause from a surface that erased it -- so
DCH-0p and DCH-2 are two independent walls rather than one. And throughput becomes
a construction rather than a threshold: a cache lineage joined to accounting that
separates newly evaluated from reused input tokens, since a rate over total input
divides mismatched subjects and yields a valid arithmetic operation that is not a
hardware measurement. "No four-digit rate is reachable" is a view, which #9946
forbids as a correctness wall.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LSzWg5t7F22xEtbd83fzQ6

---------

Co-authored-by: gunbc-ci-auto-heal <gunbc-ci-auto-heal@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
briansrls pushed a commit that referenced this pull request Sep 14, 2026
…achable; one was mine

The floor at f843706 refused structurally with three diagnostics. They are NOT one class, and
checking each against main rather than batching them is what separated them.

INHERITED, byte-identical on main, latent because no compiled closure reached these modules:

  runner_capacity_operator_access   `transport: SshShell` with no `ssh_host` -- a literal missing a
                                    required field of the variant.
  runner_service_activation_witness `observed_spark: []` in a FullHostConverge literal; that field is
                                    in no version of the type, on this branch or on main. Added by
                                    #9832.

MINE: `wt_srv3_deploy()()`. My migration regex rewrote the identifier in sites where an earlier pass had
already appended `()`, producing a double call. The compiler rendered it `function '<expr>' not found in
scope`, which reads like an unrelated resolution failure until the line is opened -- so it looked
inherited until it wasn't. Sixth scripted-edit defect on this branch, same root as the other five: a
blanket rewrite applied without checking what the text had already become.

WHY THEY APPEARED NOW. Editing those files made them changed witnesses, so the floor compiled them for
the first time. That is this corpus's reachability distinction read in the other direction: an arm
nothing reaches looks fine until something reaches it, and a type error nothing typechecks is invisible
rather than absent.

THE SshShell REPAIR GOES THROUGH THE SPEC, NOT A LITERAL. `runner_host_ssh_endpoint` reads `ssh_target`
off the host's own `RunnerHostSpec`; an unmodeled host renders an unresolvable marker rather than
borrowing a neighbour's address, because a recovery aimed at the wrong machine is worse than one that
does not start. That is the second time the spec split has paid for itself -- the first was the
production consumer that turned out to want `runner_user` all along.

Measured: runner_service_activation 16/16; runner_host_deploy 9 pass with only the inherited
`srv4_sccache_prereq_pins_version_and_digest` red. Regeneration rc=0 and moved no artifact, so the
repairs perturb no projection.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q7tG7JyxfKEDUtBUYuEw6m
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant