Skip to content

Delete the expected-red roster: resolve the 36 held reds - #9042

Merged
briansrls merged 4 commits into
mainfrom
session/crisp-eagle-411
Aug 23, 2026
Merged

briansrls merged 4 commits into
mainfrom
session/crisp-eagle-411

Conversation

@briansrls

Copy link
Copy Markdown
Contributor

Auto-opened by session-dashboard for session crisp-eagle-411.
Pushing to session/crisp-eagle-411 advances this PR.

Worker attestation

Before flipping this PR to ready for review, confirm each item:

  • Title describes the change (not the session id or branch).
  • PR body summarises what and why (replace the TODO below).
  • Tests run: name the command (e.g. npm test, cargo test) and the result.
  • If this closes a work item, the body contains a Closes #N directive.
  • No commits on this branch are surprises (no fork/cherry-pick I did not make).
  • No secrets / credentials / large binaries staged.

Summary

TODO: replace this paragraph with one or two sentences naming the change and its motivation. Reviewers read this first.

Test plan

  • TODO: list the commands that ran (or "no tests changed; relied on CI") and the outcome.

…ile the shell mock's unconditional hermetic arm as a §4b class

Nine `test fn`s under `v2.test.fixture.walk_plan_stage.` are not witnesses. They
are the SPECIMENS that `claim_executor --plan-entry
src/v2/test/fixture/walk_plan_stage/plan.dag --plan-function <recipe>` drives;
the recipes assert on them. The required floor discovered all nine anyway and
enrolled every one — two on `floor_expected_red`, seven on `floor_route_gap` —
so two ledgers that may only shrink each carry rows that can never leave.

Two are red BY CONSTRUCTION (`failure_red_member`'s body is the literal `false`;
`recursion_refusal_member` self-calls so the call-depth refusal fires, which is
what its recipe observes). The other seven perform the host effects the recipes
are about, and `stage2_marker` writes the exact marker path
`walk_plan_stage_failure_barrier_recipe` asserts must stay ABSENT — so routing
it hermetically would have the floor corrupt another fixture's oracle.

`fixture_home_prefixes()` declines them on the AUTHORED module name, the same
mechanism `long_home_prefixes()` uses and the one the 2026-08-04 ruling permits.
`v2.test.fixture.` is deliberately NOT the prefix: 27 modules declare it and only
this family authors `test fn`s. No coverage is lost — the recipes were always the
executing consumer, which is more than a route-gap row can say. The fixture's own
`common.dag` note already claimed this state of the world and had been false since
the floor cut; it is corrected rather than deleted.

Also files a discovered error class. `shell.Exec.Run` declares ONE hermetic mock
arm, `0 => { exit_code: 0, success: true, stdout: "", stderr: "" }`, keyed on an
exit status the mock itself supplies — so `exit 1` and `true` return the same
value. `RunArgv` and `Check` share the shape. Discriminated by required floor run
32645296251 (main 183e597) without a new instrument: two witnesses in one
module differing in one string, `witness_shell_on_host_success_converges` PASSED
and `witness_shell_on_host_failure_fails_closed` KNOWN-RED. `gunbc.hermetic_mock_fidelity`
records it at OutsideTheLadder with ceiling StructurallyGuaranteed and a next
trigger; it changes no behaviour, because deleting the arms is a ~75-call-site
census escalated to the operator, not a side effect of clearing roster rows.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@gunbai-bot
gunbai-bot Bot marked this pull request as ready for review August 23, 2026 19:41
@gunbai-bot

gunbai-bot Bot commented Aug 23, 2026

Copy link
Copy Markdown
Contributor

CI red here is INHERITED FROM MAIN, not from this branch, and it is worth stating exactly so nobody re-derives it.

Run 32662035240 fails in the floor's PREPARATION phase, before any witness executes:

required-ci: FAILED PHASE floor refused: subject=eca8cfe823ae3b14 modules_resolved=3849
dag/gunbc/runner_slot_provision.dag:240:28: error: field 'argv' not found in type 'ArgvCommand'
dag/gunbc/runner_slot_provision.dag:240:14: error: sole_constructor type 'ArgvCommand' cannot be constructed outside its defining module
dag/gunbc/runner_slot_provision.dag:240:14: error: missing required field 'program'
dag/gunbc/runner_slot_provision.dag:240:14: error: missing required field 'arguments'

That is the site #9031 ("Convert the one ArgvCommand site the seal missed, through a builder rather than by widening the seal") exists to fix. This branch does not touch dag/gunbc/runner_slot_provision.dag, and main's own run 32646482842 refuses the same way at the same phase. The earlier witnesses entry showing fail was run 32662021121, which was CANCELLED by the draft→ready flip rather than failing on its own.

The parse and regen phases both passed here: parse OK 51 file(s) parse-clean, required-regen: first_generation_equal=true planned=133 executed=133 declared_divergent=1 [main.rs].

Consequence for this PR, stated rather than papered over: because preparation refuses, no run has yet reported declined_fixture=9, and no run has re-joined either roster after the nine deletions. That identity join is the measurement this PR owes; it becomes takeable the moment #9031 lands. Everything else is green by execution and listed in the description.

— sent from crisp-eagle-411

Brian Searls and others added 3 commits August 23, 2026 20:26
…e each its discriminating row

Reading the class off the exit status alone described the population wrong. The
mock supplies EVERY declared output field, so each field is a separate
fabrication — and one of the three reds in the opposite direction, which is why
"the exit-1 witnesses" was never the right name for these ten rows.

  exit status            witness_shell_on_host_failure_fails_closed — "exit 1"
                         asserts NOT converged; the arm answers success.
  captured output        deploy_access_privilege_witness
                         .witness_privileged_fixture_surfaces_stdout — a
                         SUCCEEDING script asserts its stdout carries a marker,
                         and the arm's stdout is "", so this row reds on a
                         script the arm gets RIGHT.
  transport reachability witness_ssh_shell_unreachable_fails_closed — "true"
                         over SshShell to 203.0.113.1 (TEST-NET-3, reserved
                         unroutable) asserts NOT converged; the arm answers
                         success without the transport being consulted.

`FabricatedFact` is typed rather than left in the header because the three have
different consumers and different remedies: a fix discriminating only on exit
status would leave the other two standing. `Check` declares no stdout, so its
row names two facts and not three — the list is read off each operation's own
output block.

All ten held reds in the cluster were read at source and every one asserts on a
fact in that list; none is a separate defect wearing the same symptom.

`hermetic_mock_row_names_what_it_fabricates` is a second predicate rather than a
clause inside the trigger check, because an untracked row and an undiagnosed one
need different repairs. Its probe carries a trigger AND an empty list, so it also
discriminates against the FIRST check: a green there would mean the two had
collapsed into one.

Six witnesses evaluated true against a release gunbc built in the same remote
dispatch.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…fold and declined_fixture=9 becomes measurable
…eachability the one that has no hermetic remedy

Reachability is structurally different from the other two and a shared trigger
hid that. An exit status and a captured stream are functions of the input the
operation already declares, so a faithful mock could in principle serve them. No
mock can honestly answer "was 203.0.113.1 reachable" — the only truthful moves
are to consult the transport, which the hermetic floor forbids by construction,
or to refuse. So a remedy that made the arm vary with its declared input would
close two of the three, pass review, and leave the sharpest one exactly where it
is. That is now a match arm, not a sentence.

Moving the trigger onto the FACT also makes the old check unnecessary rather than
redundant. `next_trigger: String` on the row was authorable as `""`, so it needed
`hermetic_mock_row_is_tracked` to catch that; derived by a total function, a
fabricating row cannot lack a trigger and the check has no red left to find.
Deleted with the field rather than left standing permanently green and citable as
coverage (§4b). The authorable red that remains — `fabricates: []` — is a
different defect and keeps its own check, and it now matters more than before:
with the triggers hanging off the facts, an unnamed fact is a dropped remedy.

Replaces the retired check with `hermetic_mock_reachability_trigger_is_distinct`,
which reds if the three triggers ever collapse to one string.

Four witnesses evaluated true against a release gunbc built in the same remote
dispatch.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@gunbai-bot

gunbai-bot Bot commented Aug 23, 2026

Copy link
Copy Markdown
Contributor

THE MEASUREMENT THIS PR OWED IS TAKEN, and CI is green. Run 32665124117 on ca1c4b84a6 (main merged in at f49886339a, which carries #9031's ArgvCommand fix — the preparation refusal reported above is gone):

[floor-phase] phase=site-projection sites=12144 files=1589 claims=10688 declined_long=548 declined_fixture=9 declined_live=899
required-floor: offered=12144 routed=10688 declined_long=548 declined_fixture=9 declined_live=899 — every discovered site is exactly one of these
required-floor: planned=10688 executed=10688 terminal=10688 passed=10417 known_red_held=35 failed=0
              stale_quarantine=0 route_gap_unenrolled=0 route_gap_held=94 stale_route_gap=0
              known_red_now_passing=0 known_red_runtime_errored=142
required-ci: phases_run=3 failed=0

declined_fixture=9 — exactly the nine walk_plan_stage members, no more and no fewer, and the site partition still closes (12144 = 10688 + 548 + 9 + 899).

The joins the deletions had to survive, all clean. These are the counters that would have caught a mismatched pair of edits, and each is zero:

  • route_gap_unenrolled=0 — no newly-declined site reappeared as an unenrolled route gap.
  • stale_route_gap=0 — no surviving floor_route_gap row names an identity that stopped route-gapping.
  • stale_quarantine=0 / known_red_now_passing=0 — no surviving floor_expected_red row greened.
  • failed=0.

Had I deleted a roster row without declining its site, or declined a site without deleting its row, one of those four would be nonzero. That is the identity join, and it is the reason the deletions and the decline had to land in one change.

Attribution, stated precisely rather than by subtraction. route_gap_held 101 → 94 is exactly my seven. known_red_held 36 → 35 and known_red_runtime_errored 143 → 142 are each one row, mine — but main moved under this branch between the two runs (offered 12124 → 12144), so those two deltas are read off a different base than the 179-row census in the description and only the direction is attributable by arithmetic alone. The nine that are unambiguously mine are the ones the new counter names directly.

The four hermetic_mock_fidelity witnesses are in passed; the carrier changes no behaviour, as described.

— sent from crisp-eagle-411

@briansrls
briansrls merged commit 388771a into main Aug 23, 2026
2 checks passed
@briansrls
briansrls deleted the session/crisp-eagle-411 branch August 23, 2026 22:14
gunbai-bot Bot pushed a commit that referenced this pull request Aug 23, 2026
…17-row unenrolment on top

# Conflicts:
#	src/v2/workflow/floor_expected_red.dag
briansrls pushed a commit that referenced this pull request Aug 23, 2026
…stream in #9042, so this branch's copy is superseded

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
gunbai-bot Bot pushed a commit that referenced this pull request Aug 24, 2026
The floor's own removal path, run against the head that carries main:
run 32670979986 reports `stale_quarantine=82`, and stale-quarantine is
gating (`required_floor_outcome_is_clean` lists seven causes and includes
it). All 82 identities were verified present in the roster before removal
and absent after; 184 rows -> 102; every chain line brace-balanced; no
duplicate heads.

The seventeen this branch removed earlier were measured against a roster
that no longer exists -- main's #9042 rewrote it under the branch. The
count is corrected in place rather than annotated beside, because two
counts for one fact is two accounts of one fact.

WITHDRAWN WITH IT: the claim that this edit carries a same-module
discriminating half. The four generated_coproduct_exhaustiveness_* rows
were retained on the earlier reading that they sat in the same generated
module and were not reported passing. The same run names all four
explicitly as STALE-QUARANTINE, so that is false of the current
measurement; retaining them would enrol four rows this branch makes pass
-- the exact gating failure the removal exists to avoid. Removing them is
correct and the receipt I claimed for the retention is not, so the claim
is withdrawn rather than left standing unsupported.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
briansrls pushed a commit that referenced this pull request Aug 25, 2026
…alse) on main's own head, and 170 enrolled expected-red rows are ERRORING rather than failing (#9020)

* WIP: MAIN RED #4: eight runner-slot/fleet-capacity witnesses return Bool(fals

* Unenrol the seventeen rows this fix made pass, and keep the four beside them that it did not

The floor reported seventeen v2.test.claim.generated_conformance_floor identities as
STALE-QUARANTINE -- enrolled as expected-red and PASSED. That is this branch's fix working:
all seventeen were previously KNOWN-RED-RUNTIME-ERRORED, throwing `no such function` on names
that were genuinely declared, because the reference-closure collector narrowed on binding
metadata its own producer never populates. Deriving the closure from lexical binders instead
makes the calls resolve, the witnesses answer, and every one of them answers PASS.

The two arms are not symmetric and that is why this reds the build. KNOWN-RED-RUNTIME-ERRORED
is reported and deliberately NOT gating (claim_executor required_floor_outcome_is_clean lists
seven causes and that is not one of them); STALE-QUARANTINE IS one of the seven. So the fix
moved seventeen rows from a non-gating arm to a gating one, and the roster edit is the
prescribed remedy the floor's own message names.

The four generated_coproduct_exhaustiveness_* rows in the same generated module are retained
deliberately. They were not reported passing, so they are still red for their own reasons --
unenrolling the module wholesale would have dropped four live reds under cover of a repair
that never reached them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Re-derive the unenrolment at the measured population: 82 rows, not 17

The floor's own removal path, run against the head that carries main:
run 32670979986 reports `stale_quarantine=82`, and stale-quarantine is
gating (`required_floor_outcome_is_clean` lists seven causes and includes
it). All 82 identities were verified present in the roster before removal
and absent after; 184 rows -> 102; every chain line brace-balanced; no
duplicate heads.

The seventeen this branch removed earlier were measured against a roster
that no longer exists -- main's #9042 rewrote it under the branch. The
count is corrected in place rather than annotated beside, because two
counts for one fact is two accounts of one fact.

WITHDRAWN WITH IT: the claim that this edit carries a same-module
discriminating half. The four generated_coproduct_exhaustiveness_* rows
were retained on the earlier reading that they sat in the same generated
module and were not reported passing. The same run names all four
explicitly as STALE-QUARANTINE, so that is false of the current
measurement; retaining them would enrol four rows this branch makes pass
-- the exact gating failure the removal exists to avoid. Removing them is
correct and the receipt I claimed for the retention is not, so the claim
is withdrawn rather than left standing unsupported.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Restore the chunk aggregator my own unscoped edit blanked

The previous commit emptied `floor_expected_red_chunks()` to `Empty {}`,
so the roster evaluated to ZERO identities and run 32678275911 refused
with `cause=ExpectedRedRosterEmpty`. The 102 surviving rows were still in
their chunk functions; nothing referenced them.

THE DEFECT IS THE ONE I HAD JUST WRITTEN UP FOR ANOTHER LANE, committed
in the same edit I used as the example. I enumerated the population by
matching QUOTED IDENTITIES and edited by matching ANY LINE STARTING WITH
`Cons { head:`. The aggregator is a Cons chain whose heads are chunk
FUNCTION CALLS, not strings, so it was inside the edit's denominator and
outside the census's: the rebuild found zero quoted heads in it and
faithfully rewrote it as the empty chain.

So the rule is not "scope your regex better", it is the narrower one:
ENUMERATE AND EDIT THROUGH ONE MECHANISM, or diff the two sets and refuse
on non-empty symmetric difference. A count that only counts what the
census can see cannot detect damage the edit does outside it -- my
"82 removed, 0 remaining" was true and told me nothing about the
aggregator.

VERIFIED STRUCTURALLY THIS TIME, not by count: parse both revisions into
per-function head lists, then assert every LOST head is a quoted
identity in the reported-82 set, no head is GAINED, and no chunk function
appears or disappears. Result: 82 identity removals, 0 unexpected
changes, removed set == reported set.

The guard worked exactly as designed -- an empty roster makes the
partition-sum and did-not-execute checks vacuous, and the refusal says
so and names the condition rather than running a vacuous floor to green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Author the seed-growth receipt this change owed: 853 hand Rust lines, no row

Review 55459 requested changes and is right. gunbc.seed_growth_admission
`seed_growth_forward_freeze_policy_note` makes unenumerated hand-written
src/v1 Rust a STOP-LINE, requiring exact item identity, .dag authority,
reason, owning lane, deletion trigger, current boundary, and both deltas.
This change added 853 lines and authored none of it.

WHAT LANDS: gunbc.reference_closure_binder_seed_growth carrying
`reference_closure_binder_seed_growth_justification`, registered in the
CLOSED roster `seed_growth_justification_roster()` and named in the
roster note -- the policy requires both, one row beside the obligation
and one in the roster.

ENUMERATED AT ITEM GRAIN, measured rather than recalled: 9 citable
declarations (8 production, plus the test module
`reference_collector_binder_fixtures`, which carries 22 further hand
items including the two mutation controls). Hand-LOC delta +853/-21 in
cli_run.rs by `git diff --numstat` against the merge subject.

NOT NETTED: collect_node_refs, collect_node_refs_inner and
reference_resolution_facts were MODIFIED, not added, and are recorded
as ExistingSeedItemModified so they are not counted as additions. No
deletions are netted against the additions -- the policy forbids it and
none are claimed.

WHY RUST REMAINS: the collector runs inside the seed's own resolve path,
building the index a .dag authority would itself need in order to
evaluate, so a .dag realization would require the closure it computes.
The trigger names the ordering break rather than a date, and refuses a
partial migration explicitly: a split between a .dag classifier and a
seed walk would be two authorities for one decision.

Same defect crisp-boar-716 self-reported on #9095; same remedy, same
roster.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* The second consumer read the refusal and published anyway

review 55667 is correct. `reference_resolution_facts` builds an
`ExprVarClassification` and never reads `classify.refusals` or
`classify.reconciles()` before publishing edges. An unsupported binder form
means the binder set is incomplete, so a name that IS bound can be recorded
as a free reference or the reverse; the graph is then published under-bound
or widened while the refusal sits unread in a field. The floor's index build
checks both correctly -- this is the same judgment with one consumer
enforcing it and one ignoring it, which is the fail-open the classification
exists to remove.

TWO CAUSES, NOT ONE. A binder refusal names a syntax the collector must
learn; a reconciliation miss names an accounting defect in the collector
itself. Different remedies, so one reason symbol each -- collapsing them
would route both to whichever remedy the shared symbol suggested.

SKIP-AND-RECORD, matching this producer's three existing refusal arms
(`unreadable`, `no-module-line`, `parse-failed`). The file is excluded from
the published edges and its exclusion is typed, located and countable through
`reference_accounting_refusals`. That is what separates it from the
empty-observation narrow: a skipped file is visible in the refusal channel
rather than silently absent from the graph.

RECEIPT UPDATED, and the near-miss is worth recording. The hand-LOC figure
moved +853 -> +887. My first edit added a SECOND figure in the header while
the original stood at line 49 -- two counts for one fact, in the receipt
whose entire purpose is to enumerate growth, three hours after gunbc#9151
was sent back for exactly that shape in a neighbouring file. Caught before
commit by grepping the file for the thing I had just written rather than for
the thing I had in mind. One figure now, updated in place, with its basis
stated: measured against the merge base, not against main's tip, because a
diff against a moving tip counts main's own edits as this change's growth.

EVIDENCE, and its limit. `cargo check -p v1-compiler` passes. That is a
positive control only: the two new arms have NO executing discriminating
test, because reddening them needs a .dag on disk with unsupported binder
syntax plus a call over a temp root, and this producer's fixture harness
takes source strings rather than roots. The classifier's own refusal path IS
covered (the fixture asserts `refusals.is_empty()`); what is uncovered is
this consumer's response to it. Naming that gap is the honest report -- the
missing root-based fixture is the next-rung trigger, not a claim of coverage.

* Three of the nine receipt rows named referents the carrier cannot represent

review 55673 is right, and the overcount was not a miscount. classify,
refuse_binder and reconciles are methods on `impl ExprVarClassification<'_>`
(cli_run.rs, one impl block, three methods -- verified by reading it), and
std.decl_ref has no ImplMethod field. A DeclarationRef naming an impl method
is not a row that happens to be wrong; it is a row whose referent the carrier
cannot express, in a census whose entire value is exactness.

gunbc.seed_growth_admission already says impl methods are uncitable and
already carries the precedent: stage0_rust_observation_seed_growth_justification
documents `impl std::fmt::Display for Stage0CargoBinManifestParseRefusal` as
prose for this reason. So the correction follows that precedent rather than
inventing a spelling -- the three are enumerated in prose, by name, with what
each does. They are still hand Rust and still enumerated; what they are not
is citable, and stating the distinction is the point.

collect_every_var_name came out for a different reason, and it was this
receipt contradicting itself: it lives inside mod
reference_collector_binder_fixtures, which the same paragraph already names
as the enumeration unit because the module is what would be deleted. Counting
the module and one of its members double-counts.

HAND-ITEM DELTA: +9 -> +5 citable (4 production, 1 test module), plus 3
uncitable impl methods in prose. The four that survive are genuinely citable
and I checked each: ExprVarClass (enum), ExprVarClassification (struct),
binder_names_of and pattern_binder_names (free functions, not methods).

This is the second correction to this receipt in one PR -- the first was a
stale LOC figure. Both are the same class the receipt exists to prevent, in
the receipt itself, which is worth saying plainly rather than fixing quietly:
an item-grain census is only worth its exactness, and mine was wrong twice
before a reviewer had to point at it.

EVIDENCE: prose and row-list only, no behaviour. Evaluating this module's
justification on the pre-edit and post-edit trees gives byte-identical
diagnostic sets (203 lines, zero diff); the single diagnostic naming this
file is the stale local shim's inability to parse a section 4c annotation
block, present identically before and after.

---------

Co-authored-by: gunbc-ci-auto-heal <gunbc-ci-auto-heal@users.noreply.github.com>
Co-authored-by: Brian Searls <briansearls1@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant