Skip to content

Delete supply-side lens enforcement from floor discovery - #8141

Merged
briansrls merged 18 commits into
mainfrom
session/eager-cat-841
Aug 12, 2026
Merged

briansrls merged 18 commits into
mainfrom
session/eager-cat-841

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Aug 11, 2026 •

Copy link
Copy Markdown
Contributor

What

Two censuses ran inside discover_floor_witness_roster: the inert-lens reach walk and the construction-justification census. Both ask a question about who authored a lens, and both answer it by acquiring a whole-corpus module graph.

Every discovery run paid for that acquisition, and who that was changed under this branch, so it is stated by era rather than as one population: before #8140 the roster walk was unconditional, so ordinary CI, regen, the falsifier cadence and every coordinated worker all paid; #8140 made it demand-directed and regen stopped paying; from there until this PR the two censuses burdened every remaining discovery-bearing execution — all of which wanted a witness roster and nothing else. That is the shape the floor-cost work keeps finding: the unit of computation is the world, the unit of fact is one module's authorship.

Neither census had a consumer that justified the acquisition. The inert-lens half additionally reported through two host builtins whose entire .dag surface was a pair of self-recursive stubs — fn inert_lens_unreached_module_count() -> Int { inert_lens_unreached_module_count() } — reachable only because the interpreter intercepted the spelling before the body ran.

Deleted end-to-end, not merely unwired

Per the acceptance bar: zero non-historical references, zero host dispatch, zero floor execution, zero registry obligation.

  • src/v2/lens/inert_lens.dag
  • both host builtins, both interpreter dispatch registrations (free-call and v4_bridge), the generated EvalCallBridgeLensInertLens family, the 04_method.dag type-table entries, the std.primitives roster rows and the four v1_interpreter_primitive_surface arms
  • the InertLens registry variant, lens_registry_v0_inert_lens, lens_contract_inert_lens and the lens_module_gate invariant surface
  • inert_lens_modules, inert_lens_modules_legacy, lens_justification_census, unjustified_lens_modules, declares_construction_justification, and both floor refusal arms (cli_run.rs and floor_discovery_snapshot.rs)
  • dag/test/claim/long/inert_lens_hygiene_witness_test.dag and its frozen deferral row, recorded in the shrink log as a deletion-of-subject rather than a migration

Net +80 / −908 across 26 files (−828), refreshed at head 9b7626fa002 — the first three figures in this body were taken at the opening head and went stale twice as review corrections landed.

What this does NOT claim

The expensive facts build is still there. build_module_graph_facts_live has two remaining consumers on this path — effect-reach derivation, and the cross-worker snapshot transport (facts_to_snapshot → FloorDiscoverySnapshot.module_graph_facts → install_module_graph_facts_cache). Deleting both censuses does not delete the world acquisition, and an earlier draft of this work asserted it would. That was wrong; the fourth consumer is downstream in the serialization and invisible from the producer body.

The open question this leaves is the right one to leave: which exact graph facts does the coordinated scoped worker consume, and were they already produced by the ordinary execution that precedes it?

This is a scope narrowing, not a climb. A newly authored lens with no witness, and a lens recording no construction_justification, are both writable again and nothing detects either. The obligation survives as review diligence, which is strictly weaker. Declared in DESIGN §6 with its next-rung trigger rather than left to be rediscovered: an authorship fact belongs on the module's own declaration, checked at ingestion where the module is already parsed, rather than reconstructed corpus-wide by a consumer that wanted a roster.

Retained

refuse_on_module_graph_read_refusals and both of its red controls stay. The arm was written for the lens censuses, but it now guards the two live consumers; the tests moved to a module_graph_read_refusal_tests module and the unreadable-source control is renamed off its lens framing.

Citations repaired (§3)

Two live prose sites cited things this change deletes, so they are corrected in the same diff rather than left stale:

  • gunbc.roadmap_authority v1_interpreter_primitive_roster_acceptance_note cited free_call_shadowing_is_exactly_the_two_inert_lens_bridges as surviving exact-set evidence. That witness becomes free_call_carries_no_shadowed_spelling — an exact set assertion over an empty set, which is genuinely weaker, and the note now says so rather than reading as though nothing changed.
  • v2.compiler.source_authority used the inert-lens stubs as its worked example of host interception.

Verification

cargo check -p v1-compiler --all-targets clean; cargo fmt --all --check clean via hooks. The .dag corpus compile, the primitive-surface drift gate and the witness corpus run in CI — a whole-tree build is a CI job, not a session job.

🤖 Generated with Claude Code

https://claude.ai/code/session_018ZrxLJrt8ASGLegiK72fAi

gunbc-ci-auto-heal and others added 13 commits August 11, 2026 11:18
… pre-plan

claim_executor ran discover_floor_witness_roster_with_snapshot once up front,
BEFORE resolving the plan, on the stated ground that a naming violation should
be "the cheapest possible failure". Measured, that walk is the most expensive
phase in the process: 5.9 min of a 56.5-min ordinary floor (run 31477894666),
and ~6 min of a ~15-min regen whose plan has exactly two nodes.

It is expensive because "naming hygiene" is a misleading label. The four rules
in v2.workflow.floor_naming_hygiene are string predicates over file paths and
line prefixes, but the roster producer they are reached through also builds
module-graph facts, runs a second strict reference-resolution pass, computes
path indexes, and runs inert-lens reachability plus the construction-
justification census. A two-node regen plan paid all of it to discover a roster
it never reads.

Hygiene is a property of the witness ROSTER, so it is now paid by the plans that
have one: the walk moves to the existing `schedules_discovery` predicate, after
the plan's batches settle. The roster is memoized by request digest
(IN_PROCESS_ROSTER_BY_REQUEST), so plans that DO schedule discovery pay exactly
what they paid before — the corpus batch hits the memo this call fills. Plans
that do not schedule discovery pay nothing, and cannot be unhygienic: they have
no roster.

The walk-attempt id is minted unconditionally as before; it is a tracing
coordinate every later phase stamps, and it is not the expensive part.

This also closes the PRELUDE COVERAGE HOLE gunbc.ci_spec
gunbc_ci_floor_batch_wall_budget_note already names: the walk sat outside every
batch budget and could only red at the step cap. It is now inside the region the
plan accounts for, or absent.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018ZrxLJrt8ASGLegiK72fAi
…iew 51099)

review 51099 (cursor/composer-2.5, REQUEST_CHANGES) correctly caught that #8140
changed documented CI behavior without updating its authority. DESIGN.md
Building & checks stated the opposite of the new code:

  "the executor runs the zero-enrollment naming walk when no discovery batch
   is scheduled"

That clause is superseded here in gunbc.design_document (DESIGN.md is generated
from it; the heal job regenerates the projection).

Both of the reviewer's findings are recorded rather than only the first:

(a) SCOPE — discovery-free plans no longer run the corpus-wide nameability
    rules. Declared as a narrowing, with the honest coverage argument: every PR
    runs the `ci` job, whose plan schedules discovery and walks the full tree,
    so per-PR coverage is unchanged; what is deleted is a second redundant walk
    on regen-only and plan-artifact-only runs. Explicitly NOT backstopped by the
    affected-set falsifier, which has produced no green verdict since
    2026-08-03.

(b) ORDERING — for plans that DO schedule discovery the walk now runs after
    plan resolve/eval, so a naming violation pays ~0.5 min of plan resolution
    before refusing. The reviewer asked whether the trade is intentional: it is,
    and it is priced — ~0.5 min later on the refusing path against ~6 min saved
    on every regen.

A dissolve-on is recorded: the clause, the separate walk, and the `__`-basename
rule all retire when the placement rules move to canonical source ingestion and
test identities derive from parser-produced declarations.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018ZrxLJrt8ASGLegiK72fAi
…51101)

review 51101 (cursor/composer-2.5) caught claim_executor.rs:10283 still
asserting "The walk still runs before plan evaluation, so a naming violation
stays the cheapest failure" — the exact opposite of what this PR does.

Non-blocking as a defect, but it is the same class the PR itself is about: a
comment standing as authority for behavior the code no longer has. #8140's
own receipt was a block comment whose stated premise had been false for
months.

The paragraph's real subject — install output policy BEFORE the walk so the
whole-tree read is funnelled rather than emitting ~2.3k `[file] read` lines —
is unchanged and still correct; the walk simply moved further away from it.
Rewritten to say that, rather than deleted, so the ordering requirement keeps
its rationale.

Swept the rest of claim_executor.rs / cli_run.rs for other assertions of the
old ordering: the only remaining hits are this PR's own comment describing the
prior behavior in the past tense.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018ZrxLJrt8ASGLegiK72fAi
Two censuses ran inside `discover_floor_witness_roster` — the inert-lens
reach walk and the construction-justification census. Both asked a question
about who authored a lens, and both answered it by acquiring a whole-corpus
module graph. Every discovery run paid that: every PR, regen, and every
coordinated scoped worker, all of which wanted a witness roster and nothing
else. The unit of computation was the world; the unit of fact was one
module's authorship.

Deleted end-to-end, not merely unwired:

- `v2.lens.inert_lens` (its `.dag` surface was two self-recursive stubs,
  `fn f() { f() }`, reachable only because the interpreter intercepted them)
- its two host builtins, both interpreter dispatch registrations, the
  generated bridge family, the `04_method` type-table entries and the
  `std.primitives` roster rows
- the `InertLens` registry variant, its registry row, contract row and
  `lens_module_gate` invariant surface
- `inert_lens_modules`, `inert_lens_modules_legacy`, `lens_justification_census`,
  `unjustified_lens_modules`, `declares_construction_justification` and both
  floor refusal arms (`cli_run.rs` and `floor_discovery_snapshot.rs`)
- the long witness and its frozen deferral row (shrink logged)

`build_module_graph_facts_live` still runs on this path and this change does
not claim otherwise: effect-reach derivation and the cross-worker snapshot
transport both consume it. `refuse_on_module_graph_read_refusals` and its two
red controls are retained and re-homed, since the fail-closed arm now guards
those consumers rather than the deleted censuses.

This is a scope narrowing, not a climb. A new lens with no witness, and a lens
recording no `construction_justification`, are both writable again and nothing
detects either. Declared in DESIGN §6 with its next-rung trigger: authorship
belongs on the module's own declaration, checked where the module is already
parsed, rather than reconstructed corpus-wide by a consumer that wanted a
roster.

Two citations repaired rather than left stale (§3): the roadmap acceptance note
cited a shadow witness this change renames and weakens, and `source_authority`
cited the inert-lens stubs as its example of host interception.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018ZrxLJrt8ASGLegiK72fAi
…ement

Review 51113 (cursor/composer-2.5) found two second authorities still claiming
enforcement the code no longer performs — DESIGN §3 dual representation and §4b
rung honesty. Both confirmed against the current head, both fixed.

`gunbc.plans.construction_justification_rule` said the presence check "run[s] in
`discover_floor_corpus_rows`" and listed it as current status. Its retirement
condition also named that check as its trigger, so after the deletion the
condition could never fire — an unreachable lifecycle claim that structurally
cannot report itself satisfied. The plan now leads with a supersession notice,
§2/§3/§4 are marked historical rather than reworded, and a new §5 records what
was deleted, what it costs (a lens added tomorrow with no justification lands
green; the 35-module classification in §4 is a historical measurement, not a
maintained invariant), and the next-rung trigger. The retirement condition is
replaced with a reachable one and says why.

`floor_discovery_snapshot.rs`'s consumer census still listed the two gates and
still called the roster walk "pre-plan", which #8140 already made false.

Swept the rest rather than fixing only what was reported: eight comments in
`cli_run.rs` and one in `claim_executor.rs` named the inert-lens reach as a live
consumer of the observation rows, the reference-edge producer, and the selection
tier. Repointed to the consumers that remain.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018ZrxLJrt8ASGLegiK72fAi
Two bounded review corrections.

`is_top_level_lens_module` survived the census deletion with no remaining
caller — only its own definition. It compiled clean because the crate carries a
blanket allow, which is exactly why the residue needed finding by reading rather
than by warning. Deleted; the PR's bar is end-to-end with zero residue, and a
dead classifier left behind is the pattern this change exists to close.

The new DESIGN paragraph asserted the censuses charged "every discovery run — on
every PR, on regen, in every coordinated worker". True before #8140, false on
this PR's base: #8140 made the roster walk demand-directed, so a discovery-free
plan such as regen already stopped paying. Both DESIGN and the plan carrier now
split the claim by era — unconditional before #8140, regen exempt after it, the
censuses burdening every remaining discovery-bearing execution until this
deletion. A change removing stale supply-side enforcement must not land a fresh
stale assertion in the same diff.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018ZrxLJrt8ASGLegiK72fAi
@gunbai-bot

gunbai-bot Bot commented Aug 11, 2026

Copy link
Copy Markdown
Contributor Author

Floor is red on inherited main breakage, not on this PR

build, regen and heal_generated_artifacts all pass at e585778. The ci floor fails with two compile-clean classes, and neither is caused by this change:

58 × dag/{gunbc/ci_budget_tree, product/budget_tree, test/claim/budget_tree_witness_test}.dag
        error: unresolved type 'Memory'
 2 × dag/test/claim/github_app_registry_witness_test.dag
        error: function 'unsatisfied_capabilities' not found in scope

Evidence that this is inherited rather than introduced:

  • Main is red with the identical signature. Run 31518518948 on main 341d3203b fails with 92 × unresolved type 'Memory' and 20 × unsatisfied_capabilities' not found. Same classes, same files, larger counts on the wider scope.
  • None of the failing files are in this diff. git diff --name-only origin/main...HEAD contains no measure, budget_tree, or github_app_registry path.
  • std.measure is untouched here and identical on both sides. Memory is a Quantity variant imported as a type argument (Appropriation<Memory, One>); the last commit to touch dag/std/measure.dag is 341d3203bc2 (SPATIAL-2, SPATIAL-2: units grounding, signed components, affine space, exact Euclidean geometry #8091), which is exactly the commit at which main went red.

I am deliberately not repairing either class inside this PR. Both are someone else's lane, and folding an unrelated corpus repair into a bounded deletion is how a deletion stops being reviewable. This PR needs a green base, not a wider diff.

Merge criteria otherwise: 2/2 distinct approvals (claude, cursor), no open REQUEST_CHANGES, mergeable CLEAN, build/regen/heal green. The floor gate is blocked on main healing; once it does I will merge main in and re-run rather than push code.

Accounting refresh, prose only: the body's +80 / −908 across 26 files is the authored head 9b7626f. At the healed head e585778 GitHub reports +96 / −914 across 27 files, the difference being the auto-heal commit's regenerated DESIGN.md and docs/plans/construction-justification-rule.md projections.

— sent from eager-cat-841

gunbc-ci-auto-heal and others added 3 commits August 11, 2026 20:50
… census

Review 51154 caught a consistency failure in my own work: I applied the
stale-authority fix to `construction_justification_rule.dag` and then judged
`inert_layer_lens.dag` "historical prose" without running the same test on it.
It fails that test. It is a registered plan whose §3 tells a future worker that
Tier 1 is "buildable now (reuse #5433)" and to extend `inert_lens_modules` —
a function this PR deletes — and §7 goes further, advising them to extend it
behind a flag rather than fork it. Someone following that plan would go looking
for machinery that is gone.

Four rows repointed rather than deleted, since the design reasoning survives
even though its cited mechanism does not:

- §3 Tier 1 now leads with the supersession, names `v2.lens.module_graph` as the
  surviving reachability authority, and says plainly that "reuse the existing
  walk" now means "build the walk", which is a larger job than the paragraph
  reads.
- §6's reuse map repoints the transitive-reachability-BFS row off
  `cli_run.rs:2558-2606` — a positional citation into a file that has since
  moved several thousand lines, which is the §3 rot mode exactly.
- §7's seed caveat drops the extend-behind-a-flag advice and states the real
  constraint: whatever Tier 1 becomes must not reintroduce a corpus-wide walk
  inside floor discovery, because the placement was the defect, not the walk.
- §5's "fail closed exactly as #5433 does" moves to past tense; no lens-inertness
  gate runs today.

Also marked the two remaining historical citations, in the same plan's landed
doc-graph receipt and in `axiom_syllogism_lens.dag`'s precedent table, so every
surviving mention of a deleted symbol carries its deletion beside it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018ZrxLJrt8ASGLegiK72fAi
@gunbai-bot

gunbai-bot Bot commented Aug 11, 2026

Copy link
Copy Markdown
Contributor Author

Both red checks trace to one timeout, and it is not this diff

regen hit its 15-minute step cap; ci then refused without evaluating anything:

FloorUpstreamAlreadyRed build=… regen=failure heal=…
  "an upstream job already established this branch is red, so the floor refuses
   before spending its budget. The located cause is that job's own diagnostic;
   this step adds no verdict of its own."

So there is exactly one cause to explain, and the decisive evidence is that two runs exist on the identical head a9115935e:

run event regen
31537407119 workflow_dispatch success, 7m16s
31537265997 pull_request failure, 15m00s (step cap)

Same commit, same tree, opposite outcomes — content cannot explain that. The branch's own history agrees: regen has passed seven times here at 6:47–9:33, including twice after the main merge, and main's regen at the equivalent commits runs 6:33–7:35. The two runs were executing concurrently on the same SHA, and the failing run's siblings were slow too (heal 5m24s, build 2m03s against this branch's usual).

I am not pushing a fix, because there is no defect here to fix:

  • Not a code change. The tree that timed out is the same tree that passed regen in 7m16s thirty minutes earlier.
  • Not a cap raise. DESIGN is explicit — "Do not raise CI step cap first — cap lives in gunbc.ci_spec." A cap raise to absorb contention is the absorbing-fallback shape §5 forbids: it would zero the failure's frequency by construction and hide the real signal.
  • Not a third concurrent run. Starting one now would compete with the dispatch run's floor and risk reproducing exactly the contention that caused this.

The dispatch run's ci floor is executing on this SHA now and will produce the real verdict. Once it settles I will re-run the pull_request regen/ci jobs so the required checks reflect what the tree does.

One finding worth keeping past this PR: regen at 6:47–9:33 against a 15-minute cap has under 2× headroom, and contention alone is now demonstrated sufficient to exhaust it. #8140 bought that margin (14:54 → 8:26); this run is evidence the margin is thinner than the numbers suggest, and that the next regen-input growth spends it.

— sent from eager-cat-841

@gunbai-bot

gunbai-bot Bot commented Aug 11, 2026

Copy link
Copy Markdown
Contributor Author

Correction to the mechanism in my previous comment

I attributed the regen timeout to concurrency — two runs on the same SHA competing for runners. That was inference from the timing overlap; I did not check where the jobs ran. The observable fact is host identity, and the correction came from the operator thread rather than from me.

same SHA a9115935e:
  srv1-03   success   7.27 min
  srv4-03   FAILURE   >=15 min  (killed at the cap — CENSORED, not measured)

other regens, comparable content:
  srv1-05   success   9.35 min
  srv2-05   success   9.55 min
  srv4-01   success   8.47 min

The conclusion in that comment stands — the timeout is not caused by this diff, since the identical checkout passed in 7m16s. What changes is why, and the difference matters for anyone sizing budgets off it.

Two things this data does not support, stated because they are the easy misreadings:

  • Not "srv4 is slow." srv4-01 completed in 8.47 min, mid-pack. It is srv4-03 specifically, or a condition on it, not a family property.
  • Not "regen takes 15 minutes on srv4-03." The step was killed at the cap, so the true demand was never observed and cannot be quoted. It is a censored lower bound — the same honesty class gunbc.ci_spec applies to its own operator-signed raises.

The finding worth carrying past this PR. Among completing runs the fleet spread is 7.27 -> 9.55 min = 1.31x, roughly consistent with the ~1.2x spread the budget notes explicitly fund (gunbc_ci_compile_clean_clamp_note, gunbc_ci_floor_batch_wall_budget_note). srv4-03 on identical content is >=2.06x and unbounded above. Budget margins across this repository are sized on that 1.2x assumption; a host at double it is not covered by any of them, and regen at ~7-9.5 min against a 15-minute cap has the least headroom of the family.

For the record, the measured anatomy of a passing 9m21s regen, since "regen is slow" gets attributed to the wrong phase:

~2m00s   job setup (checkout, toolchain, artifact download + verify)
 1m56s   UNATTRIBUTED — nothing printed inside the regen step
 1m11s   output-policy + group-syntax install
    38s  plan resolve                      <- for a TWO-node plan
   110ms naming-hygiene walk               <- #8140; was ~6 min this morning
 2m26s   regen_verify_gate_passes          <- real work
 2m34s   self_host_realized_comparison_staleness_gate_holds  <- real work
   ~20s  receipts + upload

About five minutes is the genuine self-host fixed point. The other four are fixed prelude paid regardless of the plan asking for two nodes — the same cost shape #8140 and #8141 have been removing, surfacing in a different phase.

— sent from eager-cat-841

gunbc-ci-auto-heal and others added 2 commits August 11, 2026 22:51
The floor finally evaluated this branch — main was red before, then an upstream
timeout cascaded — and found exactly one red witness in 8757:
`v1_interpreter_primitive_dispatch_authority_acceptance_contract_holds`.

It is genuinely this PR's. Deleting `v2.lens.inert_lens` removed the last two
rows carrying `EvalCallBridgeFamilySite { module: v2.lens.inert_lens }`, so the
distinct bridge-family count went 9 -> 8 and

    distinct_bridge_family_site_count() == 9

redded. The literal was a population pin — the exact class DESIGN §5 rejects,
and the exact class #7615 removed from this same carrier's census witnesses.
The file's own `closing_contract_note` opens by claiming "The checks here are
structural properties that survive roster growth -- not population pins", so
the clause contradicted its own contract and my deletion is what surfaced it.

Decrementing 9 to 8 would restore green while preserving the defect, so the
control is derived instead: the number of distinct emit-site keys must equal the
number of distinct modules the bridge rows themselves name, and exceed one.

That is not a tautology, because the two sides come from different places — the
left from `dispatch_emit_site_key`'s keying, the right from each row's own
`module` field. Collapsing the families back onto one shared key (the defect the
clause is named for) reds it, 1 != N. Adding or removing a family does not.
The clause now measures what its name claims, in both directions.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018ZrxLJrt8ASGLegiK72fAi
@gunbai-bot

gunbai-bot Bot commented Aug 11, 2026

Copy link
Copy Markdown
Contributor Author

Third and final refinement on the regen red — the mechanism is a lost runner, not contention

My first comment said contention; my second corrected that to host identity. Both understated it. The operator thread pulled GitHub's own classification for the failing job:

"The self-hosted runner lost communication with the server" — naming runner termination, CPU/memory starvation, or network loss as possible causes.

The runner disappeared during regen. And it was not confined to this branch: main's regen failed the same way in the same window on srv3-01.

So the honest split is:

ESTABLISHED
  #8141 is not a semantic regen regression
  the failure varies by runner host
  srv3-01 and srv4-03 lost runner communication during regen
  srv1-03 completed the identical branch subject in 7:00

NOT ESTABLISHED
  that CPU contention caused it
  that the 15-minute step cap caused it
  which of memory, network, runner-service death or host pressure is the cause

I was wrong to say contention was proved. It was inference from a timing overlap, and I repeated the same shape after being corrected once — a confident answer to the adjacent question.

This changes the remedy, which is why it is worth a third comment rather than a footnote: the fix is fleet health, not CI tuning. A vanished runner is not absorbed by raising a cap, adding a retry, or adding a cache — each of those would zero the failure's frequency by construction and hide a host that is failing to stay alive.

A second over-generalization, corrected

I reported "1m56s unattributed" inside the regen step as if it were a property of regen. It is a property of the run I measured. Verified against step boundaries on both:

host step wall step start -> first print policy install genuinely unattributed
srv1-05 9m04s 187 s 71 s 116 s
srv1-03 7m00s 74 s 71 s ~3 s

Both measurements are correct; they are different runs. The thread is right that there is no separate unexplained block in the 7-minute run. And the delta is itself the finding: ~2 minutes of pre-install silence on one host versus ~3 seconds on another, for identical work, accounts for nearly the whole 7:00 vs 9:04 spread. That is the same host-variance axis as the lost-runner failure, showing up as cost instead of as a kill.

What is unchanged and worth restating: #8140 worked. The naming walk is 116 ms, and the ~7 minutes that remain are a different set of operations — ~1m50s executor setup, ~4m50s across the two fixed-point gates, ~20s receipts.

Branch state

Merged latest main (49402500b65) with a merge commit, no rebase, no force-push. Head is now 78bfd93039f, which also carries the one genuine content fix the floor found: the bridge-family red control, derived instead of pinned at a population literal. One CI attempt is running on that integrated head.

— sent from eager-cat-841

@briansrls
briansrls merged commit 210efd5 into main Aug 12, 2026
5 checks passed
@briansrls
briansrls deleted the session/eager-cat-841 branch August 12, 2026 02:04
gunbai-bot Bot pushed a commit that referenced this pull request Aug 12, 2026
Review 51377: the paragraph describing the naming-hygiene walk's measured cost
still said the roster producer "runs inert-lens reachability plus the
construction-justification census" — present tense, about two censuses gunbc#8141
deleted. `claim_executor` already says "(until gunbc#8141 deleted them)" and the §6
bullet in this same document records the deletion in past tense, so the canonical
authority contradicted both.

The measurement itself stands: those censuses WERE part of what made the walk the
most expensive phase when it was measured. What was wrong is the tense, which
asserts a superseded population as the present one — and doing that inside a
deletion diff is the exact failure #8141's own review caught, recorded a few
paragraphs above in this file. Reworded to "and — until gunbc#8141 deleted them —
ran", preserving the cost claim as the historical fact it is.

Swept for other occurrences: the §6 bullet is already past tense; this was the only
stale one.

DESIGN.md is a projection of this authority and is regenerated by
`heal_generated_artifacts`, so it is not hand-edited here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018ZrxLJrt8ASGLegiK72fAi
briansrls pushed a commit that referenced this pull request Aug 12, 2026
… discovery (#8167)

* Make the witness-roster walk demand-directed instead of unconditional pre-plan

claim_executor ran discover_floor_witness_roster_with_snapshot once up front,
BEFORE resolving the plan, on the stated ground that a naming violation should
be "the cheapest possible failure". Measured, that walk is the most expensive
phase in the process: 5.9 min of a 56.5-min ordinary floor (run 31477894666),
and ~6 min of a ~15-min regen whose plan has exactly two nodes.

It is expensive because "naming hygiene" is a misleading label. The four rules
in v2.workflow.floor_naming_hygiene are string predicates over file paths and
line prefixes, but the roster producer they are reached through also builds
module-graph facts, runs a second strict reference-resolution pass, computes
path indexes, and runs inert-lens reachability plus the construction-
justification census. A two-node regen plan paid all of it to discover a roster
it never reads.

Hygiene is a property of the witness ROSTER, so it is now paid by the plans that
have one: the walk moves to the existing `schedules_discovery` predicate, after
the plan's batches settle. The roster is memoized by request digest
(IN_PROCESS_ROSTER_BY_REQUEST), so plans that DO schedule discovery pay exactly
what they paid before — the corpus batch hits the memo this call fills. Plans
that do not schedule discovery pay nothing, and cannot be unhygienic: they have
no roster.

The walk-attempt id is minted unconditionally as before; it is a tracing
coordinate every later phase stamps, and it is not the expensive part.

This also closes the PRELUDE COVERAGE HOLE gunbc.ci_spec
gunbc_ci_floor_batch_wall_budget_note already names: the walk sat outside every
batch budget and could only red at the step cap. It is now inside the region the
plan accounts for, or absent.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018ZrxLJrt8ASGLegiK72fAi

* Update the design authority for the demand-directed hygiene walk (review 51099)

review 51099 (cursor/composer-2.5, REQUEST_CHANGES) correctly caught that #8140
changed documented CI behavior without updating its authority. DESIGN.md
Building & checks stated the opposite of the new code:

  "the executor runs the zero-enrollment naming walk when no discovery batch
   is scheduled"

That clause is superseded here in gunbc.design_document (DESIGN.md is generated
from it; the heal job regenerates the projection).

Both of the reviewer's findings are recorded rather than only the first:

(a) SCOPE — discovery-free plans no longer run the corpus-wide nameability
    rules. Declared as a narrowing, with the honest coverage argument: every PR
    runs the `ci` job, whose plan schedules discovery and walks the full tree,
    so per-PR coverage is unchanged; what is deleted is a second redundant walk
    on regen-only and plan-artifact-only runs. Explicitly NOT backstopped by the
    affected-set falsifier, which has produced no green verdict since
    2026-08-03.

(b) ORDERING — for plans that DO schedule discovery the walk now runs after
    plan resolve/eval, so a naming violation pays ~0.5 min of plan resolution
    before refusing. The reviewer asked whether the trade is intentional: it is,
    and it is priced — ~0.5 min later on the refusing path against ~6 min saved
    on every regen.

A dissolve-on is recorded: the clause, the separate walk, and the `__`-basename
rule all retire when the placement rules move to canonical source ingestion and
test identities derive from parser-produced declarations.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018ZrxLJrt8ASGLegiK72fAi

* chore: regenerate drifted generated artifacts (ci auto-heal)

* Repair the stale output-policy comment left by the walk move (review 51101)

review 51101 (cursor/composer-2.5) caught claim_executor.rs:10283 still
asserting "The walk still runs before plan evaluation, so a naming violation
stays the cheapest failure" — the exact opposite of what this PR does.

Non-blocking as a defect, but it is the same class the PR itself is about: a
comment standing as authority for behavior the code no longer has. #8140's
own receipt was a block comment whose stated premise had been false for
months.

The paragraph's real subject — install output policy BEFORE the walk so the
whole-tree read is funnelled rather than emitting ~2.3k `[file] read` lines —
is unchanged and still correct; the walk simply moved further away from it.
Rewritten to say that, rather than deleted, so the ordering requirement keeps
its rationale.

Swept the rest of claim_executor.rs / cli_run.rs for other assertions of the
old ordering: the only remaining hits are this PR's own comment describing the
prior behavior in the past tense.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018ZrxLJrt8ASGLegiK72fAi

* Delete supply-side lens enforcement from floor discovery

Two censuses ran inside `discover_floor_witness_roster` — the inert-lens
reach walk and the construction-justification census. Both asked a question
about who authored a lens, and both answered it by acquiring a whole-corpus
module graph. Every discovery run paid that: every PR, regen, and every
coordinated scoped worker, all of which wanted a witness roster and nothing
else. The unit of computation was the world; the unit of fact was one
module's authorship.

Deleted end-to-end, not merely unwired:

- `v2.lens.inert_lens` (its `.dag` surface was two self-recursive stubs,
  `fn f() { f() }`, reachable only because the interpreter intercepted them)
- its two host builtins, both interpreter dispatch registrations, the
  generated bridge family, the `04_method` type-table entries and the
  `std.primitives` roster rows
- the `InertLens` registry variant, its registry row, contract row and
  `lens_module_gate` invariant surface
- `inert_lens_modules`, `inert_lens_modules_legacy`, `lens_justification_census`,
  `unjustified_lens_modules`, `declares_construction_justification` and both
  floor refusal arms (`cli_run.rs` and `floor_discovery_snapshot.rs`)
- the long witness and its frozen deferral row (shrink logged)

`build_module_graph_facts_live` still runs on this path and this change does
not claim otherwise: effect-reach derivation and the cross-worker snapshot
transport both consume it. `refuse_on_module_graph_read_refusals` and its two
red controls are retained and re-homed, since the fail-closed arm now guards
those consumers rather than the deleted censuses.

This is a scope narrowing, not a climb. A new lens with no witness, and a lens
recording no `construction_justification`, are both writable again and nothing
detects either. Declared in DESIGN §6 with its next-rung trigger: authorship
belongs on the module's own declaration, checked where the module is already
parsed, rather than reconstructed corpus-wide by a consumer that wanted a
roster.

Two citations repaired rather than left stale (§3): the roadmap acceptance note
cited a shadow witness this change renames and weakens, and `source_authority`
cited the inert-lens stubs as its example of host interception.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018ZrxLJrt8ASGLegiK72fAi

* chore: regenerate drifted generated artifacts (ci auto-heal)

* Reconcile the plan carrier and stale comments with the deleted enforcement

Review 51113 (cursor/composer-2.5) found two second authorities still claiming
enforcement the code no longer performs — DESIGN §3 dual representation and §4b
rung honesty. Both confirmed against the current head, both fixed.

`gunbc.plans.construction_justification_rule` said the presence check "run[s] in
`discover_floor_corpus_rows`" and listed it as current status. Its retirement
condition also named that check as its trigger, so after the deletion the
condition could never fire — an unreachable lifecycle claim that structurally
cannot report itself satisfied. The plan now leads with a supersession notice,
§2/§3/§4 are marked historical rather than reworded, and a new §5 records what
was deleted, what it costs (a lens added tomorrow with no justification lands
green; the 35-module classification in §4 is a historical measurement, not a
maintained invariant), and the next-rung trigger. The retirement condition is
replaced with a reachable one and says why.

`floor_discovery_snapshot.rs`'s consumer census still listed the two gates and
still called the roster walk "pre-plan", which #8140 already made false.

Swept the rest rather than fixing only what was reported: eight comments in
`cli_run.rs` and one in `claim_executor.rs` named the inert-lens reach as a live
consumer of the observation rows, the reference-edge producer, and the selection
tier. Repointed to the consumers that remain.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018ZrxLJrt8ASGLegiK72fAi

* Delete the dead lens classifier and date-scope the cost claim

Two bounded review corrections.

`is_top_level_lens_module` survived the census deletion with no remaining
caller — only its own definition. It compiled clean because the crate carries a
blanket allow, which is exactly why the residue needed finding by reading rather
than by warning. Deleted; the PR's bar is end-to-end with zero residue, and a
dead classifier left behind is the pattern this change exists to close.

The new DESIGN paragraph asserted the censuses charged "every discovery run — on
every PR, on regen, in every coordinated worker". True before #8140, false on
this PR's base: #8140 made the roster walk demand-directed, so a discovery-free
plan such as regen already stopped paying. Both DESIGN and the plan carrier now
split the claim by era — unconditional before #8140, regen exempt after it, the
censuses burdening every remaining discovery-bearing execution until this
deletion. A change removing stale supply-side enforcement must not land a fresh
stale assertion in the same diff.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018ZrxLJrt8ASGLegiK72fAi

* chore: regenerate drifted generated artifacts (ci auto-heal)

* Repoint the sibling plan authorities that still advertise the deleted census

Review 51154 caught a consistency failure in my own work: I applied the
stale-authority fix to `construction_justification_rule.dag` and then judged
`inert_layer_lens.dag` "historical prose" without running the same test on it.
It fails that test. It is a registered plan whose §3 tells a future worker that
Tier 1 is "buildable now (reuse #5433)" and to extend `inert_lens_modules` —
a function this PR deletes — and §7 goes further, advising them to extend it
behind a flag rather than fork it. Someone following that plan would go looking
for machinery that is gone.

Four rows repointed rather than deleted, since the design reasoning survives
even though its cited mechanism does not:

- §3 Tier 1 now leads with the supersession, names `v2.lens.module_graph` as the
  surviving reachability authority, and says plainly that "reuse the existing
  walk" now means "build the walk", which is a larger job than the paragraph
  reads.
- §6's reuse map repoints the transitive-reachability-BFS row off
  `cli_run.rs:2558-2606` — a positional citation into a file that has since
  moved several thousand lines, which is the §3 rot mode exactly.
- §7's seed caveat drops the extend-behind-a-flag advice and states the real
  constraint: whatever Tier 1 becomes must not reintroduce a corpus-wide walk
  inside floor discovery, because the placement was the defect, not the walk.
- §5's "fail closed exactly as #5433 does" moves to past tense; no lens-inertness
  gate runs today.

Also marked the two remaining historical citations, in the same plan's landed
doc-graph receipt and in `axiom_syllogism_lens.dag`'s precedent table, so every
surviving mention of a deleted symbol carries its deletion beside it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018ZrxLJrt8ASGLegiK72fAi

* chore: regenerate drifted generated artifacts (ci auto-heal)

* Derive the bridge-family split control instead of pinning its population

The floor finally evaluated this branch — main was red before, then an upstream
timeout cascaded — and found exactly one red witness in 8757:
`v1_interpreter_primitive_dispatch_authority_acceptance_contract_holds`.

It is genuinely this PR's. Deleting `v2.lens.inert_lens` removed the last two
rows carrying `EvalCallBridgeFamilySite { module: v2.lens.inert_lens }`, so the
distinct bridge-family count went 9 -> 8 and

    distinct_bridge_family_site_count() == 9

redded. The literal was a population pin — the exact class DESIGN §5 rejects,
and the exact class #7615 removed from this same carrier's census witnesses.
The file's own `closing_contract_note` opens by claiming "The checks here are
structural properties that survive roster growth -- not population pins", so
the clause contradicted its own contract and my deletion is what surfaced it.

Decrementing 9 to 8 would restore green while preserving the defect, so the
control is derived instead: the number of distinct emit-site keys must equal the
number of distinct modules the bridge rows themselves name, and exceed one.

That is not a tautology, because the two sides come from different places — the
left from `dispatch_emit_site_key`'s keying, the right from each row's own
`module` field. Collapsing the families back onto one shared key (the defect the
clause is named for) reds it, 1 != N. Adding or removing a family does not.
The clause now measures what its name claims, in both directions.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018ZrxLJrt8ASGLegiK72fAi

* Delete the orphan-helper census and the `__` filename rule from floor discovery

Two more required-path phases whose cost was denominated in the corpus and
whose answer nobody consumed. Net −1446/+39.

**The `__`-basename rule.** A census found ZERO offending basenames in the tree
and no stated rationale anywhere for the ban — it guarded an empty population.
It was not free: reaching the predicate meant collecting every `.dag` path in
the corpus and then resolving a SEPARATE `.dag` entry
(`FLOOR_NAMING_HYGIENE_ENTRY` -> `floor_filename_hygiene_refusal_via_producer`)
on every discovery-bearing run, so a zero-population style rule cost a whole-tree
walk plus an entry resolve. Both are gone, along with the snapshot's
`naming_hygiene_refusal` field and the two witness controls.

**The orphan-helper census.** It walked every `*_test.dag`, parsed each one,
projected `DeclSurface`/`ModuleSurface` values, resolved a second interpreter
context, and ran a fuelled reachability fixpoint against a hand-authored
cross-module export exception roster — to decide whether a plain helper in a
test file was referenced. An unreferenced test helper is dead-code hygiene. It
is not evidence that the compiler or the tests are correct, and it did not
justify a recurring whole-corpus traversal on a required path.
`gunbc.test_module_hygiene` goes 661 -> 106 lines, the Rust bridge 835 -> ~320,
and `test_module_hygiene_scaffold.dag` deletes whole (its dissolution obligation
is discharged by the deletion, not carried forward).

**Scope narrowing, declared rather than implied.** An unreferenced test helper
and a `__` basename are both writable again and nothing detects either. The
orphan class had fourteen enrolled witnesses and they are deleted with it —
§4b permits that only because the class is ABANDONED, not climbing: there is no
higher rung for the evidence to guard, and keeping fixtures for a fold nothing
calls would be specification-without-execution one level up. Recorded in DESIGN,
in the module authority note, and in the witness file that used to hold them.

**What survives, and why:** the `test fn` placement rule, the file-grain expand
half, and `failure_receipt_companion` — the naming convention `claim_executor`
actually invokes. That last one was only ever exercised through the census's
whole-corpus walk, so it would have silently lost its executing consumer; it
gains direct unit and `.dag` witnesses here instead.

**Not in scope, deliberately:** the line-scanned test-identity derivation. That
is the second parser, and replacing it needs the canonical parser to retain the
`test` marker — which `drop_leading_test_marker` discards, because `test` is a
live module-path segment (`test.claim.*`, `extdeps.test.*`) and cannot be lexed
as a keyword. `realization_attempt.dag` already names the prerequisite: a
contextual-keyword terminal in `GrammarExpr`, which does not exist yet.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018ZrxLJrt8ASGLegiK72fAi

* Refresh the two authority notes the deletion falsified

Review 51262 found `floor_naming_hygiene_note` still asserting, in a file this
PR edits, that the module sits where it does so "orphan/filename hygiene resolve
stays free of the Filesystem service closure" and that it dissolves the hand-Rust
mirror `check_floor_filename_hygiene`. Neither survives the deletion: there is no
orphan census and no filename-hygiene resolve left to keep free of anything, and
the filename half of that dissolution obligation is discharged by deletion rather
than by dissolution — no equivalence receipt was ever owed for a rule with a
zero-row population.

The note now states what remains true instead, including the part worth carrying:
the test-decl line scan is still a second parser, and replacing it is not a
refactor of this module — it needs the canonical parser to retain the `test`
marker, which `drop_leading_test_marker` discards, and `test` cannot become a lex
keyword because it is a live module-path segment.

Swept for the same class rather than fixing only what was reported:
`floor_discovery_dissolve_trigger` still described "the producer and
filename-hygiene entries" as two typed entry values resolved through
`resolve_workspace_entry`. There is one now.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018ZrxLJrt8ASGLegiK72fAi

* chore: regenerate drifted generated artifacts (ci auto-heal)

* Carry the deleted-witness rationale as an authored annotation, not a prose data row

`prose_row_introduction_gate` refused this change: `test_module_hygiene_orphan_witness_deletion_note`
is a `data NAME_note: String` declaration introduced under
`dag/test/claim/test_module_hygiene_hand_rust_equivalence_witness_test.dag`, a path
already on `gunbc.prose_row_frontier` `prose_row_migration_scope`. DESIGN §4c: prose
is not forbidden, unclassified prose is, and a `String` declaration whose sole
purpose is commentary is misplaced data.

The rationale is irreducible — it records why fourteen witnesses were deleted with
the machinery they tested, and why §4b's dissolution-on-climb rule does not save
them (the class is abandoned, not climbed) — so it becomes a leading `//` block
attached to the declaration below it. The PR-number and section references in the
prose are dropped rather than carried across: those are exactly the machine-consumed
facts §4c says belong in a typed carrier, not in commentary.

The two other in-scope rows this change touches (`test_module_hygiene_authority_note`,
`failure_receipt_companion_note`) are pre-existing declaration names whose content was
rewritten, not introductions, and the gate does not refuse them. They are left as
rows rather than swept here.

Green by execution: the edited witness parses and
`test_module_hygiene_file_grain_empty_function_holds` returns `true`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018ZrxLJrt8ASGLegiK72fAi

* chore: regenerate drifted generated artifacts (ci auto-heal)

* Make the deleted censuses past tense in the Building & checks paragraph

Review 51377: the paragraph describing the naming-hygiene walk's measured cost
still said the roster producer "runs inert-lens reachability plus the
construction-justification census" — present tense, about two censuses gunbc#8141
deleted. `claim_executor` already says "(until gunbc#8141 deleted them)" and the §6
bullet in this same document records the deletion in past tense, so the canonical
authority contradicted both.

The measurement itself stands: those censuses WERE part of what made the walk the
most expensive phase when it was measured. What was wrong is the tense, which
asserts a superseded population as the present one — and doing that inside a
deletion diff is the exact failure #8141's own review caught, recorded a few
paragraphs above in this file. Reworded to "and — until gunbc#8141 deleted them —
ran", preserving the cost claim as the historical fact it is.

Swept for other occurrences: the §6 bullet is already past tense; this was the only
stale one.

DESIGN.md is a projection of this authority and is regenerated by
`heal_generated_artifacts`, so it is not hand-edited here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018ZrxLJrt8ASGLegiK72fAi

* Drop orphan/helpers from the snapshot consumer-census row

Review 51382: the module header still listed "Naming hygiene, orphan/helpers" as
what the demand-directed roster walk reads, in a file this PR edits. The
orphan-helper census and the `__`-basename rule are deleted here, so the row named
work that no longer happens — doc drift inside the diff that removed the subject.

The row now names what the walk actually still does: `test fn` placement hygiene,
producer roster, module-graph facts, effect-reach derivation, with the deletion
noted so a reader is not left wondering where the other two went.

The same review's DESIGN.md finding is real and is NOT fixed by hand: DESIGN.md is a
generated projection of `dag/gunbc/design_document.dag`, whose wording was corrected
in 69009a3. Editing the projection directly would author bytes no authority
produced, and `heal_generated_artifacts` reverts exactly that. Heal last ran against
the previous head (a5b441b, before the authority fix); its run on this head
re-projects the corrected sentence.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018ZrxLJrt8ASGLegiK72fAi

* chore: regenerate drifted generated artifacts (ci auto-heal)

* Repair the comment whose contrast named a deleted function

Review 51390: a comment justifying why a wiring test calls the producer seam argued
by contrast with `floor_filename_hygiene_refusal_for_paths` — which this PR deletes.
The live half of the argument still holds (the seam is where a wire-contract
violation is observable, and the rule itself is owned content-side by
`floor_discovery_equivalence_misplaced_wire_contract_refuses_holds`, so re-deriving
it here would be a second representation). Only the contrast was dangling.

Rewritten to state the positive reason, with the retired comparison noted as
history rather than silently dropped: a reader who remembers the old sentence
should find out where it went, not wonder whether the argument changed. No code,
assertion, or behavior change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018ZrxLJrt8ASGLegiK72fAi

---------

Co-authored-by: gunbc-ci-auto-heal <gunbc-ci-auto-heal@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: gunbai-bot[bot] <289086189+gunbai-bot[bot]@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant