Skip to content

Floor cost repair: one producer, not three witnesses — the compile-phase frontier standing crosses the 500ms CPU ceiling and makes the required floor a coin flip - #10133

Merged
briansrls merged 1 commit into
mainfrom
session/warm-ibex-636
Sep 2, 2026

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor

The coin flip, located

required-floor refused three of the last four main runs with interrupted_cpu_deadline in {17, 15, 2} and verdict=FloorRefused, while run 33673806092 on the same corpus passed. The refused rows are the same seventeen identities in three modules, and the run log names them BUDGET-REFUSED in 503..506ms against v2.workflow.required_floor required_floor_claim_cpu_safety_limit_ms (500):

module rows
test/claim/self_host_compile_phase_live_gate_witness 13
test/claim/self_host_compile_phase_frontier_witness 2
test/claim/compiler_frontend_program_status_witness 2

On the passing run those same rows completed at 452–496ms. Nothing about them changed between runs; they sit inside one scheduler quantum of the ceiling, so which side a PR lands on is ambient fleet load.

One root, not three

All seventeen rows spend their whole cost in gunbc.self_host_compile_phase_frontier compile_phase_frontier_standing. The first two modules call it directly; the third reaches it through gunbc.compiler_frontend_program_status milestone_status, whose SelfHostWholeCorpusPopulationDerived arm asks the same question. So the repair is in the producer, not in any witness — and v2.workflow.required_floor names exactly this remedy ("reduce what the witness reaches for"; relocation is forbidden).

What was fixed

Attributed with the instrument that already exists (claim_batch --entry <witness> --functions <rows>, whose [witness] line reports per-row CPU), using probe modules that call one callee at a time. Three terms carried it, all shapes DESIGN §6 fixes on sight rather than on the realized n:

  • newly_exposed_identity_admitted — a cheap identity guard conjoined ahead of an expensive cause derivation. && in this substrate evaluates BOTH operands (v1_interpreter eval_expr_inner evaluates left and right before applying the operator), so the guard narrowed the result without guarding the work: the derivation, two of whose arms scan the whole previous board, ran once per (added identity × disposition row) pair. Selecting the rows first is exactly equivalent — any(xs, p && q) = any(filter(xs, p), q) for boolean p, q.
  • board_added_identities / board_removed_identities / phase_ratchet — set difference answered by any over the other population per row, making the comparison count the product of two populations that grow with every receipt. Now one identity-keyed membership index per population.
  • identity_is_hop_relocation — re-derived diagnostic_semantic_key for every row of added and removed on every call. Those keys depend on the transition, not on the identity under test, so they are derived once at apply_regression_transition — the least common ancestor of the demand (DESIGN §2).

Result on this container: the worst row of each module falls from ~275–296ms to ~60–67ms.

Evidence

Every row of all three witness modules passes: 13/13, 35/35, 34/34 — including the frontier witness's discriminating reds over newly_exposed_identity_admitted, and the live gate's planted-identity and equal-cardinality-swap refusals. Those are the probes that would go red if the restructure had changed the answer.

No semantic change is claimed and none was found. The receipts' sealed identities, prefix digests and census digests are all derived, not literals, so an equal standing is an equal projection: gunbc.design_ledgers renders only latest and its output is unchanged. No consumer outside this module named any signature that moved, and the module has no stage0 Rust mirror, so no regeneration is implied.

What this does not do

It does not bring the population to the ~1ms median an ordinary witness should hold. The rows remain above required_floor_claim_cost_line_ms (100, diagnostic only) once the fleet's host factor is applied. The residue is the receipt-sealing data fold, paid once per claim — a different subject from the per-call scans repaired here.

Worth knowing beyond this diff

&& and || are strict in the interpreter and short-circuiting in the emitted Rust. The two agree on the answer and differ in what they evaluate. Every cheap_guard && expensive_derivation written in the corpus by an author carrying the Rust habit over is paying the expensive half unconditionally. This diff repairs one instance; the class is not censused.

🤖 Generated with Claude Code

https://claude.ai/code/session_01WEFMasGn2cAGCMJW2694pb

…s to ~60ms so three required-floor witness modules stop crossing the 500ms CPU ceiling

THE COIN FLIP, LOCATED. `required-floor` refused three of the last four main runs
with `interrupted_cpu_deadline` in {17, 15, 2} and `verdict=FloorRefused`, while
run 33673806092 on the same corpus passed. The refused rows are the same
seventeen identities in three modules, and the run log names them
`BUDGET-REFUSED in 503..506ms` against
`v2.workflow.required_floor` `required_floor_claim_cpu_safety_limit_ms` (500):

  test/claim/self_host_compile_phase_live_gate_witness      13 rows
  test/claim/self_host_compile_phase_frontier_witness        2 rows
  test/claim/compiler_frontend_program_status_witness        2 rows

On the passing run those same rows completed at 452-496ms. Nothing about them
changed between the runs; they sit inside one scheduler quantum of the ceiling,
so which side of it a PR lands on is ambient fleet load. That is the coin flip.

ONE ROOT, NOT THREE. All seventeen rows spend their whole cost in
`gunbc.self_host_compile_phase_frontier` `compile_phase_frontier_standing` over
the persisted receipt series. The first two modules call it directly; the third
reaches it through `gunbc.compiler_frontend_program_status` `milestone_status`,
whose `SelfHostWholeCorpusPopulationDerived` arm asks the same question. So this
is one repair, and it is in the producer rather than in any witness.

MEASURED WITH THE INSTRUMENT THAT ALREADY EXISTS -- `claim_batch --source-root dag
--source-root src/v2 --entry <witness> --functions <rows>`, whose `[witness]`
line reports per-row CPU. Attribution was done by probe modules that call one
callee at a time; the interpreter memoizes a pure call within one claim and not
across claims, so each probe pays its own cost and the deltas are marginal
figures. Three terms carried it, all of them shapes DESIGN section 6 fixes on
sight rather than on the realized n:

  `newly_exposed_identity_admitted` -- a cheap identity guard conjoined ahead of
  an expensive cause derivation. `&&` in this substrate evaluates BOTH operands
  (`v1_interpreter` `eval_expr_inner` evaluates left and right before applying
  the operator), so the guard narrowed the result without guarding the work: the
  derivation, two of whose arms scan the whole previous board, ran once per
  (added identity x disposition row) pair. Selecting the rows first is exactly
  equivalent -- for boolean p and q, `any(xs, p && q)` = `any(filter(xs, p), q)`.

  `board_added_identities` / `board_removed_identities` / `phase_ratchet` -- set
  difference answered by `any` over the other population per row, so the
  comparison count is the PRODUCT of two populations that grow with every
  receipt. Now one identity-keyed membership index per population.

  `identity_is_hop_relocation` -- re-derived `diagnostic_semantic_key` for every
  row of `added` and `removed` on every call. The keys depend on the transition,
  not on the identity under test, so they are now derived once at
  `apply_regression_transition`, which is the least common ancestor of the demand.

NO SEMANTIC CHANGE IS CLAIMED AND NONE WAS FOUND. Every row of all three witness
modules passes, including the frontier witness's discriminating reds over
`newly_exposed_identity_admitted` and the live gate's planted-identity and
equal-cardinality-swap refusals. The receipts' sealed identities, prefix digests
and census digests are all derived, not literals, so an equal standing is an
equal projection: `gunbc.design_ledgers` renders only `latest`, and its output is
unchanged. No consumer outside this module named any signature that moved, and
the module has no stage0 Rust mirror, so no regeneration is implied.

WHAT THIS DOES NOT DO. It does not lower the population toward the ~1ms median an
ordinary witness should hold; the rows land near 60ms locally and remain above
`required_floor_claim_cost_line_ms` (100) once the fleet's host factor is applied.
The residue is the receipt-sealing data fold, which is paid once per claim and is
a different subject from the per-call scans repaired here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WEFMasGn2cAGCMJW2694pb
@gunbai-bot

gunbai-bot Bot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor Author

Verified on the floor, at identity grain rather than by the green

Run 33688872492, required-witnesses-floor:

required-floor: planned=3498 executed=3498 terminal=3498 passed=3423 known_red_held=26
                failed=0 interrupted_before_verdict=0 interrupted_cpu_deadline=0
                interrupted_wall_deadline=0 completed_over_cost_requirement=0 route_gap_held=49
required-floor: verdict=FloorClean

The accounting closes exactly, which is stronger than the pass:

run planned passed interrupted_cpu_deadline passed + interrupted
main 33681059548 3498 3406 17 3423
main 33680659583 3498 3408 15 3423
this PR 3498 3423 0 3423

Same planned, same known_red_held=26, same route_gap_held=49 — so the rows that were being interrupted are the rows now passing, and nothing was dropped from the population to get there.

All seventeen rows also left the over-cost report. That report lists every row above required_floor_claim_cost_line_ms (100); the three modules appear zero times in it now, where the previously-passing run had fourteen of them at 452–496ms occupying the top of the ranking. The new worst row on the floor is v2.test.execution.rust_emit_host_call_equals_eval at cpu_ms=395 — unrelated to this diff, and now the head of that population.

One read I had to correct mid-check: my first attempt to grep this log returned 0 BUDGET-REFUSED, which looked like confirmation. The log was empty because the run had not finished, so that zero was a fact about the fetch and not about the run. The figures above come from the completed job.

— sent from warm-ibex-636

@briansrls
briansrls merged commit 7760909 into main Sep 2, 2026
6 checks passed
@briansrls
briansrls deleted the session/warm-ibex-636 branch September 2, 2026 22:45
gunbai-bot Bot pushed a commit that referenced this pull request Sep 2, 2026
gunbai-bot Bot pushed a commit that referenced this pull request Sep 2, 2026
gunbai-bot Bot pushed a commit that referenced this pull request Sep 2, 2026
gunbai-bot Bot pushed a commit that referenced this pull request Sep 2, 2026
@briansrls
briansrls restored the session/warm-ibex-636 branch September 2, 2026 22:53
gunbai-bot Bot pushed a commit that referenced this pull request Sep 2, 2026
…10133 moved

The roster admits a row on MEASURED serve below MEASURED recompute. My present-vs-absent
pair (33669782300 / 33672356296) measured the recompute side against the pre-#10133
producer, and #10133 made exactly that recompute cheaper -- so the margin the receipt
establishes is an upper bound on the margin that exists now. The measurement is not wrong;
its subject moved.

The argument for landing it anyway was that the serve side is nullary with an empty
argument row and is therefore unlikely to have crossed. That inference is the one this
roster exists to refuse: gunbc#10094 withdrew three rows enrolled on exactly that reasoning
after all three were plausible and all three were wrong on measurement. A roster whose
discipline is measured-not-inferred cannot admit its next row on a structural argument
about which side sits near the floor. Ruling by bright-ram-778, holding me to this file's
own rule rather than a new one.

So floor_pure_producer_share.dag returns to main untouched, and the accessor's comment now
says plainly that being preparation-forceable is a property of the declaration and NOT a
claim that a row exists -- an unlanded citation is indistinguishable at the citing end. It
also records that the census figures beside it predate #10133.

What remains is claims 1 and 2, whose correctness never depended on a cost result.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RuWuQWB6MPkY7sNM4jEqAy
gunbai-bot Bot pushed a commit that referenced this pull request Sep 3, 2026
…he absent arm of the share measurement (#10108)

* One home for the standing of the persisted frontier series, and the absent arm of the measurement that decides whether it may be shared

FLOOR-COST-500MS, first of two commits on this branch, and the split is the
measurement rather than tidiness. This commit carries the REWIRE ONLY. The
roster row that shares it lands in the next commit, so this branch's own floor
run is the ABSENT arm of the present-vs-absent A/B that
v2.workflow.floor_pure_producer_share requires before any row is enrolled --
two runs differing in exactly one roster line, on otherwise identical trees.

WHAT IS CONSOLIDATED. gunbc.self_host_compile_phase_frontier
compile_phase_frontier_standing judges ANY series of receipts, and its own
witnesses exercise it against perturbed ones. But there is exactly one
PERSISTED series in this repository, and "the standing of that series" was
being re-spelled at five call sites across four modules. That is one fact with
five homes, which DESIGN section 3 puts in one place; the nullary
current_compile_phase_frontier_standing is that place. The general function is
untouched and every consumer that judges a hypothetical series still calls it
directly -- that distinction is why both spellings stay.

THE DISCRIMINATING RED IS ENROLLED, because the consolidation has exactly one
failure mode and it is silent. An accessor bound to a DIFFERENT series -- a
truncated one, a fixture one, a second declaration added later -- keeps
compiling and keeps answering, about the wrong subject, and no existing claim
in either consuming closure notices, because they all read the accessor and
would agree with each other about the wrong thing.
the_persisted_standing_accessor_agrees_with_the_general_fold_on_the_named_series
is the only executing place the two spellings are compared; rebind the accessor
to the skip(n: 1) perturbation the sibling live-gate witnesses already use and
it reds while nothing else does.

WHY NOW, and this half is a consequence rather than the reason. The required
floor builds a fresh evaluation frame per claim, so this standing was
re-derived from scratch in every claim that reads it. Measured on main run
33664768371 attempt 1, required_floor_cross_claim_demand.tsv: 22 claims, 36
evaluations, 6508ms inclusive-of-callees, with phase_board_series_ratchet
(23/23) and apply_regression_transition (23/23) nested underneath. Being
NULLARY, the consolidated accessor is preparation-forceable and carries an
EMPTY argument row -- which is what makes it a candidate for the share tier at
all, since the general function's List<CompilePhaseFrontierReceipt> argument
would have to be reified and structurally verified in every consuming frame,
the shape that roster's header records LOSING. Re-derive with the instruments
named there, never from these sentences.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RuWuQWB6MPkY7sNM4jEqAy

* The present arm: enrol the persisted frontier standing as a warm share row, against this branch's own absent measurement

FLOOR-COST-500MS, second of two commits. The parent commit rewired five call
sites into one nullary accessor and changed no roster; its floor run is the
ABSENT arm. This commit adds exactly one roster line, so the two runs differ in
one line on otherwise identical trees and the present-vs-absent comparison
v2.workflow.floor_pure_producer_share requires is a property of this branch
rather than a promise in a message.

THE ABSENT ARM, measured on run 33669782300 attempt 1 (this branch at
0205a4d, required-witnesses-floor SUCCESS, required_floor_claim_cost.tsv,
3499 executed rows, cost_basis=cpu). Corpus p50 2ms; 18 rows at or above 400ms
and none at or above 450. The two closures that consume this standing:
test.claim.self_host_compile_phase_live_gate_witness 13 rows at 418-445 and
test.claim.compiler_frontend_program_status_witness 8 rows at 382-409, against
the 500ms per-claim CPU ceiling. 445 is 89% of budget, which is the state this
lane exists to end.

WHY THIS ROW IS ENROLLABLE WHERE THE CENSUS'S CANDIDATE WAS NOT. The census
ranked the GENERAL fold, whose List<CompilePhaseFrontierReceipt> argument a
serve must reify and structurally verify in every consuming frame -- the shape
this roster's own header records LOSING, where the serve cost more than the
recompute and two rows were removed. The parent commit's consolidation is what
changes that: the accessor is NULLARY, so its argument row is EMPTY. There is
nothing to reify and nothing to verify per serve, and being
preparation-forceable the fill lands outside the fold. The losing shape is
absent rather than gambled on.

ONLY THE OUTERMOST PRODUCER OF THE NEST IS ENROLLED, and the two fatter rows
underneath it are deliberately absent rather than overlooked:
phase_board_series_ratchet (23 claims / 23 evals) and
apply_regression_transition (23 / 23) are CALLEES of this standing, the
census's cost columns are inclusive of callees, and its own summary line says
they do not sum. Three roster rows would claim one saving three times.

THIS ROW IS THE FIRST ENROLLED FROM A CANDIDATE THE CENSUS PRODUCED rather
than from an incident on somebody else's pull request.
gunbc.recurring_failure_mode recurrence_ledger_scoped_below_the_recurrence
names that discovery path as the class: a roster whose rows each trace back to
a budget refusal landing on a passing lane is a roster whose discovery
instrument is blind.

WHAT IT DOES NOT DO. It does not retire gunbc.rung_drop
floor_cost_claim_qualification_unavailable. A per-closure constant is identical
on two runs of one tree by construction, so removing it removes the LEVEL; the
environment term stays and the per-claim line remains an attempt-safety
boundary rather than a claim-cost verdict. This buys headroom, which is not a
rung claim.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RuWuQWB6MPkY7sNM4jEqAy

* Take the standing, not the series: the residual re-derivation the present arm exposed, and the measurement that found it

FLOOR-COST-500MS, third commit. The present arm measured the enrolment a large
win AND left three rows crossing, and this commit is the repair for the second
half. Both halves come from the same run and neither is inferred.

WHAT THE PRESENT ARM MEASURED (run 33672356296 attempt 1, head a71dfb3,
required_floor_claim_cost.tsv, joined by identity against run 33669782300's
absent arm on the same branch). The serve is far below the recompute wherever
it applies:

  test.claim.compiler_frontend_program_status_witness   382-409ms -> 50-63ms
  test.claim.self_host_compile_phase_live_gate_witness  7 of 13 rows 418-445 -> 0-81
  corpus rows at or above 400ms                         18 -> 10
  warm fill                                             disposition=Stored, 501ms, billed to preparation

THE ROWS THAT DID NOT IMPROVE ARE THE SUBJECT OF THIS COMMIT, and they are
exactly the rows that never reached the shared value:
live_tree_frontier_verdict took a SERIES and re-derived
compile_phase_frontier_standing itself at every call, so a caller that meant
the persisted series could not reach the warm entry no matter what the roster
said.

THE PARAMETER WAS WIDER THAN THE BODY. Every use of `series` in that fold was
one call to compile_phase_frontier_standing; nothing else in it ever looked at
a receipt list. A parameter a body merely passes through to a producer is a
declared dependency on the wrong thing, and the width was not free -- it put
the shared value one level INSIDE the parameter, where nothing can serve it.
Taking the standing makes the declared dependency equal the consumed one.

IT ALSO MAKES THE TWO SUBJECTS UNCONFUSABLE AT THE CALL SITE, which is a
correctness gain rather than a cost one. Under the old signature
`series: self_host_compile_phase_frontier_receipts` and
`series: self_host_compile_phase_frontier_receipts |> skip(n: 1)` are one
character apart and mean entirely different subjects. Now a caller judging the
live tree against THE persisted frontier passes
current_compile_phase_frontier_standing(), and a caller judging a PERTURBED
series passes compile_phase_frontier_standing(series: <perturbed>) and gets a
genuine re-derivation. Every perturbation witness keeps its discriminating
power: the empty-series and dropped-genesis probes still fold their own series
and still refuse.

WHY THE ROWS THAT ROSE ARE NOT A SERVE REGRESSION, measured rather than
asserted, because this is the reading the enrolment had to survive. 72 control
rows at or above 250ms OUTSIDE the three consuming modules -- rows the roster
row cannot serve -- moved present/absent by p50 1.04, p90 1.09, max 1.17. The
consuming-module rows that rose moved 1.09-1.17, inside that band. The whole
run was hotter; the rows that rose are rows the share does not reach, and they
moved with the corpus rather than against the cache.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RuWuQWB6MPkY7sNM4jEqAy

* Put the measured receipt in the roster row, replacing the word "expected"

The row landed saying what enrolment was EXPECTED to buy, because the run that
measures it had not finished. It has: codex/gpt-5.6-sol's REQUEST_CHANGES on
gunbc#10108 (review 58865) is correct that an empty argument row establishes
cheap KEYING and establishes nothing about the cost of reifying and serving a
CompilePhaseFrontierStanding, and that this roster's own criterion demands the
present-vs-absent receipt before enrolment rather than after.

So the row now carries the two-run identity join, the control that separates a
serve measurement from a hot-and-cold pair of runs, and its own coverage
boundary. Nothing here is a new claim: the numbers are the artifacts of runs
33669782300 and 33672356296, both attempt 1, on this branch's own trees.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RuWuQWB6MPkY7sNM4jEqAy

* The last row paying for itself: ask the standing, not the ratchet, and leave exactly one payer in the corpus

FLOOR-COST-500MS, fifth commit, and it removes the corpus's single highest
claim. After the enrolment and the parameter narrowing, floor run 33675891779
attempt 1 (head 020ca87, required-witnesses-floor SUCCESS) puts ZERO rows at
or above 500ms and ONE at or above 450: this one, at 461ms -- 92% of the
ceiling, in a 3499-row corpus whose p50 is 2ms. Leaving it would hand the next
lane a smaller copy of the problem this one was chartered to remove.

IT WAS ASKING THE INNER PRODUCER. phase_board_series_ratchet(series: <the
persisted series>) re-derives the whole fold in this claim's own frame, while
the value the fold produces is exactly what the share tier now holds warm one
level up. Asking current_compile_phase_frontier_standing() reads that value.

THE ASSERTION GETS STRICTLY STRONGER RATHER THAN WEAKER, which is worth
checking rather than assuming: CompilePhaseFrontierValidated is
PhaseBoardSeriesHeld AND a latest receipt existing. A held series with no
latest, which the old spelling accepted, now refuses -- the fail-closed
direction.

WHAT IT COSTS, DECLARED. This row now trusts the served standing instead of
re-deriving it, so a store serving a WRONG value would take it green. That is
not uncovered: the_persisted_standing_accessor_agrees_with_the_general_fold_
on_the_named_series deliberately keeps paying the full recompute so exactly ONE
claim in the corpus compares the served value against a freshly folded one. One
payer, one cross-check, every other consumer reading the shared value -- the
structure the share tier exists to create, and sound only while that witness
stays enrolled, which is why its enrolment is now load-bearing rather than
incidental.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RuWuQWB6MPkY7sNM4jEqAy

* Correct an overclaim: the tier guarantees the serve, and the witness guards authoring — two hazards this lane conflated

WITHDRAWN, NOT REWORDED. The previous commit's comment said
current_persisted_compile_phase_frontier_holds is safe to read the served
standing BECAUSE the_persisted_standing_accessor_agrees_with_the_general_fold_
on_the_named_series acts as the corpus's cache cross-check -- "one payer, one
cross-check, N readers, sound only while that witness stays enrolled".

That is wrong in a way worth naming rather than quietly editing, because it
would have made v2.workflow.floor_pure_producer_share's correctness depend on
ONE TEST IN ONE MODULE, and any later lane deleting or weakening that test
would have been silently deleting a guarantee it had no reason to know it held.

THE SERVE IS THE TIER'S OWN GUARANTEE AND ALWAYS WAS. That roster keys every
entry on FN-NODE IDENTITY plus a content hash of the full argument row and
serves only after the stored row verifies structurally equal in portable form,
a hash collision degrading to recompute rather than to a wrong value. For a
NULLARY producer the argument row is empty, so a wrong serve would require one
declaration's value returned under another declaration's node identity. That
invariant carries the tier's own evidence; re-asserting it in a witness would
be a second authority for it, which is the defect this whole branch exists to
remove.

TWO HAZARDS WERE CONFLATED. The witness's actual subject is AUTHORING --
whether the accessor names the series its name claims -- which is decidable
without the cache and says nothing about serving. Its removal is an ordinary
loss of a discriminating red, which is a real loss, and not a silent conversion
of a verified share into an unverified one.

Nothing executable changes. The overclaim was in prose, which is exactly the
kind of claim DESIGN 4b(1) calls rung inflation on the compiler's own
self-description, and it is retracted at the carrier that made it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RuWuQWB6MPkY7sNM4jEqAy

* The sharing half, measured by a second instrument on a second branch

This roster admits a row on measured recompute AND measured sharing -- a
non-sharing producer is cache population rather than a saving, which is the
ground formal_production_for_lhs_exact was REMOVED on. This row carried the
recompute and the cost drop and did NOT carry a direct observation of the
sharing, which is a gap in its own justification rather than a decoration.

gunbc#10068's lane closed their PR in favour of this one and handed over the
observation that fills it: the required floor's own `[floor-shared-fill]`
ledger, green run 33649705841, recording this producer at
`fill_ms=362 inclusive_ms=362 consumer_claims=22 consumer_modules=3
disposition=shared`, beside four `fill_ms=0 disposition=exclusive` fills for
the perturbed-series arguments that correctly populate nothing.

IT IS INDEPENDENT IN EVERY RESPECT THAT COULD HAVE MADE IT CIRCULAR: different
branch, different instrument, and a different KEY SHAPE -- that lane enrolled
the general declaration claim-forced on the reified receipt series, where this
row is nullary. Two instruments sharing no accounting and agreeing on 22
consumer claims across 3 modules is what rules out the reading that the sharing
is an artifact of one census's bookkeeping.

Recorded WITH ITS ORIGIN rather than absorbed into this lane's figures, because
a corroboration whose provenance is dropped is indistinguishable from a second
derivation by the same author -- which would make two observations read as one
and inflate exactly the confidence the corroboration was supposed to earn.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RuWuQWB6MPkY7sNM4jEqAy

* Hoist the flatten out of the loop: the same parameter-wider-than-its-body defect, second instance in one module

FLOOR-COST-500MS. The payer this branch declared as residue CROSSED on head
62989fd -- run 33683717936, failed=0, interrupted_cpu_deadline=1, the single
interrupted row being
the_persisted_standing_accessor_agrees_with_the_general_fold_on_the_named_series,
the row named in advance -- on a head differing from a green one by a COMMENT
BLOCK IN A .dag FILE. Predicted before it was observed, on a diff that cannot
have caused it. This commit removes the cost rather than rerolling the run.

NOT REROLLED, DELIBERATELY. gunbc.rung_drop
floor_cost_claim_qualification_unavailable admits a bounded reroll for exactly
this signature, and that admission is for lanes refused by SOMEONE ELSE'S
marginal row. This row is this branch's own, its crossing was forecast, and
rerolling it until it greens would be the absorbing arm under the name of the
lane funded to remove it.

TWO DEFECTS, ONE SHAPE, BOTH IN gunbc.self_host_compile_phase_frontier.

(1) THE SET DIFFERENCES RE-FLATTENED PER ELEMENT. board_added_identities and
board_removed_identities wrote `board_all_identities(board: other)` INSIDE the
filter predicate, so a flat_map over every row of one board ran once per
identity of the other -- O(|current| x |previous|) flattens where one flatten
per side suffices. The inner list is now bound once.

(2) THE PARAMETER WAS WIDER THAN THE BODY -- SECOND INSTANCE ON THIS BRANCH,
and the first one is what settled the #10068 collision. newly_exposed_identity_admitted
took `previous: RustcPhaseBoard` and `current: RustcPhaseBoard` and never
looked at a board: it used each for the flattened identity list and for
.furthest_phase_reached, nothing else. Its caller invokes it ONCE PER ADDED
IDENTITY through `added |> filter(...)`, and a second caller through
`all(ids, ...)`, so both boards were re-flattened per element, twice over,
while the flattened lists are constant across the whole call. Taking the
identity list and the phase lets each caller bind them once.

NARROWED RATHER THAN WIDENED, BOTH TIMES, AND THAT IS THE RESOLUTION RATHER
THAN A STYLE CHOICE: passing the board AND its flattened identities would put
two representations of one fact in one signature, which is the defect and not
the fix. The general form -- a parameter a body only funnels into a producer
puts the shared value one level INSIDE the parameter, where nothing can hoist
or share it -- is now twice-instanced in one module and is a class rather than
an anecdote.

DESIGN section 6's bare-minimum-cost rule fixes a proven cost shape regardless
of the realized n. This n is realized: these folds put their callers at 89-93%
of the floor's per-claim CPU ceiling.

MEASURED, NOT INFERRED, AND THE NEGATIVE IS ADMISSIBLE: the before is this
branch's own 448ms (run 33675891779) and 467ms (run 33683717936, where it
crossed); the after is the next run on this branch, joined at identity, with
the 72-row ambient control re-derived on the new pairing. If the row does not
move, the honest answer becomes a declared expected-over-budget row and this
lane reports that rather than attempting again.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RuWuQWB6MPkY7sNM4jEqAy

* Hoist the two body-interior annotations to module-item grain: §4c admits only a leading block

The parse phase refused #10108's head at self_host_compile_phase_frontier.dag:986-987 --
two `//` lines sitting inside apply_regression_transition's body. §4c models annotation
capture at module-item grain only; a body-interior comment has no subject to attach to, so
accepting it would either fabricate an attachment or discard the text, and the compiler
refuses instead of choosing. All three witness-lane phases failed from that one cause
(namespace-wave-admission had no head index, the floor had no subject), so the run carried
no cost opinion at all -- this red is not the class the PR exists to remove.

The rationale survives the move rather than being reworded into something vague: the
leading block now states what the declaration does at its own grain -- both boards are
fixed for the transition, so flattening per added identity would be authored duplication
(DESIGN §2), not a recurrence a cache could discharge.

Reported by bright-ram-778, who read the job log before assuming the cost class.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RuWuQWB6MPkY7sNM4jEqAy

* Split claim 3 out: the share row's admission rests on a denominator #10133 moved

The roster admits a row on MEASURED serve below MEASURED recompute. My present-vs-absent
pair (33669782300 / 33672356296) measured the recompute side against the pre-#10133
producer, and #10133 made exactly that recompute cheaper -- so the margin the receipt
establishes is an upper bound on the margin that exists now. The measurement is not wrong;
its subject moved.

The argument for landing it anyway was that the serve side is nullary with an empty
argument row and is therefore unlikely to have crossed. That inference is the one this
roster exists to refuse: gunbc#10094 withdrew three rows enrolled on exactly that reasoning
after all three were plausible and all three were wrong on measurement. A roster whose
discipline is measured-not-inferred cannot admit its next row on a structural argument
about which side sits near the floor. Ruling by bright-ram-778, holding me to this file's
own rule rather than a new one.

So floor_pure_producer_share.dag returns to main untouched, and the accessor's comment now
says plainly that being preparation-forceable is a property of the declaration and NOT a
claim that a row exists -- an unlanded citation is indistinguishable at the citing end. It
also records that the census figures beside it predate #10133.

What remains is claims 1 and 2, whose correctness never depended on a cost result.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RuWuQWB6MPkY7sNM4jEqAy

---------

Co-authored-by: Brian Searls <briansearls1@gmail.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
gunbai-bot Bot added a commit that referenced this pull request Sep 3, 2026
…au is real, the promotion mechanism is not, and #10133 was an outlier (#10143)

* Measure the floor cost distribution near the 500ms ceiling: the plateau is real, the promotion mechanism is not, and #10133 was an outlier rather than one step of a treadmill

MEASUREMENT, NOT A REPAIR, commissioned after gunbc#10133 to decide whether further
per-producer work is a treadmill. Nothing here ran a floor: every figure is a reading
of `required-floor-claim-cost`, `required-floor-cross-claim-demand` and the
`required-floor:` summary already published by ten main runs.

THREE FINDINGS.

The band IS dense. 87 rows sit above 300ms and 32 above 400ms, and repairing modules
top-down returns about 10ms per module below the top of the band -- the treadmill,
quantified.

The new crossers do NOT share a root the way #10133's three did. That trio reported
612,662-658,336 eval steps, a 1.07x spread that IS the signature of one shared dominant
producer. The families now near the ceiling span 63k-197k, a 3.1x spread, in three
distinct shapes. One shared-producer lead exists (`bind_outcome`) and is explicitly not
a plan: the artifact ranking it declares its own columns inclusive-of-callees and
non-summable, and they total 14x the run's actual claim CPU.

The margin is measurable and uniform. Same-identity run-to-run inflation is median
1.16x, p90 1.35x, max ~2.0x, and it does NOT vary with cost bucket or with host-call
heaviness -- both hypotheses tested and refuted, which is what makes one global margin
defensible. PR runs reach ~1.6x, implying a clean-run budget near 312ms.

AND ONE CORRECTION TO THE FRAMING THIS WAS COMMISSIONED UNDER. Promotion is refuted:
rank does not cause a crossing and removing rows above a row does not make it slower.
The five rows that crossed post-repair went ABOVE their own pre-repair maxima (median
343ms, max 472ms over 45 observations, against 504-533ms), on the most contended run in
the sample. Replaying all ten runs with the repaired modules excluded takes the refusal
rate from 5/10 to 1/10, and across the nine pre-repair runs no row outside those three
modules ever crossed. They were 17 of the 22 distinct identities that crossed anywhere:
an outlier, not a sample of the plateau.

Raising the ceiling is named as not-an-option, because `v2.workflow.required_floor`
already litigated it twice and repudiated both raises.

An eleventh run's artifact downloaded empty and is excluded; recorded because it
silently emptied a ten-way intersection before it was caught.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WEFMasGn2cAGCMJW2694pb

* Give the floor cost analysis an entry point: the derivations become a witnessed authority and the memo cites a producer instead of carrying figures (review 58983)

REVIEW 58983 REFUSED #10143 AND WAS RIGHT. The memo named
`gunbc.witness_floor_workflow` as its instrument. That workflow produces the RAW
ROWS and performs none of the analysis: the ten-run intersection, the band
histogram, the counterfactual replay and the inflation percentiles existed only
as prose plus numbers, so a reader could check the arithmetic of nothing. DESIGN
section 6 is explicit -- name the instrument, never transcribe its output, and if
a measurement is worth re-deriving it is worth an entry point.

WHAT LANDS. `gunbc.floor_cost_distribution` carries the derivations as pure
functions over rows supplied by the caller; `tools.floor_cost_distribution_instrument`
reads the artifacts and renders the report. That is the section 3 split between
what the analysis IS and where the bytes come from, and it is what lets
`test.claim.floor_cost_distribution_witness` exercise every derivation on a
hand-built fixture with no run artifacts present -- 14 rows, each with a
discriminating red beside it.

VERIFIED BY REPRODUCTION, not by inspection. Run against the ten artifacts the
memo names, the instrument returns every figure the memo carries: the band
histogram (2231/502/441/148/89/55/27/5), 87 rows over 300ms and 32 over 400ms,
5 runs refused as observed against 1 with the repaired modules excluded, worst
row 533ms, the full ten-step ladder (525, 511, 490, 487, 486, 466, 458, 450, 434,
429), and inflation percentiles of 1147/1345/1433/1643 permille implying
435/371/348/304ms clean-run budgets. Those figures were originally derived by an
independent ad-hoc script; two implementations agreeing is the cross-check.

THE REFUSAL ARMS ARE PRODUCED, NOT DECLARED. An absent or empty artifact refuses
and is named rather than contributing an empty population -- exactly the failure
that silently emptied a ten-way identity intersection during the original
analysis, where a run reading as zero rows is indistinguishable from a run in
which nothing crossed. Both arms are witnessed, with a discriminating red proving
a complete sample produces no refusal line, and the actuator's refusal is
exercised by execution in both directions.

A NOTE ON WHY THIS IS AN INSTRUMENT AND NOT A SCAFFOLD, since the distinction
decided the shape: a scaffold is what the terminal architecture throws away, an
instrument is what it keeps and re-runs. This analysis will be re-derived after
the next floor change, so it earns a modeled home under dag/gunbc/instruments/
rather than a committed shell or python script.

The memo keeps its figures as illustrative of what the instrument returned on the
runs it names, and says plainly that the instrument is the authority where the
two disagree.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WEFMasGn2cAGCMJW2694pb

* Route the floor cost analysis's whole time axis through std.measure Millisecond (review 59000)

REVIEW 59000 IS RIGHT AND THE SCOPE IT NAMES IS THE RIGHT SCOPE. The finding is
not six leaky fields, it is that the module carried a domain unit vocabulary
end-to-end -- rows, bands, ladder, threshold, budget -- as bare `Int`
milliseconds with no `std.measure` import at all. That is a parallel
representation of a quantity this repository already models, which is DESIGN
section 2's "net concepts must not grow by re-invention" and section 3's single
authority. The corpus had already ruled on exactly this shape: gunbc.witness_row_cost's
own seed note records constructing Millisecond through the interpreter
specifically "so no flat-scalar duration crosses the seam".

WHAT CHANGED. `FloorCostRow.cpu_ms`, `CostBand.lower_ms`/`upper_ms`,
`LadderStep.worst_ms`, every ceiling and threshold parameter, the band edges, and
the implied budget are `Millisecond`. The parser wraps at the boundary, where the
artifact's integer column becomes a carried quantity. Comparison goes through
`millisecond_count` rather than over the carrier, because the carrier is what
stops a figure read from one clock being compared against a figure read from
another and the count is the projection a comparison is entitled to.

WHAT DELIBERATELY STAYS `Int`, so this is not a blanket sweep: row counts and
refused-run counts are cardinalities, and the inflation permille and its
percentile argument are a DIMENSIONLESS RATIO. Giving a ratio a time unit would
be the same error in the other direction.

`implied_clean_run_budget_ms` is renamed `implied_clean_run_budget` -- the `_ms`
suffix was the field name doing the unit's job, which is the thing the carrier
replaces.

VERIFIED BY REPRODUCTION, WHICH IS THE POINT OF DOING IT THIS WAY. A unit
refactor is exactly where an off-by-conversion hides, so the control is that the
instrument's output over the ten named artifacts is BYTE-IDENTICAL to the run
before this commit -- same bands, same 87/32, same 5 and 1, same ladder, same
1147/1345/1433/1643 permille and same 435/371/348/304ms budgets. All 14 witness
rows still pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WEFMasGn2cAGCMJW2694pb

---------

Co-authored-by: gunbc-ci-auto-heal <gunbc-ci-auto-heal@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
gunbai-bot Bot pushed a commit that referenced this pull request Sep 3, 2026
Forward-delta adjudication of this PR's tested base against current main found
one MATERIAL commit that no path-overlap check would have surfaced: gunbc#10143
measures the near-ceiling population this document's screening section ranks, and
landed after the base.

Two results of that measurement bound what any screening here can buy:

1. THE BAND IS A PLATEAU, no gap below the ceiling. Repairing top-down, the first
   three modules buy 43ms and the next seven buy 61ms -- ~10ms per module, about
   one fiftieth of the ceiling. So surfacing the next-most-expensive producer has
   a small payoff BY CONSTRUCTION, and this document's own ordering (safety-relevant
   rows ahead of milliseconds) is the better one for reasons now measured.

2. THE CROSSERS DO NOT GENERALLY SHARE A ROOT. #10143 records that the modules
   newly near the ceiling do not share a root the way #10133's three did, so the
   shared-derivation bucket described here is the exception, not the common case.

Also records the one worked counterexample, gunbc#10170: four rust_body_add_emit
rows reaching five canonical fixtures through a ~5k-line module, cut from
506/425/416/362 to 180/174/172/168. The SPREAD collapsing 144ms -> 12ms is what
distinguishes a genuine shared root from four rows each getting cheaper -- a
constant subtracted from four independent costs leaves the spread.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019LhF5WCbZqrZHPqsnjpkYu
briansrls pushed a commit that referenced this pull request Sep 3, 2026
…xt reader of "the floor is a coin flip" will hit it (#10196)

* Measure the floor cost distribution near the 500ms ceiling: the plateau is real, the promotion mechanism is not, and #10133 was an outlier rather than one step of a treadmill

MEASUREMENT, NOT A REPAIR, commissioned after gunbc#10133 to decide whether further
per-producer work is a treadmill. Nothing here ran a floor: every figure is a reading
of `required-floor-claim-cost`, `required-floor-cross-claim-demand` and the
`required-floor:` summary already published by ten main runs.

THREE FINDINGS.

The band IS dense. 87 rows sit above 300ms and 32 above 400ms, and repairing modules
top-down returns about 10ms per module below the top of the band -- the treadmill,
quantified.

The new crossers do NOT share a root the way #10133's three did. That trio reported
612,662-658,336 eval steps, a 1.07x spread that IS the signature of one shared dominant
producer. The families now near the ceiling span 63k-197k, a 3.1x spread, in three
distinct shapes. One shared-producer lead exists (`bind_outcome`) and is explicitly not
a plan: the artifact ranking it declares its own columns inclusive-of-callees and
non-summable, and they total 14x the run's actual claim CPU.

The margin is measurable and uniform. Same-identity run-to-run inflation is median
1.16x, p90 1.35x, max ~2.0x, and it does NOT vary with cost bucket or with host-call
heaviness -- both hypotheses tested and refuted, which is what makes one global margin
defensible. PR runs reach ~1.6x, implying a clean-run budget near 312ms.

AND ONE CORRECTION TO THE FRAMING THIS WAS COMMISSIONED UNDER. Promotion is refuted:
rank does not cause a crossing and removing rows above a row does not make it slower.
The five rows that crossed post-repair went ABOVE their own pre-repair maxima (median
343ms, max 472ms over 45 observations, against 504-533ms), on the most contended run in
the sample. Replaying all ten runs with the repaired modules excluded takes the refusal
rate from 5/10 to 1/10, and across the nine pre-repair runs no row outside those three
modules ever crossed. They were 17 of the 22 distinct identities that crossed anywhere:
an outlier, not a sample of the plateau.

Raising the ceiling is named as not-an-option, because `v2.workflow.required_floor`
already litigated it twice and repudiated both raises.

An eleventh run's artifact downloaded empty and is excluded; recorded because it
silently emptied a ten-way intersection before it was caught.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WEFMasGn2cAGCMJW2694pb

* Give the floor cost analysis an entry point: the derivations become a witnessed authority and the memo cites a producer instead of carrying figures (review 58983)

REVIEW 58983 REFUSED #10143 AND WAS RIGHT. The memo named
`gunbc.witness_floor_workflow` as its instrument. That workflow produces the RAW
ROWS and performs none of the analysis: the ten-run intersection, the band
histogram, the counterfactual replay and the inflation percentiles existed only
as prose plus numbers, so a reader could check the arithmetic of nothing. DESIGN
section 6 is explicit -- name the instrument, never transcribe its output, and if
a measurement is worth re-deriving it is worth an entry point.

WHAT LANDS. `gunbc.floor_cost_distribution` carries the derivations as pure
functions over rows supplied by the caller; `tools.floor_cost_distribution_instrument`
reads the artifacts and renders the report. That is the section 3 split between
what the analysis IS and where the bytes come from, and it is what lets
`test.claim.floor_cost_distribution_witness` exercise every derivation on a
hand-built fixture with no run artifacts present -- 14 rows, each with a
discriminating red beside it.

VERIFIED BY REPRODUCTION, not by inspection. Run against the ten artifacts the
memo names, the instrument returns every figure the memo carries: the band
histogram (2231/502/441/148/89/55/27/5), 87 rows over 300ms and 32 over 400ms,
5 runs refused as observed against 1 with the repaired modules excluded, worst
row 533ms, the full ten-step ladder (525, 511, 490, 487, 486, 466, 458, 450, 434,
429), and inflation percentiles of 1147/1345/1433/1643 permille implying
435/371/348/304ms clean-run budgets. Those figures were originally derived by an
independent ad-hoc script; two implementations agreeing is the cross-check.

THE REFUSAL ARMS ARE PRODUCED, NOT DECLARED. An absent or empty artifact refuses
and is named rather than contributing an empty population -- exactly the failure
that silently emptied a ten-way identity intersection during the original
analysis, where a run reading as zero rows is indistinguishable from a run in
which nothing crossed. Both arms are witnessed, with a discriminating red proving
a complete sample produces no refusal line, and the actuator's refusal is
exercised by execution in both directions.

A NOTE ON WHY THIS IS AN INSTRUMENT AND NOT A SCAFFOLD, since the distinction
decided the shape: a scaffold is what the terminal architecture throws away, an
instrument is what it keeps and re-runs. This analysis will be re-derived after
the next floor change, so it earns a modeled home under dag/gunbc/instruments/
rather than a committed shell or python script.

The memo keeps its figures as illustrative of what the instrument returned on the
runs it names, and says plainly that the instrument is the authority where the
two disagree.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WEFMasGn2cAGCMJW2694pb

* Route the floor cost analysis's whole time axis through std.measure Millisecond (review 59000)

REVIEW 59000 IS RIGHT AND THE SCOPE IT NAMES IS THE RIGHT SCOPE. The finding is
not six leaky fields, it is that the module carried a domain unit vocabulary
end-to-end -- rows, bands, ladder, threshold, budget -- as bare `Int`
milliseconds with no `std.measure` import at all. That is a parallel
representation of a quantity this repository already models, which is DESIGN
section 2's "net concepts must not grow by re-invention" and section 3's single
authority. The corpus had already ruled on exactly this shape: gunbc.witness_row_cost's
own seed note records constructing Millisecond through the interpreter
specifically "so no flat-scalar duration crosses the seam".

WHAT CHANGED. `FloorCostRow.cpu_ms`, `CostBand.lower_ms`/`upper_ms`,
`LadderStep.worst_ms`, every ceiling and threshold parameter, the band edges, and
the implied budget are `Millisecond`. The parser wraps at the boundary, where the
artifact's integer column becomes a carried quantity. Comparison goes through
`millisecond_count` rather than over the carrier, because the carrier is what
stops a figure read from one clock being compared against a figure read from
another and the count is the projection a comparison is entitled to.

WHAT DELIBERATELY STAYS `Int`, so this is not a blanket sweep: row counts and
refused-run counts are cardinalities, and the inflation permille and its
percentile argument are a DIMENSIONLESS RATIO. Giving a ratio a time unit would
be the same error in the other direction.

`implied_clean_run_budget_ms` is renamed `implied_clean_run_budget` -- the `_ms`
suffix was the field name doing the unit's job, which is the thing the carrier
replaces.

VERIFIED BY REPRODUCTION, WHICH IS THE POINT OF DOING IT THIS WAY. A unit
refactor is exactly where an off-by-conversion hides, so the control is that the
instrument's output over the ten named artifacts is BYTE-IDENTICAL to the run
before this commit -- same bands, same 87/32, same 5 and 1, same ladder, same
1147/1345/1433/1643 permille and same 435/371/348/304ms budgets. All 14 witness
rows still pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WEFMasGn2cAGCMJW2694pb

* Record that the floor-crosser premise is measured false, where the next reader of "the floor is a coin flip" will hit it

THE BELIEF THIS RETIRES IS THE ONE THAT COSTS A NIGHT. The work item behind
gunbc#10133's follow-up was "three (later seven) self-host witness modules sit
ABOVE the 500ms hard CPU ceiling". Measured, no row in that family sits above the
ceiling, both remedies section 5 names are refused for it on evidence, and the
ceiling is correctly denominated. That result lived only in a session thread, and
session threads archive -- so it is written where section 5's open question is,
because the reader who re-forms this belief will arrive from the same symptom: a
red floor lane naming a handful of emit rows.

THREE MISREADINGS, EACH WITH THE CONTROL THAT EXPOSES IT, because the numbers are
less reusable than the way they mislead. An `interrupted_before_verdict` figure is
the BUDGET, not the row -- every interrupted row reports approximately whatever
ceiling stopped it, so a set of them looks like a tight cluster "just over"
whatever the true costs are. The crossing set is NOT the expensive set -- the row
that passed on CI is the most expensive of its four siblings, so repairing "the
rows that crossed" fits noise, and the non-crossing siblings are what reveal the
band. And a one-anchor local-to-CI factor does not transfer: one taken from a
single identity predicted a row comfortably under the ceiling that CI had in fact
interrupted.

NAMED, NOT TRANSCRIBED (section 6). The join has no instrument module -- it needs
a second axis the section 5 instrument does not read -- so the recipe is given at
identity grain and the absence of an owning authority is stated as a gap rather
than papered over. Ratios are described by their shape and stability across named
run classes; re-derive them rather than quote this page.

THE HEDGE IS PRESERVED EXACTLY. The projection term is recorded as measured at the
`target_project_arrow_body_to_value_expression` FRAME and attributed BY
ELIMINATION to the primitive-apply step beneath it, with the neighbouring
candidates named as measuring zero and the two table-shaped ones noted as tested
with inputs that would have exposed a build even on a miss. The live lead is left
UNDETERMINED between a memo keyed above the body and a one-time warm, because
those want different providers and the mechanism must be separated before any
provider is proposed. Both prior sharing attempts are recorded with their
refutations so neither is re-proposed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WEFMasGn2cAGCMJW2694pb

* Retire the stable-subject verb: no fixed above-ceiling identity set, not "no row sits above the ceiling" (review 5098047617)

THE REVIEW IS RIGHT AND THE VERB WAS THE WHOLE DEFECT. The section opened with
"Measured, no row in that family sits above the ceiling", and its own later
evidence refutes that sentence: failed-run inflation carries the top of the band
across 500ms, and four completed-over-cost rows were observed at 506-515ms.
"Sits" asserts a STABLE ROW PROPERTY while the section's central finding is that
crossing is ATTEMPT-DEPENDENT -- so the opening claimed the one shape the body
spends four subsections dismantling.

WHAT THE SECTION NOW SAYS is the proposition actually established: no fixed
above-ceiling identity set exists; completed green runs place the family's band
below 500ms, while run variance can carry its top across. What is retired is that
the family is INTRINSICALLY above the ceiling -- a row does not stably sit on
either side of an attempt-safety boundary -- and explicitly NOT the observed
crossings, which are real and recur. The 506-515ms observations are named in the
retraction itself so the correction cannot be read as denying them.

The heading changes for the same reason: "the crossers are not above the ceiling"
was the same stable-subject shape one line up.

NOTHING ELSE MOVES. The two other scoped statements were already bounded
correctly -- the restated subject says "on a normal run" and "a property of the
run, not of the rows", and the emit_host_fold example is scoped to "every green
run in the join". The re-derivation recipe, the three misreading controls, the
eval_steps axis and the producer-share caveat all stand.

WHY THIS CLASS IS WORTH THE COMMIT MESSAGE: when a subject's true form is a
DISTRIBUTION, ordinary language keeps snapping it back to a PROPERTY, and it does
so in both directions -- the same data produced "the rows are genuinely under
budget" elsewhere within the hour. Neither sentence survives its own evidence.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WEFMasGn2cAGCMJW2694pb

---------

Co-authored-by: gunbc-ci-auto-heal <gunbc-ci-auto-heal@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
briansrls pushed a commit that referenced this pull request Sep 3, 2026
… discriminator, and the signature screen (#10114)

* Share the compile-phase frontier fold instead of re-folding it once per witness

The floor's own cross-claim demand census names one derivation that 23 required
claims each recompute and none reuse: gunbc.self_host_compile_phase_frontier
phase_board_series_ratchet, at 23 claims against 23 evals across three modules
in required_floor_cross_claim_demand.tsv (main run 33668368846). The two modules
it dominates are that run's two largest populations over the 100ms cost line in
required_floor_claim_cost.tsv, and inside each of them every over-line row lands
within a few percent of ~615k eval steps -- the shape of one fold repeated, not
of N witnesses doing their own work.

So the remedy is the roster, not the witnesses: splitting them would only
re-attribute the fold to whichever fragment ran first, which is
witness_row_cost's standing
witness_decomposition_does_not_reduce_entry_cost_note.

Enrolled claim-forced at the root rather than at compile_phase_frontier_standing
-- the census shows the standing's cost is inclusive of this fold, so one row
discharges both and reaches one more claim. The value is PhaseBoardSeriesVerdict,
a nullary variant on the held arm, so the serve walks nothing; that is the
admission criterion the two rust target models failed, and the reason they stay
out.

Present-vs-absent is re-derived from required_floor_claim_cost.tsv and
required_floor_cross_claim_demand.tsv on this branch against run 33668368846 --
the instruments named in the roster's own admission criterion.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CDKZ5pDN7cXTSEvYST2y78

* The lane's ownership question has an instrument: point the census doc at the cross-claim demand ledger

Everything in this document infers ownership from timings and concludes a
timings-only census cannot see whose work a cost is. The floor now emits
required_floor_cross_claim_demand.tsv on every required run -- claims against
evals per producer identity, which answers that question directly -- and a
reader landing on this document was not being told it exists.

Also records the two findings from re-deriving the census on run 33668368846:
the log's over-cost head is capped at 25 rows against 295 in the artifact
(instrument_output_read_as_subject_content, the second occurrence on this lane),
and eval_steps rather than milliseconds is the within-module discriminator,
because steps are deterministic and so a tight step cluster cannot be a quiet
runner -- the confound that demoted the variance screen and the cluster prior.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CDKZ5pDN7cXTSEvYST2y78

* Supply the controlled present-vs-absent receipt the admission criterion requires

review 58870 was right: the row was enrolled with the recompute half measured
and the serve half only promised, which is the test the two rust target models
are excluded on. The recipe also named a run at a different head, which cannot
answer the serve question at all.

The pair is now controlled -- main run 33671815204 at 20caf4e, my branch merged
onto that same head as run 33673657876, both executed=3498, differing by this
one line -- and the whole-run total is the conjunct that answers it, because the
serve lands on the consuming claims and a per-row drop alone cannot see it:
claim_cpu_total_ms 110998 -> 93110, the ratchet 23 evals -> 1, the three modules
11833ms -> 4260ms with all 21 over-line rows now under it, and failed /
unexpected_failures 0/0 on both arms.

The result that matters more than the milliseconds: main reports cpu_deadline=2
and both preempted rows are this module's discriminating REDs, cut off by the
500ms deadline before reaching a verdict. This branch reports cpu_deadline=0 and
both answer. The fold was not merely expensive, it was suppressing the two
controls the witness exists for.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CDKZ5pDN7cXTSEvYST2y78

* Screen census candidates by signature, not by cost: the ratio rule and three by-inspection exclusions

The demand census ranks by cost and cost is the wrong sort key. Three large
classes on it cannot be shared at all, and the fn-typed-parameter class refuses
at PUBLICATION rather than losing a measurement -- so a controlled pair spent on
one produces noise, not a negative result. I proposed such a candidate in this
lane and withdrew it on reading the signature; the screening test is what stops
the next person paying two CI runs for that.

The rule is a ratio, not a size heuristic: serve scales with argument plus value
(hash, structural verify, container rebuild), recompute scales with the
derivation between them. phase_board_series_ratchet carries a whole receipt
series as its key and still admits, because 585,854 eval steps amortise the
verify.

The fn-in-key exclusion is grounded in the seed rather than inferred from the
roster's prose: reification refuses Value::Fn as OriginBoundNode, and arguments
reify on the same path via portable_args_from_ctx -> RefusedArgsNotPortable.

Also records that the deadline-preempted rows on that pair's absent arm were the
witness's discriminating REDs, so this population is ordered by which refusals
are not executing, not only by milliseconds.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CDKZ5pDN7cXTSEvYST2y78

* Record the circularity: the repair for the deadline class travels through the gate the class refuses

This is the session's most consequential structural finding and it existed only
in a message thread, which dies. A preemption is runner-dependent, so any PR can
draw a refusal from two rows that are not broken -- including the PR carrying the
repair. Observed in situ rather than argued: a 138-line documentation-only change
with no .dag, no Rust and no workflow was refused by
a_live_tree_that_gained_an_identity_refuses_and_names_it at 501ms/500ms while the
consolidation that would make that fold cheap sat open with its own floor lane
running.

Also states plainly why the obvious exit is closed. A re-run does not make a row
cheaper, it redraws the runner, and the green it buys is a green over refusals
that did not execute -- the interrupted rows ARE the discriminating REDs, so the
mark asserts a verdict for precisely the claims that were preempted. cpu_deadline
is a safety counter, not a performance one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CDKZ5pDN7cXTSEvYST2y78

* Three method claims that outran their instruments (review 5096548942)

All three are prose defects in the appendix, not CI failures. The exact-head run
is terminal success and remains a DRAW, not a repair.

1. SAFETY RANKING WAS CLAIMED, AND IS NOT AVAILABLE. The text said to rank the
   population by which refusals are not executing, and that the interruption
   diagnostics say whether a discriminating RED is the one cut off. They do not.
   They name the victim; they do not author what the victim exists to prove.
   std.witness_purpose is explicit that purpose is AUTHORED, NOT INFERRED, and it
   currently has zero declarers and zero consumers. The two live-gate rows are
   refusal probes because their source was read one at a time -- that does not
   generalise to a corpus-wide procedure by reading names off a diagnostic.
   Now stated as three separate claims with only the first two discharged:
   victim identity is observable; these two were independently verified; ranking
   by safety relevance REMAINS BLOCKED until authored purpose can be joined
   against verdict_reached.

2. THE SIGNATURE SCREEN CONTRADICTED ITS OWN RULE. It said the rule "screens OUT
   -- it never admits" and then called one shape "the shape that admits" and
   another "admissible". In this repository admit is a loaded verb: only the
   controlled present-versus-absent receipt gets to say it. Both now say the
   shape SURVIVES THE EXCLUSION SCREEN and is worth measuring.

3. eval_steps WAS ASKED AN OWNERSHIP QUESTION. It measures WORK. Near-identical
   step counts across rows do not identify a shared producer -- the same total can
   arise from unrelated derivations of similar size. Ownership is established by
   the cross-claim demand artifact, which is keyed by PRODUCER IDENTITY. The
   shared-derivation conclusion is now a JOIN of the two instruments: demand names
   the shared producer, eval_steps characterises the work reaching it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019LhF5WCbZqrZHPqsnjpkYu

* File the third truncation: module_sample caps at 8 while the row reports modules=71

The two print caps this document already records have trailers. This one does
not: required_floor_cross_claim_demand.tsv's module_sample column caps at 8
modules and 1044 of 26317 rows exceed it, so a row lists eight consumers beside a
modules=71 and a reader joining producers to modules gets a silently partial
answer.

It bites exactly where the join matters -- the truncated rows are the widely
shared producers, which are the ones most likely to explain why a whole module's
rows cluster -- so a corpus-scale module-to-producer join is not performable from
this artifact today, and the per-module reading here is sound only where the join
was done by hand. That qualification is owed: a 73%-of-over-line-CPU
classification was reported off the step ratio with the demand join performed for
only the two modules acted on, which is a step-cluster inference at that strength
rather than an ownership result.

Named as an obligation, deliberately not undertaken here so it can be sized
against other work rather than absorbed into this lane.

Three truncations on one lane is a pattern, and the reusable shape is stated: an
instrument can report a cap honestly at the top level while the column a consumer
reads silently answers for fewer.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CDKZ5pDN7cXTSEvYST2y78

* Bound the screening by what gunbc#10143 measured

Forward-delta adjudication of this PR's tested base against current main found
one MATERIAL commit that no path-overlap check would have surfaced: gunbc#10143
measures the near-ceiling population this document's screening section ranks, and
landed after the base.

Two results of that measurement bound what any screening here can buy:

1. THE BAND IS A PLATEAU, no gap below the ceiling. Repairing top-down, the first
   three modules buy 43ms and the next seven buy 61ms -- ~10ms per module, about
   one fiftieth of the ceiling. So surfacing the next-most-expensive producer has
   a small payoff BY CONSTRUCTION, and this document's own ordering (safety-relevant
   rows ahead of milliseconds) is the better one for reasons now measured.

2. THE CROSSERS DO NOT GENERALLY SHARE A ROOT. #10143 records that the modules
   newly near the ceiling do not share a root the way #10133's three did, so the
   shared-derivation bucket described here is the exception, not the common case.

Also records the one worked counterexample, gunbc#10170: four rust_body_add_emit
rows reaching five canonical fixtures through a ~5k-line module, cut from
506/425/416/362 to 180/174/172/168. The SPREAD collapsing 144ms -> 12ms is what
distinguishes a genuine shared root from four rows each getting cheaper -- a
constant subtracted from four independent costs leaves the spread.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019LhF5WCbZqrZHPqsnjpkYu

* Restore the union: my eae0cca reverted the b561d7c corrections, and the re-fix dropped the truncation

MY FAULT FIRST. eae0cca was written by a script that read the file from a STALE
WORKING TREE, so it silently reverted all three of b561d7c's corrections -- the
work-is-not-ownership paragraph, the survives-the-screen wording, and the
safety-ranking-remains-blocked statement. I then reported "I am done pushing"
while having undone someone else's fix without noticing. Nothing in my diff
review caught it because I diffed my own change, not the file's resulting state.

9bdd91c restored those corrections from its own copy and, by the same mechanism
in the opposite direction, dropped the module_sample truncation section eae0cca
had added. So no commit on this branch has ever carried both.

This one does, verified by marker rather than by reading the diff: the three
corrections, the #10143 bounds, and the truncation finding are all present at
once. The check that would have caught either revert is grepping the RESULT for
every claim the file is supposed to carry -- a diff shows what you changed, never
what you clobbered.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CDKZ5pDN7cXTSEvYST2y78

* Retract the #10170 counterexample: this document's own discriminator refutes it

I added gunbc#10170 as the one worked example of a genuine shared root, on
506/425/416/362 -> 180/174/172/168 with the spread collapsing 144ms -> 12ms. Two
floor artifacts refute it, and the instrument that refutes it is the one this
section argues for.

  pre  (run 33707763185)  cpu 314 305 290 318   steps 171091 166669 162507 171124
  post (run 33711894017)  cpu 339 327 317 344   steps 171052 166630 162468 171085

eval_steps moved 0.02% -- 39 steps out of 171091. The work did not change, so the
relocation removed no evaluation work from these claims. CPU is HIGHER after, which
is within noise, so the honest reading is no measured effect in either direction.

The 506 baseline was a contended-run outlier: the rows were already at
314/305/290/318 with a spread of 28 before anything was touched, so both the drop
and the spread collapse were artifacts of the baseline I compared against.

And 180/174/172/168 came from a targeted claim_batch -- a fraction of the corpus and
a different execution envelope. I asked the author to report from the artifact
rather than the log, then accepted a number that came from neither, and committed
it here.

STEPS ARE DETERMINISTIC WHERE MILLISECONDS ARE NOT. A cost story that milliseconds
support and steps refute is the clock talking. The same discriminator that says work
is not ownership says here that a millisecond drop is not a work reduction.

Records the measured lead in its place, found by the instrument that establishes
ownership rather than by step clustering: the cross-claim demand census keyed by
PRODUCER IDENTITY names target_project_arrow_body_to_value_expression at 95 claims /
117 evals, 4396ms total with 4350ms cross-claim across the emit family including all
four rows. Admission still requires the controlled present-versus-absent pair, which
has not been run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019LhF5WCbZqrZHPqsnjpkYu

* The lead my retraction named is disqualified thirty lines below it

Caught by lively-stag-270, against my own paragraph. My retraction of the #10170
counterexample named target_project_arrow_body_to_value_expression as "the measured
lead" -- and this same document's fn-typed-parameter exclusion names that exact
function as excluded, carrying handle_transform: fn(Node, Node, TargetModel) -> ...

Reification refuses the argument (portable_args_from_ctx ->
RefusedArgsNotPortable), so the key cannot be formed and the store refuses AT
PUBLICATION. It is not a candidate whose serve might lose; it is a row that never
stores. The controlled present-versus-absent pair my paragraph called for would
measure noise against nothing, because the absent arm is the only arm.

SAME SHAPE AS THE CONTRADICTION THE REVIEWER FOUND IN THIS FILE EARLIER: a screen
that says it only excludes, followed by a sentence reading as an admission. Here it
is a section that excludes a function, preceded by a paragraph offering it as the
lead. A reader taking that as a worklist spends two CI runs on a row that cannot
store -- the exact waste the screening section exists to prevent, and it would have
been the sixth candidate withdrawn on this class.

Keeps the measurement and the point about which instrument found it: the demand
artifact keyed by producer identity reports 95 claims / 117 evals, 4396ms total,
4350ms cross-claim. Changes only the disposition, from a lead to a recorded fact
about the serve mechanism's COVERAGE -- the largest cross-claim producer in this
family is unreachable by the mechanism, which is not a gap in the census.

Verified by asserting on the RESULT: all seven required markers present after the
edit, per the discipline in stale_buffer_write_reverts_outside_its_own_diff.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019LhF5WCbZqrZHPqsnjpkYu

---------

Co-authored-by: gunbc-ci-auto-heal <gunbc-ci-auto-heal@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant