Repository navigation
Floor cost repair: one producer, not three witnesses — the compile-phase frontier standing crosses the 500ms CPU ceiling and makes the required floor a coin flip - #10133
Conversation
…s to ~60ms so three required-floor witness modules stop crossing the 500ms CPU ceiling
THE COIN FLIP, LOCATED. `required-floor` refused three of the last four main runs
with `interrupted_cpu_deadline` in {17, 15, 2} and `verdict=FloorRefused`, while
run 33673806092 on the same corpus passed. The refused rows are the same
seventeen identities in three modules, and the run log names them
`BUDGET-REFUSED in 503..506ms` against
`v2.workflow.required_floor` `required_floor_claim_cpu_safety_limit_ms` (500):
test/claim/self_host_compile_phase_live_gate_witness 13 rows
test/claim/self_host_compile_phase_frontier_witness 2 rows
test/claim/compiler_frontend_program_status_witness 2 rows
On the passing run those same rows completed at 452-496ms. Nothing about them
changed between the runs; they sit inside one scheduler quantum of the ceiling,
so which side of it a PR lands on is ambient fleet load. That is the coin flip.
ONE ROOT, NOT THREE. All seventeen rows spend their whole cost in
`gunbc.self_host_compile_phase_frontier` `compile_phase_frontier_standing` over
the persisted receipt series. The first two modules call it directly; the third
reaches it through `gunbc.compiler_frontend_program_status` `milestone_status`,
whose `SelfHostWholeCorpusPopulationDerived` arm asks the same question. So this
is one repair, and it is in the producer rather than in any witness.
MEASURED WITH THE INSTRUMENT THAT ALREADY EXISTS -- `claim_batch --source-root dag
--source-root src/v2 --entry <witness> --functions <rows>`, whose `[witness]`
line reports per-row CPU. Attribution was done by probe modules that call one
callee at a time; the interpreter memoizes a pure call within one claim and not
across claims, so each probe pays its own cost and the deltas are marginal
figures. Three terms carried it, all of them shapes DESIGN section 6 fixes on
sight rather than on the realized n:
`newly_exposed_identity_admitted` -- a cheap identity guard conjoined ahead of
an expensive cause derivation. `&&` in this substrate evaluates BOTH operands
(`v1_interpreter` `eval_expr_inner` evaluates left and right before applying
the operator), so the guard narrowed the result without guarding the work: the
derivation, two of whose arms scan the whole previous board, ran once per
(added identity x disposition row) pair. Selecting the rows first is exactly
equivalent -- for boolean p and q, `any(xs, p && q)` = `any(filter(xs, p), q)`.
`board_added_identities` / `board_removed_identities` / `phase_ratchet` -- set
difference answered by `any` over the other population per row, so the
comparison count is the PRODUCT of two populations that grow with every
receipt. Now one identity-keyed membership index per population.
`identity_is_hop_relocation` -- re-derived `diagnostic_semantic_key` for every
row of `added` and `removed` on every call. The keys depend on the transition,
not on the identity under test, so they are now derived once at
`apply_regression_transition`, which is the least common ancestor of the demand.
NO SEMANTIC CHANGE IS CLAIMED AND NONE WAS FOUND. Every row of all three witness
modules passes, including the frontier witness's discriminating reds over
`newly_exposed_identity_admitted` and the live gate's planted-identity and
equal-cardinality-swap refusals. The receipts' sealed identities, prefix digests
and census digests are all derived, not literals, so an equal standing is an
equal projection: `gunbc.design_ledgers` renders only `latest`, and its output is
unchanged. No consumer outside this module named any signature that moved, and
the module has no stage0 Rust mirror, so no regeneration is implied.
WHAT THIS DOES NOT DO. It does not lower the population toward the ~1ms median an
ordinary witness should hold; the rows land near 60ms locally and remain above
`required_floor_claim_cost_line_ms` (100) once the fleet's host factor is applied.
The residue is the receipt-sealing data fold, which is paid once per claim and is
a different subject from the per-call scans repaired here.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WEFMasGn2cAGCMJW2694pb
Verified on the floor, at identity grain rather than by the greenRun 33688872492, The accounting closes exactly, which is stronger than the pass:
Same All seventeen rows also left the over-cost report. That report lists every row above One read I had to correct mid-check: my first attempt to grep this log returned — sent from warm-ibex-636 |
…-172-j1j2-citation-cause-split
…-172-j3-lease-plan-refusal
…-172-j4-durability-gate
…10133 moved The roster admits a row on MEASURED serve below MEASURED recompute. My present-vs-absent pair (33669782300 / 33672356296) measured the recompute side against the pre-#10133 producer, and #10133 made exactly that recompute cheaper -- so the margin the receipt establishes is an upper bound on the margin that exists now. The measurement is not wrong; its subject moved. The argument for landing it anyway was that the serve side is nullary with an empty argument row and is therefore unlikely to have crossed. That inference is the one this roster exists to refuse: gunbc#10094 withdrew three rows enrolled on exactly that reasoning after all three were plausible and all three were wrong on measurement. A roster whose discipline is measured-not-inferred cannot admit its next row on a structural argument about which side sits near the floor. Ruling by bright-ram-778, holding me to this file's own rule rather than a new one. So floor_pure_producer_share.dag returns to main untouched, and the accessor's comment now says plainly that being preparation-forceable is a property of the declaration and NOT a claim that a row exists -- an unlanded citation is indistinguishable at the citing end. It also records that the census figures beside it predate #10133. What remains is claims 1 and 2, whose correctness never depended on a cost result. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RuWuQWB6MPkY7sNM4jEqAy
…he absent arm of the share measurement (#10108) * One home for the standing of the persisted frontier series, and the absent arm of the measurement that decides whether it may be shared FLOOR-COST-500MS, first of two commits on this branch, and the split is the measurement rather than tidiness. This commit carries the REWIRE ONLY. The roster row that shares it lands in the next commit, so this branch's own floor run is the ABSENT arm of the present-vs-absent A/B that v2.workflow.floor_pure_producer_share requires before any row is enrolled -- two runs differing in exactly one roster line, on otherwise identical trees. WHAT IS CONSOLIDATED. gunbc.self_host_compile_phase_frontier compile_phase_frontier_standing judges ANY series of receipts, and its own witnesses exercise it against perturbed ones. But there is exactly one PERSISTED series in this repository, and "the standing of that series" was being re-spelled at five call sites across four modules. That is one fact with five homes, which DESIGN section 3 puts in one place; the nullary current_compile_phase_frontier_standing is that place. The general function is untouched and every consumer that judges a hypothetical series still calls it directly -- that distinction is why both spellings stay. THE DISCRIMINATING RED IS ENROLLED, because the consolidation has exactly one failure mode and it is silent. An accessor bound to a DIFFERENT series -- a truncated one, a fixture one, a second declaration added later -- keeps compiling and keeps answering, about the wrong subject, and no existing claim in either consuming closure notices, because they all read the accessor and would agree with each other about the wrong thing. the_persisted_standing_accessor_agrees_with_the_general_fold_on_the_named_series is the only executing place the two spellings are compared; rebind the accessor to the skip(n: 1) perturbation the sibling live-gate witnesses already use and it reds while nothing else does. WHY NOW, and this half is a consequence rather than the reason. The required floor builds a fresh evaluation frame per claim, so this standing was re-derived from scratch in every claim that reads it. Measured on main run 33664768371 attempt 1, required_floor_cross_claim_demand.tsv: 22 claims, 36 evaluations, 6508ms inclusive-of-callees, with phase_board_series_ratchet (23/23) and apply_regression_transition (23/23) nested underneath. Being NULLARY, the consolidated accessor is preparation-forceable and carries an EMPTY argument row -- which is what makes it a candidate for the share tier at all, since the general function's List<CompilePhaseFrontierReceipt> argument would have to be reified and structurally verified in every consuming frame, the shape that roster's header records LOSING. Re-derive with the instruments named there, never from these sentences. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RuWuQWB6MPkY7sNM4jEqAy * The present arm: enrol the persisted frontier standing as a warm share row, against this branch's own absent measurement FLOOR-COST-500MS, second of two commits. The parent commit rewired five call sites into one nullary accessor and changed no roster; its floor run is the ABSENT arm. This commit adds exactly one roster line, so the two runs differ in one line on otherwise identical trees and the present-vs-absent comparison v2.workflow.floor_pure_producer_share requires is a property of this branch rather than a promise in a message. THE ABSENT ARM, measured on run 33669782300 attempt 1 (this branch at 0205a4d, required-witnesses-floor SUCCESS, required_floor_claim_cost.tsv, 3499 executed rows, cost_basis=cpu). Corpus p50 2ms; 18 rows at or above 400ms and none at or above 450. The two closures that consume this standing: test.claim.self_host_compile_phase_live_gate_witness 13 rows at 418-445 and test.claim.compiler_frontend_program_status_witness 8 rows at 382-409, against the 500ms per-claim CPU ceiling. 445 is 89% of budget, which is the state this lane exists to end. WHY THIS ROW IS ENROLLABLE WHERE THE CENSUS'S CANDIDATE WAS NOT. The census ranked the GENERAL fold, whose List<CompilePhaseFrontierReceipt> argument a serve must reify and structurally verify in every consuming frame -- the shape this roster's own header records LOSING, where the serve cost more than the recompute and two rows were removed. The parent commit's consolidation is what changes that: the accessor is NULLARY, so its argument row is EMPTY. There is nothing to reify and nothing to verify per serve, and being preparation-forceable the fill lands outside the fold. The losing shape is absent rather than gambled on. ONLY THE OUTERMOST PRODUCER OF THE NEST IS ENROLLED, and the two fatter rows underneath it are deliberately absent rather than overlooked: phase_board_series_ratchet (23 claims / 23 evals) and apply_regression_transition (23 / 23) are CALLEES of this standing, the census's cost columns are inclusive of callees, and its own summary line says they do not sum. Three roster rows would claim one saving three times. THIS ROW IS THE FIRST ENROLLED FROM A CANDIDATE THE CENSUS PRODUCED rather than from an incident on somebody else's pull request. gunbc.recurring_failure_mode recurrence_ledger_scoped_below_the_recurrence names that discovery path as the class: a roster whose rows each trace back to a budget refusal landing on a passing lane is a roster whose discovery instrument is blind. WHAT IT DOES NOT DO. It does not retire gunbc.rung_drop floor_cost_claim_qualification_unavailable. A per-closure constant is identical on two runs of one tree by construction, so removing it removes the LEVEL; the environment term stays and the per-claim line remains an attempt-safety boundary rather than a claim-cost verdict. This buys headroom, which is not a rung claim. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RuWuQWB6MPkY7sNM4jEqAy * Take the standing, not the series: the residual re-derivation the present arm exposed, and the measurement that found it FLOOR-COST-500MS, third commit. The present arm measured the enrolment a large win AND left three rows crossing, and this commit is the repair for the second half. Both halves come from the same run and neither is inferred. WHAT THE PRESENT ARM MEASURED (run 33672356296 attempt 1, head a71dfb3, required_floor_claim_cost.tsv, joined by identity against run 33669782300's absent arm on the same branch). The serve is far below the recompute wherever it applies: test.claim.compiler_frontend_program_status_witness 382-409ms -> 50-63ms test.claim.self_host_compile_phase_live_gate_witness 7 of 13 rows 418-445 -> 0-81 corpus rows at or above 400ms 18 -> 10 warm fill disposition=Stored, 501ms, billed to preparation THE ROWS THAT DID NOT IMPROVE ARE THE SUBJECT OF THIS COMMIT, and they are exactly the rows that never reached the shared value: live_tree_frontier_verdict took a SERIES and re-derived compile_phase_frontier_standing itself at every call, so a caller that meant the persisted series could not reach the warm entry no matter what the roster said. THE PARAMETER WAS WIDER THAN THE BODY. Every use of `series` in that fold was one call to compile_phase_frontier_standing; nothing else in it ever looked at a receipt list. A parameter a body merely passes through to a producer is a declared dependency on the wrong thing, and the width was not free -- it put the shared value one level INSIDE the parameter, where nothing can serve it. Taking the standing makes the declared dependency equal the consumed one. IT ALSO MAKES THE TWO SUBJECTS UNCONFUSABLE AT THE CALL SITE, which is a correctness gain rather than a cost one. Under the old signature `series: self_host_compile_phase_frontier_receipts` and `series: self_host_compile_phase_frontier_receipts |> skip(n: 1)` are one character apart and mean entirely different subjects. Now a caller judging the live tree against THE persisted frontier passes current_compile_phase_frontier_standing(), and a caller judging a PERTURBED series passes compile_phase_frontier_standing(series: <perturbed>) and gets a genuine re-derivation. Every perturbation witness keeps its discriminating power: the empty-series and dropped-genesis probes still fold their own series and still refuse. WHY THE ROWS THAT ROSE ARE NOT A SERVE REGRESSION, measured rather than asserted, because this is the reading the enrolment had to survive. 72 control rows at or above 250ms OUTSIDE the three consuming modules -- rows the roster row cannot serve -- moved present/absent by p50 1.04, p90 1.09, max 1.17. The consuming-module rows that rose moved 1.09-1.17, inside that band. The whole run was hotter; the rows that rose are rows the share does not reach, and they moved with the corpus rather than against the cache. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RuWuQWB6MPkY7sNM4jEqAy * Put the measured receipt in the roster row, replacing the word "expected" The row landed saying what enrolment was EXPECTED to buy, because the run that measures it had not finished. It has: codex/gpt-5.6-sol's REQUEST_CHANGES on gunbc#10108 (review 58865) is correct that an empty argument row establishes cheap KEYING and establishes nothing about the cost of reifying and serving a CompilePhaseFrontierStanding, and that this roster's own criterion demands the present-vs-absent receipt before enrolment rather than after. So the row now carries the two-run identity join, the control that separates a serve measurement from a hot-and-cold pair of runs, and its own coverage boundary. Nothing here is a new claim: the numbers are the artifacts of runs 33669782300 and 33672356296, both attempt 1, on this branch's own trees. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RuWuQWB6MPkY7sNM4jEqAy * The last row paying for itself: ask the standing, not the ratchet, and leave exactly one payer in the corpus FLOOR-COST-500MS, fifth commit, and it removes the corpus's single highest claim. After the enrolment and the parameter narrowing, floor run 33675891779 attempt 1 (head 020ca87, required-witnesses-floor SUCCESS) puts ZERO rows at or above 500ms and ONE at or above 450: this one, at 461ms -- 92% of the ceiling, in a 3499-row corpus whose p50 is 2ms. Leaving it would hand the next lane a smaller copy of the problem this one was chartered to remove. IT WAS ASKING THE INNER PRODUCER. phase_board_series_ratchet(series: <the persisted series>) re-derives the whole fold in this claim's own frame, while the value the fold produces is exactly what the share tier now holds warm one level up. Asking current_compile_phase_frontier_standing() reads that value. THE ASSERTION GETS STRICTLY STRONGER RATHER THAN WEAKER, which is worth checking rather than assuming: CompilePhaseFrontierValidated is PhaseBoardSeriesHeld AND a latest receipt existing. A held series with no latest, which the old spelling accepted, now refuses -- the fail-closed direction. WHAT IT COSTS, DECLARED. This row now trusts the served standing instead of re-deriving it, so a store serving a WRONG value would take it green. That is not uncovered: the_persisted_standing_accessor_agrees_with_the_general_fold_ on_the_named_series deliberately keeps paying the full recompute so exactly ONE claim in the corpus compares the served value against a freshly folded one. One payer, one cross-check, every other consumer reading the shared value -- the structure the share tier exists to create, and sound only while that witness stays enrolled, which is why its enrolment is now load-bearing rather than incidental. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RuWuQWB6MPkY7sNM4jEqAy * Correct an overclaim: the tier guarantees the serve, and the witness guards authoring — two hazards this lane conflated WITHDRAWN, NOT REWORDED. The previous commit's comment said current_persisted_compile_phase_frontier_holds is safe to read the served standing BECAUSE the_persisted_standing_accessor_agrees_with_the_general_fold_ on_the_named_series acts as the corpus's cache cross-check -- "one payer, one cross-check, N readers, sound only while that witness stays enrolled". That is wrong in a way worth naming rather than quietly editing, because it would have made v2.workflow.floor_pure_producer_share's correctness depend on ONE TEST IN ONE MODULE, and any later lane deleting or weakening that test would have been silently deleting a guarantee it had no reason to know it held. THE SERVE IS THE TIER'S OWN GUARANTEE AND ALWAYS WAS. That roster keys every entry on FN-NODE IDENTITY plus a content hash of the full argument row and serves only after the stored row verifies structurally equal in portable form, a hash collision degrading to recompute rather than to a wrong value. For a NULLARY producer the argument row is empty, so a wrong serve would require one declaration's value returned under another declaration's node identity. That invariant carries the tier's own evidence; re-asserting it in a witness would be a second authority for it, which is the defect this whole branch exists to remove. TWO HAZARDS WERE CONFLATED. The witness's actual subject is AUTHORING -- whether the accessor names the series its name claims -- which is decidable without the cache and says nothing about serving. Its removal is an ordinary loss of a discriminating red, which is a real loss, and not a silent conversion of a verified share into an unverified one. Nothing executable changes. The overclaim was in prose, which is exactly the kind of claim DESIGN 4b(1) calls rung inflation on the compiler's own self-description, and it is retracted at the carrier that made it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RuWuQWB6MPkY7sNM4jEqAy * The sharing half, measured by a second instrument on a second branch This roster admits a row on measured recompute AND measured sharing -- a non-sharing producer is cache population rather than a saving, which is the ground formal_production_for_lhs_exact was REMOVED on. This row carried the recompute and the cost drop and did NOT carry a direct observation of the sharing, which is a gap in its own justification rather than a decoration. gunbc#10068's lane closed their PR in favour of this one and handed over the observation that fills it: the required floor's own `[floor-shared-fill]` ledger, green run 33649705841, recording this producer at `fill_ms=362 inclusive_ms=362 consumer_claims=22 consumer_modules=3 disposition=shared`, beside four `fill_ms=0 disposition=exclusive` fills for the perturbed-series arguments that correctly populate nothing. IT IS INDEPENDENT IN EVERY RESPECT THAT COULD HAVE MADE IT CIRCULAR: different branch, different instrument, and a different KEY SHAPE -- that lane enrolled the general declaration claim-forced on the reified receipt series, where this row is nullary. Two instruments sharing no accounting and agreeing on 22 consumer claims across 3 modules is what rules out the reading that the sharing is an artifact of one census's bookkeeping. Recorded WITH ITS ORIGIN rather than absorbed into this lane's figures, because a corroboration whose provenance is dropped is indistinguishable from a second derivation by the same author -- which would make two observations read as one and inflate exactly the confidence the corroboration was supposed to earn. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RuWuQWB6MPkY7sNM4jEqAy * Hoist the flatten out of the loop: the same parameter-wider-than-its-body defect, second instance in one module FLOOR-COST-500MS. The payer this branch declared as residue CROSSED on head 62989fd -- run 33683717936, failed=0, interrupted_cpu_deadline=1, the single interrupted row being the_persisted_standing_accessor_agrees_with_the_general_fold_on_the_named_series, the row named in advance -- on a head differing from a green one by a COMMENT BLOCK IN A .dag FILE. Predicted before it was observed, on a diff that cannot have caused it. This commit removes the cost rather than rerolling the run. NOT REROLLED, DELIBERATELY. gunbc.rung_drop floor_cost_claim_qualification_unavailable admits a bounded reroll for exactly this signature, and that admission is for lanes refused by SOMEONE ELSE'S marginal row. This row is this branch's own, its crossing was forecast, and rerolling it until it greens would be the absorbing arm under the name of the lane funded to remove it. TWO DEFECTS, ONE SHAPE, BOTH IN gunbc.self_host_compile_phase_frontier. (1) THE SET DIFFERENCES RE-FLATTENED PER ELEMENT. board_added_identities and board_removed_identities wrote `board_all_identities(board: other)` INSIDE the filter predicate, so a flat_map over every row of one board ran once per identity of the other -- O(|current| x |previous|) flattens where one flatten per side suffices. The inner list is now bound once. (2) THE PARAMETER WAS WIDER THAN THE BODY -- SECOND INSTANCE ON THIS BRANCH, and the first one is what settled the #10068 collision. newly_exposed_identity_admitted took `previous: RustcPhaseBoard` and `current: RustcPhaseBoard` and never looked at a board: it used each for the flattened identity list and for .furthest_phase_reached, nothing else. Its caller invokes it ONCE PER ADDED IDENTITY through `added |> filter(...)`, and a second caller through `all(ids, ...)`, so both boards were re-flattened per element, twice over, while the flattened lists are constant across the whole call. Taking the identity list and the phase lets each caller bind them once. NARROWED RATHER THAN WIDENED, BOTH TIMES, AND THAT IS THE RESOLUTION RATHER THAN A STYLE CHOICE: passing the board AND its flattened identities would put two representations of one fact in one signature, which is the defect and not the fix. The general form -- a parameter a body only funnels into a producer puts the shared value one level INSIDE the parameter, where nothing can hoist or share it -- is now twice-instanced in one module and is a class rather than an anecdote. DESIGN section 6's bare-minimum-cost rule fixes a proven cost shape regardless of the realized n. This n is realized: these folds put their callers at 89-93% of the floor's per-claim CPU ceiling. MEASURED, NOT INFERRED, AND THE NEGATIVE IS ADMISSIBLE: the before is this branch's own 448ms (run 33675891779) and 467ms (run 33683717936, where it crossed); the after is the next run on this branch, joined at identity, with the 72-row ambient control re-derived on the new pairing. If the row does not move, the honest answer becomes a declared expected-over-budget row and this lane reports that rather than attempting again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RuWuQWB6MPkY7sNM4jEqAy * Hoist the two body-interior annotations to module-item grain: §4c admits only a leading block The parse phase refused #10108's head at self_host_compile_phase_frontier.dag:986-987 -- two `//` lines sitting inside apply_regression_transition's body. §4c models annotation capture at module-item grain only; a body-interior comment has no subject to attach to, so accepting it would either fabricate an attachment or discard the text, and the compiler refuses instead of choosing. All three witness-lane phases failed from that one cause (namespace-wave-admission had no head index, the floor had no subject), so the run carried no cost opinion at all -- this red is not the class the PR exists to remove. The rationale survives the move rather than being reworded into something vague: the leading block now states what the declaration does at its own grain -- both boards are fixed for the transition, so flattening per added identity would be authored duplication (DESIGN §2), not a recurrence a cache could discharge. Reported by bright-ram-778, who read the job log before assuming the cost class. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RuWuQWB6MPkY7sNM4jEqAy * Split claim 3 out: the share row's admission rests on a denominator #10133 moved The roster admits a row on MEASURED serve below MEASURED recompute. My present-vs-absent pair (33669782300 / 33672356296) measured the recompute side against the pre-#10133 producer, and #10133 made exactly that recompute cheaper -- so the margin the receipt establishes is an upper bound on the margin that exists now. The measurement is not wrong; its subject moved. The argument for landing it anyway was that the serve side is nullary with an empty argument row and is therefore unlikely to have crossed. That inference is the one this roster exists to refuse: gunbc#10094 withdrew three rows enrolled on exactly that reasoning after all three were plausible and all three were wrong on measurement. A roster whose discipline is measured-not-inferred cannot admit its next row on a structural argument about which side sits near the floor. Ruling by bright-ram-778, holding me to this file's own rule rather than a new one. So floor_pure_producer_share.dag returns to main untouched, and the accessor's comment now says plainly that being preparation-forceable is a property of the declaration and NOT a claim that a row exists -- an unlanded citation is indistinguishable at the citing end. It also records that the census figures beside it predate #10133. What remains is claims 1 and 2, whose correctness never depended on a cost result. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RuWuQWB6MPkY7sNM4jEqAy --------- Co-authored-by: Brian Searls <briansearls1@gmail.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…au is real, the promotion mechanism is not, and #10133 was an outlier (#10143) * Measure the floor cost distribution near the 500ms ceiling: the plateau is real, the promotion mechanism is not, and #10133 was an outlier rather than one step of a treadmill MEASUREMENT, NOT A REPAIR, commissioned after gunbc#10133 to decide whether further per-producer work is a treadmill. Nothing here ran a floor: every figure is a reading of `required-floor-claim-cost`, `required-floor-cross-claim-demand` and the `required-floor:` summary already published by ten main runs. THREE FINDINGS. The band IS dense. 87 rows sit above 300ms and 32 above 400ms, and repairing modules top-down returns about 10ms per module below the top of the band -- the treadmill, quantified. The new crossers do NOT share a root the way #10133's three did. That trio reported 612,662-658,336 eval steps, a 1.07x spread that IS the signature of one shared dominant producer. The families now near the ceiling span 63k-197k, a 3.1x spread, in three distinct shapes. One shared-producer lead exists (`bind_outcome`) and is explicitly not a plan: the artifact ranking it declares its own columns inclusive-of-callees and non-summable, and they total 14x the run's actual claim CPU. The margin is measurable and uniform. Same-identity run-to-run inflation is median 1.16x, p90 1.35x, max ~2.0x, and it does NOT vary with cost bucket or with host-call heaviness -- both hypotheses tested and refuted, which is what makes one global margin defensible. PR runs reach ~1.6x, implying a clean-run budget near 312ms. AND ONE CORRECTION TO THE FRAMING THIS WAS COMMISSIONED UNDER. Promotion is refuted: rank does not cause a crossing and removing rows above a row does not make it slower. The five rows that crossed post-repair went ABOVE their own pre-repair maxima (median 343ms, max 472ms over 45 observations, against 504-533ms), on the most contended run in the sample. Replaying all ten runs with the repaired modules excluded takes the refusal rate from 5/10 to 1/10, and across the nine pre-repair runs no row outside those three modules ever crossed. They were 17 of the 22 distinct identities that crossed anywhere: an outlier, not a sample of the plateau. Raising the ceiling is named as not-an-option, because `v2.workflow.required_floor` already litigated it twice and repudiated both raises. An eleventh run's artifact downloaded empty and is excluded; recorded because it silently emptied a ten-way intersection before it was caught. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WEFMasGn2cAGCMJW2694pb * Give the floor cost analysis an entry point: the derivations become a witnessed authority and the memo cites a producer instead of carrying figures (review 58983) REVIEW 58983 REFUSED #10143 AND WAS RIGHT. The memo named `gunbc.witness_floor_workflow` as its instrument. That workflow produces the RAW ROWS and performs none of the analysis: the ten-run intersection, the band histogram, the counterfactual replay and the inflation percentiles existed only as prose plus numbers, so a reader could check the arithmetic of nothing. DESIGN section 6 is explicit -- name the instrument, never transcribe its output, and if a measurement is worth re-deriving it is worth an entry point. WHAT LANDS. `gunbc.floor_cost_distribution` carries the derivations as pure functions over rows supplied by the caller; `tools.floor_cost_distribution_instrument` reads the artifacts and renders the report. That is the section 3 split between what the analysis IS and where the bytes come from, and it is what lets `test.claim.floor_cost_distribution_witness` exercise every derivation on a hand-built fixture with no run artifacts present -- 14 rows, each with a discriminating red beside it. VERIFIED BY REPRODUCTION, not by inspection. Run against the ten artifacts the memo names, the instrument returns every figure the memo carries: the band histogram (2231/502/441/148/89/55/27/5), 87 rows over 300ms and 32 over 400ms, 5 runs refused as observed against 1 with the repaired modules excluded, worst row 533ms, the full ten-step ladder (525, 511, 490, 487, 486, 466, 458, 450, 434, 429), and inflation percentiles of 1147/1345/1433/1643 permille implying 435/371/348/304ms clean-run budgets. Those figures were originally derived by an independent ad-hoc script; two implementations agreeing is the cross-check. THE REFUSAL ARMS ARE PRODUCED, NOT DECLARED. An absent or empty artifact refuses and is named rather than contributing an empty population -- exactly the failure that silently emptied a ten-way identity intersection during the original analysis, where a run reading as zero rows is indistinguishable from a run in which nothing crossed. Both arms are witnessed, with a discriminating red proving a complete sample produces no refusal line, and the actuator's refusal is exercised by execution in both directions. A NOTE ON WHY THIS IS AN INSTRUMENT AND NOT A SCAFFOLD, since the distinction decided the shape: a scaffold is what the terminal architecture throws away, an instrument is what it keeps and re-runs. This analysis will be re-derived after the next floor change, so it earns a modeled home under dag/gunbc/instruments/ rather than a committed shell or python script. The memo keeps its figures as illustrative of what the instrument returned on the runs it names, and says plainly that the instrument is the authority where the two disagree. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WEFMasGn2cAGCMJW2694pb * Route the floor cost analysis's whole time axis through std.measure Millisecond (review 59000) REVIEW 59000 IS RIGHT AND THE SCOPE IT NAMES IS THE RIGHT SCOPE. The finding is not six leaky fields, it is that the module carried a domain unit vocabulary end-to-end -- rows, bands, ladder, threshold, budget -- as bare `Int` milliseconds with no `std.measure` import at all. That is a parallel representation of a quantity this repository already models, which is DESIGN section 2's "net concepts must not grow by re-invention" and section 3's single authority. The corpus had already ruled on exactly this shape: gunbc.witness_row_cost's own seed note records constructing Millisecond through the interpreter specifically "so no flat-scalar duration crosses the seam". WHAT CHANGED. `FloorCostRow.cpu_ms`, `CostBand.lower_ms`/`upper_ms`, `LadderStep.worst_ms`, every ceiling and threshold parameter, the band edges, and the implied budget are `Millisecond`. The parser wraps at the boundary, where the artifact's integer column becomes a carried quantity. Comparison goes through `millisecond_count` rather than over the carrier, because the carrier is what stops a figure read from one clock being compared against a figure read from another and the count is the projection a comparison is entitled to. WHAT DELIBERATELY STAYS `Int`, so this is not a blanket sweep: row counts and refused-run counts are cardinalities, and the inflation permille and its percentile argument are a DIMENSIONLESS RATIO. Giving a ratio a time unit would be the same error in the other direction. `implied_clean_run_budget_ms` is renamed `implied_clean_run_budget` -- the `_ms` suffix was the field name doing the unit's job, which is the thing the carrier replaces. VERIFIED BY REPRODUCTION, WHICH IS THE POINT OF DOING IT THIS WAY. A unit refactor is exactly where an off-by-conversion hides, so the control is that the instrument's output over the ten named artifacts is BYTE-IDENTICAL to the run before this commit -- same bands, same 87/32, same 5 and 1, same ladder, same 1147/1345/1433/1643 permille and same 435/371/348/304ms budgets. All 14 witness rows still pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WEFMasGn2cAGCMJW2694pb --------- Co-authored-by: gunbc-ci-auto-heal <gunbc-ci-auto-heal@users.noreply.github.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Forward-delta adjudication of this PR's tested base against current main found one MATERIAL commit that no path-overlap check would have surfaced: gunbc#10143 measures the near-ceiling population this document's screening section ranks, and landed after the base. Two results of that measurement bound what any screening here can buy: 1. THE BAND IS A PLATEAU, no gap below the ceiling. Repairing top-down, the first three modules buy 43ms and the next seven buy 61ms -- ~10ms per module, about one fiftieth of the ceiling. So surfacing the next-most-expensive producer has a small payoff BY CONSTRUCTION, and this document's own ordering (safety-relevant rows ahead of milliseconds) is the better one for reasons now measured. 2. THE CROSSERS DO NOT GENERALLY SHARE A ROOT. #10143 records that the modules newly near the ceiling do not share a root the way #10133's three did, so the shared-derivation bucket described here is the exception, not the common case. Also records the one worked counterexample, gunbc#10170: four rust_body_add_emit rows reaching five canonical fixtures through a ~5k-line module, cut from 506/425/416/362 to 180/174/172/168. The SPREAD collapsing 144ms -> 12ms is what distinguishes a genuine shared root from four rows each getting cheaper -- a constant subtracted from four independent costs leaves the spread. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019LhF5WCbZqrZHPqsnjpkYu
…xt reader of "the floor is a coin flip" will hit it (#10196) * Measure the floor cost distribution near the 500ms ceiling: the plateau is real, the promotion mechanism is not, and #10133 was an outlier rather than one step of a treadmill MEASUREMENT, NOT A REPAIR, commissioned after gunbc#10133 to decide whether further per-producer work is a treadmill. Nothing here ran a floor: every figure is a reading of `required-floor-claim-cost`, `required-floor-cross-claim-demand` and the `required-floor:` summary already published by ten main runs. THREE FINDINGS. The band IS dense. 87 rows sit above 300ms and 32 above 400ms, and repairing modules top-down returns about 10ms per module below the top of the band -- the treadmill, quantified. The new crossers do NOT share a root the way #10133's three did. That trio reported 612,662-658,336 eval steps, a 1.07x spread that IS the signature of one shared dominant producer. The families now near the ceiling span 63k-197k, a 3.1x spread, in three distinct shapes. One shared-producer lead exists (`bind_outcome`) and is explicitly not a plan: the artifact ranking it declares its own columns inclusive-of-callees and non-summable, and they total 14x the run's actual claim CPU. The margin is measurable and uniform. Same-identity run-to-run inflation is median 1.16x, p90 1.35x, max ~2.0x, and it does NOT vary with cost bucket or with host-call heaviness -- both hypotheses tested and refuted, which is what makes one global margin defensible. PR runs reach ~1.6x, implying a clean-run budget near 312ms. AND ONE CORRECTION TO THE FRAMING THIS WAS COMMISSIONED UNDER. Promotion is refuted: rank does not cause a crossing and removing rows above a row does not make it slower. The five rows that crossed post-repair went ABOVE their own pre-repair maxima (median 343ms, max 472ms over 45 observations, against 504-533ms), on the most contended run in the sample. Replaying all ten runs with the repaired modules excluded takes the refusal rate from 5/10 to 1/10, and across the nine pre-repair runs no row outside those three modules ever crossed. They were 17 of the 22 distinct identities that crossed anywhere: an outlier, not a sample of the plateau. Raising the ceiling is named as not-an-option, because `v2.workflow.required_floor` already litigated it twice and repudiated both raises. An eleventh run's artifact downloaded empty and is excluded; recorded because it silently emptied a ten-way intersection before it was caught. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WEFMasGn2cAGCMJW2694pb * Give the floor cost analysis an entry point: the derivations become a witnessed authority and the memo cites a producer instead of carrying figures (review 58983) REVIEW 58983 REFUSED #10143 AND WAS RIGHT. The memo named `gunbc.witness_floor_workflow` as its instrument. That workflow produces the RAW ROWS and performs none of the analysis: the ten-run intersection, the band histogram, the counterfactual replay and the inflation percentiles existed only as prose plus numbers, so a reader could check the arithmetic of nothing. DESIGN section 6 is explicit -- name the instrument, never transcribe its output, and if a measurement is worth re-deriving it is worth an entry point. WHAT LANDS. `gunbc.floor_cost_distribution` carries the derivations as pure functions over rows supplied by the caller; `tools.floor_cost_distribution_instrument` reads the artifacts and renders the report. That is the section 3 split between what the analysis IS and where the bytes come from, and it is what lets `test.claim.floor_cost_distribution_witness` exercise every derivation on a hand-built fixture with no run artifacts present -- 14 rows, each with a discriminating red beside it. VERIFIED BY REPRODUCTION, not by inspection. Run against the ten artifacts the memo names, the instrument returns every figure the memo carries: the band histogram (2231/502/441/148/89/55/27/5), 87 rows over 300ms and 32 over 400ms, 5 runs refused as observed against 1 with the repaired modules excluded, worst row 533ms, the full ten-step ladder (525, 511, 490, 487, 486, 466, 458, 450, 434, 429), and inflation percentiles of 1147/1345/1433/1643 permille implying 435/371/348/304ms clean-run budgets. Those figures were originally derived by an independent ad-hoc script; two implementations agreeing is the cross-check. THE REFUSAL ARMS ARE PRODUCED, NOT DECLARED. An absent or empty artifact refuses and is named rather than contributing an empty population -- exactly the failure that silently emptied a ten-way identity intersection during the original analysis, where a run reading as zero rows is indistinguishable from a run in which nothing crossed. Both arms are witnessed, with a discriminating red proving a complete sample produces no refusal line, and the actuator's refusal is exercised by execution in both directions. A NOTE ON WHY THIS IS AN INSTRUMENT AND NOT A SCAFFOLD, since the distinction decided the shape: a scaffold is what the terminal architecture throws away, an instrument is what it keeps and re-runs. This analysis will be re-derived after the next floor change, so it earns a modeled home under dag/gunbc/instruments/ rather than a committed shell or python script. The memo keeps its figures as illustrative of what the instrument returned on the runs it names, and says plainly that the instrument is the authority where the two disagree. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WEFMasGn2cAGCMJW2694pb * Route the floor cost analysis's whole time axis through std.measure Millisecond (review 59000) REVIEW 59000 IS RIGHT AND THE SCOPE IT NAMES IS THE RIGHT SCOPE. The finding is not six leaky fields, it is that the module carried a domain unit vocabulary end-to-end -- rows, bands, ladder, threshold, budget -- as bare `Int` milliseconds with no `std.measure` import at all. That is a parallel representation of a quantity this repository already models, which is DESIGN section 2's "net concepts must not grow by re-invention" and section 3's single authority. The corpus had already ruled on exactly this shape: gunbc.witness_row_cost's own seed note records constructing Millisecond through the interpreter specifically "so no flat-scalar duration crosses the seam". WHAT CHANGED. `FloorCostRow.cpu_ms`, `CostBand.lower_ms`/`upper_ms`, `LadderStep.worst_ms`, every ceiling and threshold parameter, the band edges, and the implied budget are `Millisecond`. The parser wraps at the boundary, where the artifact's integer column becomes a carried quantity. Comparison goes through `millisecond_count` rather than over the carrier, because the carrier is what stops a figure read from one clock being compared against a figure read from another and the count is the projection a comparison is entitled to. WHAT DELIBERATELY STAYS `Int`, so this is not a blanket sweep: row counts and refused-run counts are cardinalities, and the inflation permille and its percentile argument are a DIMENSIONLESS RATIO. Giving a ratio a time unit would be the same error in the other direction. `implied_clean_run_budget_ms` is renamed `implied_clean_run_budget` -- the `_ms` suffix was the field name doing the unit's job, which is the thing the carrier replaces. VERIFIED BY REPRODUCTION, WHICH IS THE POINT OF DOING IT THIS WAY. A unit refactor is exactly where an off-by-conversion hides, so the control is that the instrument's output over the ten named artifacts is BYTE-IDENTICAL to the run before this commit -- same bands, same 87/32, same 5 and 1, same ladder, same 1147/1345/1433/1643 permille and same 435/371/348/304ms budgets. All 14 witness rows still pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WEFMasGn2cAGCMJW2694pb * Record that the floor-crosser premise is measured false, where the next reader of "the floor is a coin flip" will hit it THE BELIEF THIS RETIRES IS THE ONE THAT COSTS A NIGHT. The work item behind gunbc#10133's follow-up was "three (later seven) self-host witness modules sit ABOVE the 500ms hard CPU ceiling". Measured, no row in that family sits above the ceiling, both remedies section 5 names are refused for it on evidence, and the ceiling is correctly denominated. That result lived only in a session thread, and session threads archive -- so it is written where section 5's open question is, because the reader who re-forms this belief will arrive from the same symptom: a red floor lane naming a handful of emit rows. THREE MISREADINGS, EACH WITH THE CONTROL THAT EXPOSES IT, because the numbers are less reusable than the way they mislead. An `interrupted_before_verdict` figure is the BUDGET, not the row -- every interrupted row reports approximately whatever ceiling stopped it, so a set of them looks like a tight cluster "just over" whatever the true costs are. The crossing set is NOT the expensive set -- the row that passed on CI is the most expensive of its four siblings, so repairing "the rows that crossed" fits noise, and the non-crossing siblings are what reveal the band. And a one-anchor local-to-CI factor does not transfer: one taken from a single identity predicted a row comfortably under the ceiling that CI had in fact interrupted. NAMED, NOT TRANSCRIBED (section 6). The join has no instrument module -- it needs a second axis the section 5 instrument does not read -- so the recipe is given at identity grain and the absence of an owning authority is stated as a gap rather than papered over. Ratios are described by their shape and stability across named run classes; re-derive them rather than quote this page. THE HEDGE IS PRESERVED EXACTLY. The projection term is recorded as measured at the `target_project_arrow_body_to_value_expression` FRAME and attributed BY ELIMINATION to the primitive-apply step beneath it, with the neighbouring candidates named as measuring zero and the two table-shaped ones noted as tested with inputs that would have exposed a build even on a miss. The live lead is left UNDETERMINED between a memo keyed above the body and a one-time warm, because those want different providers and the mechanism must be separated before any provider is proposed. Both prior sharing attempts are recorded with their refutations so neither is re-proposed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WEFMasGn2cAGCMJW2694pb * Retire the stable-subject verb: no fixed above-ceiling identity set, not "no row sits above the ceiling" (review 5098047617) THE REVIEW IS RIGHT AND THE VERB WAS THE WHOLE DEFECT. The section opened with "Measured, no row in that family sits above the ceiling", and its own later evidence refutes that sentence: failed-run inflation carries the top of the band across 500ms, and four completed-over-cost rows were observed at 506-515ms. "Sits" asserts a STABLE ROW PROPERTY while the section's central finding is that crossing is ATTEMPT-DEPENDENT -- so the opening claimed the one shape the body spends four subsections dismantling. WHAT THE SECTION NOW SAYS is the proposition actually established: no fixed above-ceiling identity set exists; completed green runs place the family's band below 500ms, while run variance can carry its top across. What is retired is that the family is INTRINSICALLY above the ceiling -- a row does not stably sit on either side of an attempt-safety boundary -- and explicitly NOT the observed crossings, which are real and recur. The 506-515ms observations are named in the retraction itself so the correction cannot be read as denying them. The heading changes for the same reason: "the crossers are not above the ceiling" was the same stable-subject shape one line up. NOTHING ELSE MOVES. The two other scoped statements were already bounded correctly -- the restated subject says "on a normal run" and "a property of the run, not of the rows", and the emit_host_fold example is scoped to "every green run in the join". The re-derivation recipe, the three misreading controls, the eval_steps axis and the producer-share caveat all stand. WHY THIS CLASS IS WORTH THE COMMIT MESSAGE: when a subject's true form is a DISTRIBUTION, ordinary language keeps snapping it back to a PROPERTY, and it does so in both directions -- the same data produced "the rows are genuinely under budget" elsewhere within the hour. Neither sentence survives its own evidence. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WEFMasGn2cAGCMJW2694pb --------- Co-authored-by: gunbc-ci-auto-heal <gunbc-ci-auto-heal@users.noreply.github.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
… discriminator, and the signature screen (#10114) * Share the compile-phase frontier fold instead of re-folding it once per witness The floor's own cross-claim demand census names one derivation that 23 required claims each recompute and none reuse: gunbc.self_host_compile_phase_frontier phase_board_series_ratchet, at 23 claims against 23 evals across three modules in required_floor_cross_claim_demand.tsv (main run 33668368846). The two modules it dominates are that run's two largest populations over the 100ms cost line in required_floor_claim_cost.tsv, and inside each of them every over-line row lands within a few percent of ~615k eval steps -- the shape of one fold repeated, not of N witnesses doing their own work. So the remedy is the roster, not the witnesses: splitting them would only re-attribute the fold to whichever fragment ran first, which is witness_row_cost's standing witness_decomposition_does_not_reduce_entry_cost_note. Enrolled claim-forced at the root rather than at compile_phase_frontier_standing -- the census shows the standing's cost is inclusive of this fold, so one row discharges both and reaches one more claim. The value is PhaseBoardSeriesVerdict, a nullary variant on the held arm, so the serve walks nothing; that is the admission criterion the two rust target models failed, and the reason they stay out. Present-vs-absent is re-derived from required_floor_claim_cost.tsv and required_floor_cross_claim_demand.tsv on this branch against run 33668368846 -- the instruments named in the roster's own admission criterion. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDKZ5pDN7cXTSEvYST2y78 * The lane's ownership question has an instrument: point the census doc at the cross-claim demand ledger Everything in this document infers ownership from timings and concludes a timings-only census cannot see whose work a cost is. The floor now emits required_floor_cross_claim_demand.tsv on every required run -- claims against evals per producer identity, which answers that question directly -- and a reader landing on this document was not being told it exists. Also records the two findings from re-deriving the census on run 33668368846: the log's over-cost head is capped at 25 rows against 295 in the artifact (instrument_output_read_as_subject_content, the second occurrence on this lane), and eval_steps rather than milliseconds is the within-module discriminator, because steps are deterministic and so a tight step cluster cannot be a quiet runner -- the confound that demoted the variance screen and the cluster prior. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDKZ5pDN7cXTSEvYST2y78 * Supply the controlled present-vs-absent receipt the admission criterion requires review 58870 was right: the row was enrolled with the recompute half measured and the serve half only promised, which is the test the two rust target models are excluded on. The recipe also named a run at a different head, which cannot answer the serve question at all. The pair is now controlled -- main run 33671815204 at 20caf4e, my branch merged onto that same head as run 33673657876, both executed=3498, differing by this one line -- and the whole-run total is the conjunct that answers it, because the serve lands on the consuming claims and a per-row drop alone cannot see it: claim_cpu_total_ms 110998 -> 93110, the ratchet 23 evals -> 1, the three modules 11833ms -> 4260ms with all 21 over-line rows now under it, and failed / unexpected_failures 0/0 on both arms. The result that matters more than the milliseconds: main reports cpu_deadline=2 and both preempted rows are this module's discriminating REDs, cut off by the 500ms deadline before reaching a verdict. This branch reports cpu_deadline=0 and both answer. The fold was not merely expensive, it was suppressing the two controls the witness exists for. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDKZ5pDN7cXTSEvYST2y78 * Screen census candidates by signature, not by cost: the ratio rule and three by-inspection exclusions The demand census ranks by cost and cost is the wrong sort key. Three large classes on it cannot be shared at all, and the fn-typed-parameter class refuses at PUBLICATION rather than losing a measurement -- so a controlled pair spent on one produces noise, not a negative result. I proposed such a candidate in this lane and withdrew it on reading the signature; the screening test is what stops the next person paying two CI runs for that. The rule is a ratio, not a size heuristic: serve scales with argument plus value (hash, structural verify, container rebuild), recompute scales with the derivation between them. phase_board_series_ratchet carries a whole receipt series as its key and still admits, because 585,854 eval steps amortise the verify. The fn-in-key exclusion is grounded in the seed rather than inferred from the roster's prose: reification refuses Value::Fn as OriginBoundNode, and arguments reify on the same path via portable_args_from_ctx -> RefusedArgsNotPortable. Also records that the deadline-preempted rows on that pair's absent arm were the witness's discriminating REDs, so this population is ordered by which refusals are not executing, not only by milliseconds. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDKZ5pDN7cXTSEvYST2y78 * Record the circularity: the repair for the deadline class travels through the gate the class refuses This is the session's most consequential structural finding and it existed only in a message thread, which dies. A preemption is runner-dependent, so any PR can draw a refusal from two rows that are not broken -- including the PR carrying the repair. Observed in situ rather than argued: a 138-line documentation-only change with no .dag, no Rust and no workflow was refused by a_live_tree_that_gained_an_identity_refuses_and_names_it at 501ms/500ms while the consolidation that would make that fold cheap sat open with its own floor lane running. Also states plainly why the obvious exit is closed. A re-run does not make a row cheaper, it redraws the runner, and the green it buys is a green over refusals that did not execute -- the interrupted rows ARE the discriminating REDs, so the mark asserts a verdict for precisely the claims that were preempted. cpu_deadline is a safety counter, not a performance one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDKZ5pDN7cXTSEvYST2y78 * Three method claims that outran their instruments (review 5096548942) All three are prose defects in the appendix, not CI failures. The exact-head run is terminal success and remains a DRAW, not a repair. 1. SAFETY RANKING WAS CLAIMED, AND IS NOT AVAILABLE. The text said to rank the population by which refusals are not executing, and that the interruption diagnostics say whether a discriminating RED is the one cut off. They do not. They name the victim; they do not author what the victim exists to prove. std.witness_purpose is explicit that purpose is AUTHORED, NOT INFERRED, and it currently has zero declarers and zero consumers. The two live-gate rows are refusal probes because their source was read one at a time -- that does not generalise to a corpus-wide procedure by reading names off a diagnostic. Now stated as three separate claims with only the first two discharged: victim identity is observable; these two were independently verified; ranking by safety relevance REMAINS BLOCKED until authored purpose can be joined against verdict_reached. 2. THE SIGNATURE SCREEN CONTRADICTED ITS OWN RULE. It said the rule "screens OUT -- it never admits" and then called one shape "the shape that admits" and another "admissible". In this repository admit is a loaded verb: only the controlled present-versus-absent receipt gets to say it. Both now say the shape SURVIVES THE EXCLUSION SCREEN and is worth measuring. 3. eval_steps WAS ASKED AN OWNERSHIP QUESTION. It measures WORK. Near-identical step counts across rows do not identify a shared producer -- the same total can arise from unrelated derivations of similar size. Ownership is established by the cross-claim demand artifact, which is keyed by PRODUCER IDENTITY. The shared-derivation conclusion is now a JOIN of the two instruments: demand names the shared producer, eval_steps characterises the work reaching it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019LhF5WCbZqrZHPqsnjpkYu * File the third truncation: module_sample caps at 8 while the row reports modules=71 The two print caps this document already records have trailers. This one does not: required_floor_cross_claim_demand.tsv's module_sample column caps at 8 modules and 1044 of 26317 rows exceed it, so a row lists eight consumers beside a modules=71 and a reader joining producers to modules gets a silently partial answer. It bites exactly where the join matters -- the truncated rows are the widely shared producers, which are the ones most likely to explain why a whole module's rows cluster -- so a corpus-scale module-to-producer join is not performable from this artifact today, and the per-module reading here is sound only where the join was done by hand. That qualification is owed: a 73%-of-over-line-CPU classification was reported off the step ratio with the demand join performed for only the two modules acted on, which is a step-cluster inference at that strength rather than an ownership result. Named as an obligation, deliberately not undertaken here so it can be sized against other work rather than absorbed into this lane. Three truncations on one lane is a pattern, and the reusable shape is stated: an instrument can report a cap honestly at the top level while the column a consumer reads silently answers for fewer. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDKZ5pDN7cXTSEvYST2y78 * Bound the screening by what gunbc#10143 measured Forward-delta adjudication of this PR's tested base against current main found one MATERIAL commit that no path-overlap check would have surfaced: gunbc#10143 measures the near-ceiling population this document's screening section ranks, and landed after the base. Two results of that measurement bound what any screening here can buy: 1. THE BAND IS A PLATEAU, no gap below the ceiling. Repairing top-down, the first three modules buy 43ms and the next seven buy 61ms -- ~10ms per module, about one fiftieth of the ceiling. So surfacing the next-most-expensive producer has a small payoff BY CONSTRUCTION, and this document's own ordering (safety-relevant rows ahead of milliseconds) is the better one for reasons now measured. 2. THE CROSSERS DO NOT GENERALLY SHARE A ROOT. #10143 records that the modules newly near the ceiling do not share a root the way #10133's three did, so the shared-derivation bucket described here is the exception, not the common case. Also records the one worked counterexample, gunbc#10170: four rust_body_add_emit rows reaching five canonical fixtures through a ~5k-line module, cut from 506/425/416/362 to 180/174/172/168. The SPREAD collapsing 144ms -> 12ms is what distinguishes a genuine shared root from four rows each getting cheaper -- a constant subtracted from four independent costs leaves the spread. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019LhF5WCbZqrZHPqsnjpkYu * Restore the union: my eae0cca reverted the b561d7c corrections, and the re-fix dropped the truncation MY FAULT FIRST. eae0cca was written by a script that read the file from a STALE WORKING TREE, so it silently reverted all three of b561d7c's corrections -- the work-is-not-ownership paragraph, the survives-the-screen wording, and the safety-ranking-remains-blocked statement. I then reported "I am done pushing" while having undone someone else's fix without noticing. Nothing in my diff review caught it because I diffed my own change, not the file's resulting state. 9bdd91c restored those corrections from its own copy and, by the same mechanism in the opposite direction, dropped the module_sample truncation section eae0cca had added. So no commit on this branch has ever carried both. This one does, verified by marker rather than by reading the diff: the three corrections, the #10143 bounds, and the truncation finding are all present at once. The check that would have caught either revert is grepping the RESULT for every claim the file is supposed to carry -- a diff shows what you changed, never what you clobbered. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDKZ5pDN7cXTSEvYST2y78 * Retract the #10170 counterexample: this document's own discriminator refutes it I added gunbc#10170 as the one worked example of a genuine shared root, on 506/425/416/362 -> 180/174/172/168 with the spread collapsing 144ms -> 12ms. Two floor artifacts refute it, and the instrument that refutes it is the one this section argues for. pre (run 33707763185) cpu 314 305 290 318 steps 171091 166669 162507 171124 post (run 33711894017) cpu 339 327 317 344 steps 171052 166630 162468 171085 eval_steps moved 0.02% -- 39 steps out of 171091. The work did not change, so the relocation removed no evaluation work from these claims. CPU is HIGHER after, which is within noise, so the honest reading is no measured effect in either direction. The 506 baseline was a contended-run outlier: the rows were already at 314/305/290/318 with a spread of 28 before anything was touched, so both the drop and the spread collapse were artifacts of the baseline I compared against. And 180/174/172/168 came from a targeted claim_batch -- a fraction of the corpus and a different execution envelope. I asked the author to report from the artifact rather than the log, then accepted a number that came from neither, and committed it here. STEPS ARE DETERMINISTIC WHERE MILLISECONDS ARE NOT. A cost story that milliseconds support and steps refute is the clock talking. The same discriminator that says work is not ownership says here that a millisecond drop is not a work reduction. Records the measured lead in its place, found by the instrument that establishes ownership rather than by step clustering: the cross-claim demand census keyed by PRODUCER IDENTITY names target_project_arrow_body_to_value_expression at 95 claims / 117 evals, 4396ms total with 4350ms cross-claim across the emit family including all four rows. Admission still requires the controlled present-versus-absent pair, which has not been run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019LhF5WCbZqrZHPqsnjpkYu * The lead my retraction named is disqualified thirty lines below it Caught by lively-stag-270, against my own paragraph. My retraction of the #10170 counterexample named target_project_arrow_body_to_value_expression as "the measured lead" -- and this same document's fn-typed-parameter exclusion names that exact function as excluded, carrying handle_transform: fn(Node, Node, TargetModel) -> ... Reification refuses the argument (portable_args_from_ctx -> RefusedArgsNotPortable), so the key cannot be formed and the store refuses AT PUBLICATION. It is not a candidate whose serve might lose; it is a row that never stores. The controlled present-versus-absent pair my paragraph called for would measure noise against nothing, because the absent arm is the only arm. SAME SHAPE AS THE CONTRADICTION THE REVIEWER FOUND IN THIS FILE EARLIER: a screen that says it only excludes, followed by a sentence reading as an admission. Here it is a section that excludes a function, preceded by a paragraph offering it as the lead. A reader taking that as a worklist spends two CI runs on a row that cannot store -- the exact waste the screening section exists to prevent, and it would have been the sixth candidate withdrawn on this class. Keeps the measurement and the point about which instrument found it: the demand artifact keyed by producer identity reports 95 claims / 117 evals, 4396ms total, 4350ms cross-claim. Changes only the disposition, from a lead to a recorded fact about the serve mechanism's COVERAGE -- the largest cross-claim producer in this family is unreachable by the mechanism, which is not a gap in the census. Verified by asserting on the RESULT: all seven required markers present after the edit, per the discipline in stale_buffer_write_reverts_outside_its_own_diff. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019LhF5WCbZqrZHPqsnjpkYu --------- Co-authored-by: gunbc-ci-auto-heal <gunbc-ci-auto-heal@users.noreply.github.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
The coin flip, located
required-floorrefused three of the last fourmainruns withinterrupted_cpu_deadlinein {17, 15, 2} andverdict=FloorRefused, while run33673806092on the same corpus passed. The refused rows are the same seventeen identities in three modules, and the run log names themBUDGET-REFUSED in 503..506msagainstv2.workflow.required_floorrequired_floor_claim_cpu_safety_limit_ms(500):test/claim/self_host_compile_phase_live_gate_witnesstest/claim/self_host_compile_phase_frontier_witnesstest/claim/compiler_frontend_program_status_witnessOn the passing run those same rows completed at 452–496ms. Nothing about them changed between runs; they sit inside one scheduler quantum of the ceiling, so which side a PR lands on is ambient fleet load.
One root, not three
All seventeen rows spend their whole cost in
gunbc.self_host_compile_phase_frontiercompile_phase_frontier_standing. The first two modules call it directly; the third reaches it throughgunbc.compiler_frontend_program_statusmilestone_status, whoseSelfHostWholeCorpusPopulationDerivedarm asks the same question. So the repair is in the producer, not in any witness — andv2.workflow.required_floornames exactly this remedy ("reduce what the witness reaches for"; relocation is forbidden).What was fixed
Attributed with the instrument that already exists (
claim_batch --entry <witness> --functions <rows>, whose[witness]line reports per-row CPU), using probe modules that call one callee at a time. Three terms carried it, all shapes DESIGN §6 fixes on sight rather than on the realized n:newly_exposed_identity_admitted— a cheap identity guard conjoined ahead of an expensive cause derivation.&&in this substrate evaluates BOTH operands (v1_interpretereval_expr_innerevaluates left and right before applying the operator), so the guard narrowed the result without guarding the work: the derivation, two of whose arms scan the whole previous board, ran once per (added identity × disposition row) pair. Selecting the rows first is exactly equivalent —any(xs, p && q)=any(filter(xs, p), q)for boolean p, q.board_added_identities/board_removed_identities/phase_ratchet— set difference answered byanyover the other population per row, making the comparison count the product of two populations that grow with every receipt. Now one identity-keyed membership index per population.identity_is_hop_relocation— re-deriveddiagnostic_semantic_keyfor every row ofaddedandremovedon every call. Those keys depend on the transition, not on the identity under test, so they are derived once atapply_regression_transition— the least common ancestor of the demand (DESIGN §2).Result on this container: the worst row of each module falls from ~275–296ms to ~60–67ms.
Evidence
Every row of all three witness modules passes: 13/13, 35/35, 34/34 — including the frontier witness's discriminating reds over
newly_exposed_identity_admitted, and the live gate's planted-identity and equal-cardinality-swap refusals. Those are the probes that would go red if the restructure had changed the answer.No semantic change is claimed and none was found. The receipts' sealed identities, prefix digests and census digests are all derived, not literals, so an equal standing is an equal projection:
gunbc.design_ledgersrenders onlylatestand its output is unchanged. No consumer outside this module named any signature that moved, and the module has no stage0 Rust mirror, so no regeneration is implied.What this does not do
It does not bring the population to the ~1ms median an ordinary witness should hold. The rows remain above
required_floor_claim_cost_line_ms(100, diagnostic only) once the fleet's host factor is applied. The residue is the receipt-sealing data fold, paid once per claim — a different subject from the per-call scans repaired here.Worth knowing beyond this diff
&&and||are strict in the interpreter and short-circuiting in the emitted Rust. The two agree on the answer and differ in what they evaluate. Everycheap_guard && expensive_derivationwritten in the corpus by an author carrying the Rust habit over is paying the expensive half unconditionally. This diff repairs one instance; the class is not censused.🤖 Generated with Claude Code
https://claude.ai/code/session_01WEFMasGn2cAGCMJW2694pb