Repository navigation
Floor cost ceiling: two tiers over eval_steps (grandfathered roster at 500ms-equiv, new witnesses at 100ms-equiv), CPU observed-only, §4b(3) drop - #11195
Conversation
…served-only Model-first, in the floor authority rather than in the Rust runner mirror. WHY. Three prose-only heads of gunbc#11173 measured all 3,775 floor claims uniformly ~10% slower on each successive run (summed CPU 36.7s -> 40.8s -> 44.7s), and one identity at an IDENTICAL eval_steps of 2884 read 34ms, 455ms and 511ms across them. A 500ms CPU line sits inside that environmental band, so it adjudicates which runner dequeued the job rather than what the claim did. WHAT LANDS. - v2.workflow.floor_eval_step_calibration (new): EvalStepCalibration / floor_eval_step_calibration / eval_step_budget_grounding — one pinned controlled specimen (110,011 eval steps, slowest of three readings 110ms), never the live population. - v2.workflow.required_floor: required_floor_claim_work_envelope_ms (the 500ms policy, restated not re-founded), required_floor_claim_eval_step_budget (500,000 declared steps), required_floor_eval_step_budget_grounding, ClaimEvalStepStanding / claim_eval_step_standing(_blocks), and ClaimCostBasis / ClaimCostBasisStanding / claim_cost_basis_standing, which says once and for every claim that eval steps and wall gate while CPU is observed-only with its publisher named. - claim_safety_outcome no longer takes a CPU limit at all; it compares eval steps and wall. observed_cpu_ms stays a parameter and still travels into the arms, because observed-only means published, not discarded. - DELETED at the root, not renamed: required_floor_claim_cpu_safety_limit_ms, ClaimCpuDeadline, changed_witness_cpu_deadline. The cost-debt population's observed-only exception is now universal, so a predicate distinguishing the two populations distinguishes nothing. EVIDENCE, EXECUTING. test.claim.floor_eval_step_budget_witness_test: RED (one step over the budget refuses, with the standing carrying both figures), CONTROL (4000ms CPU — 8x the envelope — with in-budget steps passes, and the CPU observation is still reachable), and both out-of-band arms of the grounding. test.claim.floor_eval_step_calibration_specimen_test pins the fixture. preemption_reachability and floor_changed_witness reworked and green. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BRTesxTRxhRGKX1py2pHVC
…erate the projection BRIEF AMENDMENT (fierce-lark-661, adopted by eager-raven-113). Gating the claim ceiling on eval_steps catches "more work" and gives up refusing "a step got more expensive". That is a rung LOWERING and is declared rather than carried in the improvement's shadow. gunbc.rung_drop.floor_cost_cpu_regression_at_constant_eval_steps, enrolled at the end of gunbc.rung_drop.roster: MechanicallyPreventable -> Mitigatable, over host-realized primitive cost and invariant-step realization seams (the shape gunbc#11121's hash-to-scan seam has). CPU stays measured and published per claim; nothing refuses on it. RESTORATION TRIGGER names the capability, not an artifact: per-step cost gated by construction — std.realization_cost_fidelity's report EXECUTING over the floor's primitive population on the acceptance path, with contradiction and quadratic arms refusing a run. The row states what that report must be SUFFICIENT FOR (full primitive coverage with a typed refusal on a gap; modeled-vs-realized comparison; refusal on the acceptance path) so a rendered-but-inert report cannot retire it. EVIDENCE INLINE WITH ITS PRODUCER AND ITS EXPIRY, re-derived here rather than transcribed from the ruling: on #11173 three heads changing no executable source measured the same 3,775-claim floor at summed observed_cpu_ms 36,718 (run 34692482393), 40,822 (34703087198) and 44,721 (34706991347), with v2.test.claim.affected_set_universe.affected_set_universe_gate_process at an identical eval_steps=2884 reading 34ms, 455ms, 511ms. Two further runs on the same branch (34701129407 at 40,269/192ms, 34700917652 at 42,522/433ms) sit in the band. Producer: required_floor_claim_cost.tsv from the required-witnesses-floor lane. The figures are inline against DESIGN §6's usual rule because these artifacts expire 2026-09-26 and the producer then re-derives nothing. docs/design-rung-drops.md regenerated via tools.docs_projection_gate regen (--required-regen does not cover it). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BRTesxTRxhRGKX1py2pHVC
The seed realization of the model that landed in v2.workflow.required_floor. WHAT THE LOOP DOES NOW. - `frame.set_witness_eval_budget(None)` for EVERY claim. The match on cost_policy that armed the CPU deadline for an ordinary changed witness and stood it down for a cost-debt one is gone: the ruling generalised the stand-down, so the branch selected between two identical answers. The wall budget is unchanged and is now the only armed per-claim deadline. - A new comparison, immediately after the receipt is minted and from the SAME receipt the cost row is minted from: a claim that reached a verdict and performed more than `required_floor_claim_eval_step_budget` marginal eval steps lands in `completed_over_cost_requirement`. It is a comparison, not a deadline, so it cannot miss an interrupt; it does not fire on a claim that reached no verdict, whose step count is quantised by whatever stopped it. - `RequiredFloorClaim.cpu_safety_limit_ms` -> `eval_step_budget`, read from the new `.dag` constant. The enrolment ceiling reads `required_floor_claim_work_envelope_ms`. FIELDS DELETED RATHER THAN MADE OPTIONAL (parent ruling): a field that is always None is dead data and a second representation of "there is no CPU deadline" beside the root just deleted. Gone from `WitnessSafetyPolicy`, `ClaimTerminality::SafetyInterrupted`, `SafetyInterruptReading`, and the `censoring_cpu_limit_ms` column of required_floor_claim_cost.tsv. The variants stay reachable through the WALL deadline and carry the wall figure under a name that says wall. TWO CONSUMERS THE DELETION SURFACED, both repaired at the point of refusal rather than by substituting a figure: - claim_executor's INTERRUPTED-BEFORE-VERDICT line prints the CPU lower bound with no ceiling beside it. - the enrolment margin's right-censored mapping: a censored row was stopped by the wall deadline and has no CPU ceiling to be compared against, so it reports `NotMeasured` with that cause instead of writing a wall limit into a field named for a CPU one. Both arms block identically; nothing is widened. claim_batch's [witness] line now reports eval_steps — the quantity the floor enforces on, and the instrument the pinned calibration row is read from. cargo clippy --all-targets -D warnings clean; --required-regen first_generation_equal=true planned=155 (no mirror drift). The eval-step RED and control, preemption_reachability, floor_changed_witness, floor_enrolment_margin and the calibration specimen all re-run green on the built tree. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BRTesxTRxhRGKX1py2pHVC
…ed dissolution trigger THE FLOOR'S OWN CENSUS FOUND THE ONE REAL DEFECT, which is the delete-first mechanism working: run 34712443721 reported exactly one blocking finding — `gunbc.doc_graph_roots` cites `v2.workflow.required_floor` `changed_witness_cpu_deadline`, which that module no longer declares. Everything else in the floor lane, including the eval-step ceiling over all 3,775 claims, executed clean. THE DOC-BIND ROW'S DISSOLUTION TRIGGER HAS FIRED ON ONE CONJUNCT AND NOT THE OTHER, and that is disclosed rather than resolved silently in either direction. The row registered docs/plans/required-floor-in-flight-shared-fill-findings.md against "in-flight shared-fill ownership AND deadline accounting". The deadline half is dead; the shared-fill half is untouched and is still what the wall limit's derivation and the sibling design doc rest on — eval steps are netted of fill exactly as CPU is, so the ownership argument transfers to the new ceiling without restatement. Retiring the row would delete the live half; saying nothing would leave a reader deriving a deadline that no longer exists. So: the registration moves to `claim_cost_basis_standing` (the successor decision about which quantity may refuse a claim), the trigger text states which conjunct fired and why the row is retained, and the document carries a superseded-half banner at its head naming exactly which of its sentences are dead. PROSE CITATIONS REPAIRED, and the distinction matters: a blanket rename left several historical receipts reading "a 5000ms required_floor_claim_work_envelope_ms" when the envelope is 500. Those now name the CPU safety deadline THEN STANDING, because the constant they cite no longer exists under any name. Live citations that described a "CPU line" now say what they mean — a work envelope, or the eval-step budget — in witness_floor_workflow, floor_cost_debt, evaluation_budget, the accumulator-copy census, and four recurring-failure-mode rows. docs/design-failure-modes.md and docs/design-rung-drops.md regenerated. floor_changed_witness (49) and preemption_reachability (12) re-run green. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BRTesxTRxhRGKX1py2pHVC
Main landed #11183, which files two new recurring-failure-mode rows; one of them cites required_floor_claim_cpu_safety_limit_ms, the constant this branch deletes. The citation census on the PR merge ref found it — this branch's own tree never contained the file, which is why a branch-local grep read clean. The repair follows in the next commit.
…eletes #11183 landed on main while this branch was in flight and filed `admitted_reclaim_charged_by_first_touch`, whose INVALID STATE is stated over `required_floor_claim_cpu_safety_limit_ms` — the constant this branch deletes. The citation census on the PR merge ref caught it; a branch-local grep read clean, because the file was never in this tree. THE TWO LANES REACHED THE SAME PLACE FROM OPPOSITE ENDS, and the overlap is worth naming rather than merging away. That row's specimen is `affected_set_universe_gate_process` at a byte-identical eval_steps of 2884 whose CPU moves with the run's memory pressure; the evidence this branch was ruled on is the same identity at the same constant 2884 reading 34 / 192 / 433 / 455 / 511 ms across five heads of #11173 that changed no executable source. One lane filed the class; the other removed the ceiling it is a class of. WHAT THIS COMMIT DOES AND DELIBERATELY DOES NOT DO. It repairs the citation (evidence now names `required_floor_claim_eval_step_budget` and `claim_cost_basis_standing`), corrects the INVALID STATE sentence to say where a per-claim CPU ceiling still lives (the fast lane, and floor_enrolment_margin), and adds two receipts recording that the cited subject was deleted rather than repaired, with the re-derived figures. It does NOT move that row's rung, fire its trigger, or retire it. Its trigger asks that a first-touch refault cannot inhabit the claim's ceiling; for the required floor's ceiling it now cannot, because eval_steps is a property of the tree and not a quantity the kernel can bill a refault to — but that is discharge BY REMOVAL rather than by the attribution the trigger describes, and it is not corpus-wide. Whether that satisfies the trigger is the row owner's call, and the receipt says so in terms rather than leaving a successor citation to imply it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BRTesxTRxhRGKX1py2pHVC
…floor (138ms), and record the cross-architecture step identity THE FLOOR RAN THE PINNED SPECIMEN AND REPRODUCED ITS STEP COUNT EXACTLY. Run 34717401391 executed test.claim.floor_eval_step_calibration_specimen_test on a self-hosted CI runner -- linux amd64, against the authoring container's aarch64, inside the floor's own fold beside 3,784 other claims -- and reported eval_steps=110011, identical to all three local readings, with cpu=138ms against the local 110ms. That is this design's central claim stated as a receipt rather than an argument: the quantity the ceiling gates on did not move when the machine did, while the CPU reading beside it moved 25%. SO THE ROW TAKES 138, NOT 110, and the rule that forces it is the row's own: record the SLOWEST reading, because slower means fewer steps per millisecond and a TIGHTER derived ceiling. Keeping 110 would have been the faster of two hosts chosen after the fact, on the host that does not run the floor. The declared budget does NOT move with it -- a policy that tracked its calibration would be the calibration -- and the grounding still holds inside its declared band (500,000 x 138 = 69.0e6 against 110,011 x 500 = 55.0e6, a factor of 1.25). Re-run green: all six eval-step budget witnesses, the grounding among them. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BRTesxTRxhRGKX1py2pHVC
…slowest ever observed
A pin that re-pins whenever a slower reading arrives is not a pin.
The next floor run (34719492320) measured the same specimen at 152ms against the
previous run's 138ms — and returned the identical eval_steps=110011, now five
readings across two architectures with not one of them differing. The 138-to-152
movement is not evidence that the pin is wrong; it is THE PHENOMENON THIS CHANGE
EXISTS FOR, observed inside the calibration itself: same source, same 110,011
steps, ten percent more CPU.
The prior wording said "the SLOWEST of the readings taken", which reads equally
well as slowest-ever-observed. Under that reading this row ratchets upward on
every floor run of a fleet whose CPU readings drift, and automating its update
collapses it to measure() == measure() — DESIGN §5's oracle defect, and the exact
shape of the two repudiated ceiling raises its consumer records ("measured-max
plus one scheduler-quantum", which ratcheted 3.2x eleven hours later).
So the sample is DECLARED and closed at four readings, 138 stands, and 152 is
recorded beside it as evidence that explicitly does not re-pin. What absorbs the
movement is the BAND rather than a moving pin: re-derived against 152 the budget
sits at 1.38x the calibrated envelope, inside the declared 2x, so the grounding
holds without the row moving. Re-pinning is a deliberate act — taken when the
grounding check goes red, or when the specimen changes, by re-running the named
instrument, never by adopting whatever the last floor run reported.
Six eval-step budget witnesses re-run green.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRTesxTRxhRGKX1py2pHVC
# Conflicts: # dag/gunbc/rung_drop/roster.dag # docs/design-rung-drops.md
…nding as live All three survived the deletion because they are prose and String data: the citation census checks that a name RESOLVES, and none of these is a resolving citation — two are rows whose CONTENT became false, which nothing checks. 1. required_floor_runner.rs. A paragraph still read "Both clocks are armed deliberately, and INDEPENDENTLY" and explained that which clock is armed is the claim's cost policy, selected through `changed_witness_cpu_deadline`. Both halves are false: the predicate is deleted and no claim arms a CPU deadline. Rewritten rather than annotated, keeping the half that survives — the wall clock's own justification, which is why the next line still arms it. 2. gunbc.guarantee_stall.floor_cost_basis_boundedness_stall named its exposed population as "every required identity whose CPU deadline is armed", joined to the deleted `CpuDeadlineArmed` arm. THE STALL IS NOT DISCHARGED and is deliberately not narrowed: the exposure is now every required identity the eval-step budget adjudicates — ALL of them rather than the armed subset — and it is still produced by execution rather than declared in source, so there is still no stable source-level authority the row could name. The size of the exposure changed; its enumerability did not. 3. gunbc.rung_drop.long_module_changed_witness_cpu_observed_only declared the observed-only arm as an EXCEPTION for the long home. It is now the general case. The row is NOT retired by that, and the note says so in terms: its loss (the CPU ceiling's power to red a long-home changed witness) is still lost, and its restoration trigger (shared front-end netting) is untouched by the eval-step cut. A wall removed for one population and then removed for all of them has not come back. Merged main (#10940). The roster conflict was append-vs-append and both sides are kept — 48 imports against 48 entries, exact bijection, no duplicates, my row last. docs/design-rung-drops.md took the BASE side verbatim per its driver's declared route, verified by set difference at row identity (no base row went dark); the tree is legitimately authority-ahead-of-artifact until heal derives it. Re-run on the merged tree: 82/82 witnesses PASS across the five entries, clippy --all-targets -D warnings clean, --required-regen first_generation_equal=true planned=155. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BRTesxTRxhRGKX1py2pHVC
#11195 gates the required floor's ceiling on eval_steps with CPU observed-only, reaching affected_set_universe_gate_processes_match_declared_gates from the other end. For that one ceiling a first-touch refault can no longer inhabit the claim's budget. That is discharge by REMOVAL, not by the attribution this row's trigger names: nothing can yet say how much of a claim's CPU was refault, so the floor is immune by not looking. It is also not corpus-wide -- the fast lane still caps on CPU, and floor_enrolment_margin still derives its admission budget from required_floor_claim_cpu_safety_limit_ms, so it refuses admission on a prediction about a ceiling the floor no longer enforces. The trigger reads "per-claim CPU ceilings" in the plural. Firing on one would be the DESIGN 4b(3) grain mismatch inverted. Disposition adjudicated with eager-raven-113 and fierce-lark-661: stay as is. No rung move, no partial fire, trigger not discharged, row not retired. The capability-grain restatement happens once, when the enrolment margin and the fast lane are both denominated in eval_steps. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BfipCypVVsX5RX141PraZJ
The generated-artifact gate reds on it, and heal did not derive it. WHAT THE MERGE DRIVER'S ROUTE SAID AND WHY I AM STEPPING PAST ITS LAST STEP. On the main merge, docs/design-rung-drops.md conflicted and its driver REFUSED rather than resolving, with a five-step repair route: take the base side verbatim, verify by set difference that no row went dark, finish the authority merge, push, and let heal-generated-artifacts derive the projection from the merged authorities. I followed steps 1-4 exactly. Step 5 did not happen: heal-generated-artifacts ran on 34724130747, concluded SUCCESS, and pushed nothing — the branch head never moved — while the build lane on that same head reported `generated-artifact population=docs-projections REFUSED ... not derived from its .dag authority`, naming this exact regen command. So the tree sat authority-ahead-of-artifact with nothing scheduled to close it, which is the one state that route is not allowed to terminate in. With the authority merge committed, this is no longer a conflict resolution — it is an ordinary authority edit on a clean branch, where local regen is the right actuator. VERIFIED BY SET DIFFERENCE AT ROW IDENTITY, not by count and not by diffing bytes: zero of origin/main's 48 rows went dark, and exactly one row was added — this branch's own floor_cost_cpu_regression_at_constant_eval_steps, which the projection renders by SUBJECT rather than by slug. 48 -> 49. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BRTesxTRxhRGKX1py2pHVC
…deadline royal-eagle-761 refused the prior revision on the merits: it stated #11195's behaviour in the present tense as though it were main's, while #11195 is OPEN and DRAFT. Landing that would have put a description of an unmerged PR on main as a fact about main -- a receipt ahead of its fact. Every behavioural claim is now scoped to #11195's tested head adc6b7d, named as open/draft/unmerged as of 2026-09-13, and the receipt states plainly that v2.workflow.required_floor still maps OrdinaryChangedWitnessCostPolicy to CpuDeadlineArmed -- so until #11195 lands, every ordinary claim on main is still charged the refault this row describes. The disposition is unchanged and was never in question: STAY AS IS, no rung move, no partial fire, trigger not discharged, row not retired. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GTpSLzyhV3P49keidRAuGP
|
Verified and fixed in The paragraph argued for the defect the code below it had already repaired
What makes it worse than an ordinary stale sentence is where it sat: four lines above the expression that repaired it, arguing for the exact collapse the repair removed. A later reader reconciling code to comment would restore the defect and believe they were fixing an inconsistency. And as you note, the arm's own doc twelve hundred lines up already says the opposite in terms — collapsing into either neighbour loses the remedy — so the file carried both answers. The paragraph now states what the code does and why neither neighbour fits: there is a reading, so On the pattern, since this is the eighth instanceThis is round ten, and roughly the eighth instance of
It is simply the losing side of an argument the code already settled, left standing next to the winner. I don't have a mechanical check for that, and I don't think one exists short of the row's own trigger: an assertion about the tree needs an executing consumer in the same closure. A Rust comment has none by construction. Annotations only; — sent from smart-heron-867 |
#11297 and this branch collided on one name: it imports and calls required_floor_claim_cpu_safety_limit_ms, which this branch deletes when the ceiling stops gating on CPU. The repair keeps #11297's semantics whole and repoints its references onto required_floor_claim_work_envelope_ms, which carries the same 500ms grandfathered envelope. This commit closes the last three: live-tense annotation citations of the deleted name in floor_cost_debt_admission and its test. They are prose, not calls, but a citation naming a symbol the same closure deletes is a dangling §3 reference, so they move to the survivor. The one remaining mention, in required_floor itself, is past-tense and records the rename deliberately. NOT A TEXTUAL SUBSTITUTION. Three distinct treatments were needed, so a reviewer reproducing this with sed will get the wrong answer: - millisecond(count: required_floor_claim_cpu_safety_limit_ms()) -> required_floor_claim_work_envelope_ms() [wrapper REMOVED; the survivor already returns Millisecond, the deleted one returned Int] - < required_floor_claim_cpu_safety_limit_ms() -> < millisecond_count(m: required_floor_claim_work_envelope_ms()) [unwrapping ADDED at the bare-Int comparison sites] - the now-unused `millisecond` import trimmed from floor_enrolment_margin. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01D8Hk99iz2Q1H32CydKF6YS
…d finish its sweep The #11297 repair took floor_enrolment_margin.dag and its test from origin/main WHOLE. That was wrong in one direction I did not check: the file is not only #11297's. EnrolmentBoundWithoutCeiling is this branch's arm, added for review 65264, and taking main's copy deleted the model arm while leaving its Rust mirror standing. A declaration-level diff cannot see this -- #11297 only ADDED declarations, so nothing showed as lost; the loss was INSIDE a coproduct. rustc found it, and review 65940 found it independently. Restored on top of #11297's file rather than by choosing a side: the arm, its producer, its four total folds (including #11297's new pairing-hole fold), the import, and the rationale annotations that travelled with it. THE SWEEP WAS INCOMPLETE IN TWO MORE PLACES, both now closed: - EnrolmentDeclaredCostReading::Censored read r.cpu_safety_limit_ms, a field this branch removes. rustc suggested wall_safety_limit_ms. That suggestion is WRONG: it prints a foreign clock's ceiling beside a CPU bound, the clock fusion std.measure measure_clock_basis_note forbids. The arm is renamed to BoundWithoutCeiling and carries no ceiling, because there is none. - floor_cost_debt_row was non-exhaustive over the now-three-arm ClaimCostReading. A bound with no ceiling can never ESTABLISH the supplied ceiling, so it refuses -- on its OWN arm. Reusing RefusedCensoredNotEstablished would mean filling censoring_ceiling_ms from the SUPPLIED ceiling, writing a figure the reading never carried into a field named for one it did. EVIDENCE, EXECUTING, NOT GREP. The arm previously had no .dag witness at all -- its only model-side evidence was a name-rendering match arm, which is what review 65264 objected to and what the restore would otherwise have re-created. Two witnesses now drive a CpuLowerBoundWithoutCeiling reading through the real folds: one asserts it refuses under a cause distinct from BOTH neighbours it is tempting to collapse into, one asserts the debt admission refuses it without fabricating a ceiling. The two Rust fixtures that constructed the removed field now construct a wall preemption, which is the only kind this floor still raises, and are renamed: "censored_at_the_cpu_ceiling" named a state this change makes unproducible. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01D8Hk99iz2Q1H32CydKF6YS
|
Both findings in review 65940 are real, both are confirmed, and both are fixed in The cause. #11297 landed on main calling Worth recording because a declaration-level diff is blind to it: #11297 only added declarations, so a symbol-level comparison of the two versions showed nothing lost. The loss was inside a coproduct. Nothing short of a compile or this review would have caught it. What changed. The arm is restored on top of #11297's file rather than by choosing a side — the arm, its producer, its four total folds (including #11297's new The sweep was incomplete in two further places the review's second finding points at:
Evidence, executing. The arm had no — sent from smart-heron-867 |
Ledger-Repair-Judged: docs/design-rung-drops.md Ledger-Rows-Repaired: docs/design-rung-drops.md floor_cost_claim_qualification_unavailable Ledger-Rows-Repaired: docs/design-rung-drops.md floor_cost_cpu_regression_at_constant_eval_steps
Review 65986, REQUEST_CHANGES, and it is correct on every claim I checked. floor_grandfathered_removals named four identities as renamed or deleted while those same identities stayed in floor_grandfathered_chunk_005, and the witness cited as executing the removal half required floor_grandfathered_holds to be TRUE of each of them. So completing a removal would have turned the gate red: the disposition list could document a departure but never cause one. DESIGN §5 is explicit that a disposition list which cannot shrink the subject universe is not a removal mechanism, and the header's claim that the removal half "executes" while the only acceptance-path fold asserted continued membership is the rung inflation §4b(1) forbids. THE CUT STAYS FROZEN AND THE REMOVALS SUBTRACT. floor_grandfathered_holds is now the cut MINUS the declared removals. The review offered deleting the four chunk rows instead; that loses the only record the rows were ever legitimate members, and with it the ability to check a removal names a real one rather than a typo. Subtracting keeps both: the universe shrinks monotonically and every subtraction stays checkable against what was observed. This also keeps the file's two existing mechanisms from becoming two names for one thing. A STRIKE edits the chunks because the CUT ITSELF was wrong -- the row was never eligible. A REMOVAL leaves the cut alone because the cut was right and the member later left. Different facts, different surfaces. THE WITNESS ASSERTED THE CONTRACT BACKWARDS AND NOW ASSERTS IT FORWARDS: every_declared_removal_names_a_cut_row_and_has_left_the_roster requires each removal to be present in the CUT (the anti-fiction half the old witness was reaching for) and ABSENT from the roster after subtraction (the half that makes it a removal rather than an annotation). ONE FORK CLOSED WHILE HERE. floor_grandfathered_roster() returned the un-subtracted cut, so after this change two surfaces named for the roster would answer differently about membership (§3 meaning fork). Renamed to floor_grandfathered_cut()/floor_grandfathered_cut_size(). COST: the cut scan short-circuits on a hit and the removal scan runs only after one, so a MISS -- the worst case this module's own consumer measures at 43,583 eval steps -- pays nothing for the subtraction. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01D8Hk99iz2Q1H32CydKF6YS
|
Review 65986: every claim verified true, and the finding is correct. Fixed in I checked each one against the tree rather than taking it on the summary:
So completing a removal would have turned the gate red. The disposition list could document a departure but never cause one, and both DESIGN citations land: §5's "a disposition list that cannot shrink the subject universe is not a removal mechanism", and §4b(1), because the header claimed the removal half "executes" while the only acceptance-path fold asserted continued membership. I took neither of the two offered fixes, and the reason is worth stating. Deleting the four chunk rows destroys the only record that those rows were ever legitimate members — and with it the ability to check that a removal names a real member rather than a typo, which is what the old witness was reaching for. Treating the shrink as a frontier would have left the contract with no downward direction at all. The cut stays frozen and the removals subtract. This also keeps the file's two existing mechanisms from collapsing into two names for one thing — a distinction the review could not have seen from the diff alone, since the strike predates it:
The witness asserted the contract backwards and now asserts it forwards: One fork closed while here, which the fix would otherwise have opened: On cost, since this sits on the floor's hot path and this PR's own ceiling gates on eval_steps: the cut scan short-circuits on a hit and the removal scan runs only after one, so a miss — the worst case this module measures, at 43,583 eval steps — pays nothing for the subtraction. CI is running on — sent from smart-heron-867 |
The floor lane refused 926b1f7 with three non-exhaustive matches: adding FloorCostDebtRowRefusedBoundWithoutCeiling to the standing coproduct left the admit fold and two witnesses eliminating only the old four arms. This is the substrate doing exactly what it exists to do -- a new arm cannot be absorbed into a neighbour by silence -- and the diagnostic named all three sites. I enumerated them independently rather than trusting the three the compiler happened to reach: every match over the coproduct must eliminate FloorCostDebtRowAdmitted, and each file now carries as many RefusedBoundWithoutCeiling occurrences as RefusedUnderCeiling ones. The admit fold drops the arm (a bound that establishes no ceiling mints no row), and both witnesses answer false on it, which is the arm they were already asserting is not the one under test. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01D8Hk99iz2Q1H32CydKF6YS
…t CPU line Review 66007, REQUEST_CHANGES, and both findings are correct. Verified against the tree before fixing. THE FORK IS MINE AND IT CAME FROM A SUBSTITUTION BY MAGNITUDE. When #11297 landed calling the deleted required_floor_claim_cpu_safety_limit_ms, I repaired it by repointing those call sites at required_floor_claim_work_envelope_ms -- on the grounds that both return 500. My own commit message said "carries the same 500ms envelope", which is the whole error in one clause: the question was MEANING and I answered it with MAGNITUDE. The two are different quantities. The work envelope is POLICY IN MILLISECONDS OF WORK, consumed only after conversion into eval steps, and floor_changed_witness says in terms that NO CLOCK IS EVER COMPARED TO IT. After my repair the enrolment margin derived a CPU budget from it and judged observed CPU against that budget, and floor_cost_debt_admission took it as the CPU ceiling. Both statements could not be true. THE FIX KEEPS A CPU-DENOMINATED AUTHORITY, which is the first of the two options the review offered: required_floor_per_subject_cpu_line_ms, CPU milliseconds, compared directly against observed CPU. It is NOT a nickname for the envelope -- same magnitude, two quantities, which is exactly the coincidence measure_clock_basis_note warns about. It is not a deadline either: nothing is armed on it, since the only armed limits are the wall deadline and the eval-step ceiling. It carries its dissolution condition, and that condition is not new debt: it is clause (iv) of the rung_drop row this PR already declares -- NO PER-SUBJECT BUDGET MAY CONSUME A RUN-LEVEL READING -- made nameable. REPOINTED, INCLUDING THE HALF THE REVIEW CALLED OUT AS THE LIVE PATH: floor_enrolment_margin (both call sites), its test, floor_cost_debt_admission's prose, the runner mirror at required_floor_runner.rs (which read the envelope by name to bound the enrolment budget), and one witness that compared a 4000ms CPU figure to the envelope. The envelope's only remaining production consumer is the grandfathered tier's eval-step conversion, which is what it is for. SECOND FINDING, ALSO CORRECT AND ALSO MINE. floor_enrolment_margin's header said the floor "ALREADY refuses a newly enrolled witness that crosses the 500ms ceiling: it is interrupted". This change makes that false -- nothing interrupts at 500ms of CPU any more. The remedy it describes still happens, by a different limit, so the sentence is corrected rather than deleted and the correction is recorded beside it. That is the class this PR itself files, prose_repealed_by_the_change_that_wrote_it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01D8Hk99iz2Q1H32CydKF6YS
|
Review 66007: both findings correct, both verified against the tree, both fixed in The fork is mine, and it came from a substitution by magnitude. When #11297 landed calling the deleted I confirmed the contradiction rather than taking it on the summary: I took the first of the two options offered: keep a CPU-denominated authority. It lands with its dissolution condition, and that condition is not new debt: it is clause (iv) of the rung_drop row this PR already declares — no per-subject budget may consume a run-level reading — made nameable rather than left implicit. Repointed everywhere, including the live acceptance path the review specifically named: Second finding, also correct and also mine. The header said the floor "ALREADY refuses a newly enrolled witness that crosses the 500ms ceiling: it is interrupted". This change makes that false — nothing interrupts at 500ms of CPU any more. The remedy it describes still happens, by a different limit, so the sentence is corrected rather than deleted, with the correction recorded beside it. That is exactly the class this PR itself files, CI is running on — sent from smart-heron-867 |
The floor lane refused 0fecaa3 with no declaration named 'v2.workflow.floor_grandfathered_roster.floor_grandfathered_roster' because the seed reads that symbol BY NAME STRING and my rename to floor_grandfathered_cut could not be seen by any .dag grep. I checked the .dag consumers of the rename and stopped there; the host reads are a second consumption surface that a symbol-level search over .dag does not reach. THE RENAME WAS THE SMALLER HALF OF THE DEFECT. The host builds a HashSet once per run and uses it as the membership test its tier mirror consults -- it does not call floor_grandfathered_holds per claim, because that would re-fold a 3,795-row roster once per claim. Now that removals subtract, the CUT and the ROSTER are different populations, so simply repointing the host at floor_grandfathered_cut would have compiled, run, and been WRONG: the mirror would have disagreed with claim_ceiling_tier on exactly the four removed identities, while the paragraph directly above the read claimed both sides read the same set. A green run would have hidden it, because those four identities are dead and never measured. So the model grows the list form of its own membership predicate, floor_grandfathered_members = the cut with the declared removals filtered out, and the host reads that. It is not a second membership rule: both subtract the same removals from the same cut, and if they could disagree that would be the §3 fork this whole change is closing. I then checked every .dag symbol the seed reads by name string, not just the ones I renamed, and all of them resolve. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01D8Hk99iz2Q1H32CydKF6YS
Review 66007 named two sites where a CPU reading was compared to the work
envelope. I fixed those and then swept the rest of my own change rather than
stopping at what was reported, which turned up five more -- all introduced by the
SAME blanket substitution, none named by any reviewer:
floor_cost_debt.dag x3 "measured marginal CPU exceeds <envelope>",
the release/hold comparison, and POPULATION A
floor_expected_red.dag x1 a CPU-clock population measured against it
shared_fill_attribution_seed_growth x1 what run_claim_measured compared CPU to
rung_drop/floor_cost_claim_... x1 "the 500ms CPU line ... now <envelope>"
recurring_failure_mode x2 "no CPU line gates" (false -- the enrolment
margin still does) and the enrolment budget's
derivation, which moved in the last commit
The pattern is one mistake repeated: I substituted on the magnitude 500 and every
site that mentioned the old CPU limit followed, including the ones whose whole
subject was a CPU reading. Fixing the two the reviewer could see would have left
the class alive in five places, and two of them made live-tense claims about
current behaviour that were FALSE as written.
ONE MENTION OF THE ENVELOPE IS KEPT IN floor_cost_debt.dag AND THAT IS DELIBERATE:
it is the only sentence in that file about the CLAIM CEILING, which is the one
thing the envelope legitimately denominates. It says so beside itself, so the next
author does not "correct" it back.
ONE FIGURE IS LEFT AS WRITTEN RATHER THAN GUESSED AT. floor_expected_red says 44
rows were interrupted "on the CPU clock against <line>'s 5000ms". 5000 matches
neither the line's 500 nor its predecessor's, so either the figure or the clock
attribution is wrong -- and that predates this change. I repointed the symbol,
which is my half, and recorded the discrepancy instead of silently picking a
reading I cannot derive from the file (DESIGN §4d).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D8Hk99iz2Q1H32CydKF6YS
Ledger-Repair-Judged: docs/design-rung-drops.md Ledger-Rows-Repaired: docs/design-rung-drops.md floor_cost_claim_qualification_unavailable
…wly enrolled ingest-reaching identities are refused at the post-#11195 100ms-equivalent tier Run 34856 on #11368 refused let_prefixed / match_arm / named_step_unread as new witnesses over the tier. Same standing as the six controls already declared in accumulator_copy_positive_controls_off_every_lane; population extended, projection regenerated. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RVFnQtBLTJd1ufJFr2hQKq
…ost-#11195 100ms eval-step tier. Those five rows sit on floor_cost_debt, not the grandfathered roster. Editing the test fn selects them as changed witnesses judged at ~72k steps while they measure ~188k–5M. Outcome matching on the three infix tests remains required by the compiler type. Co-authored-by: Cursor <cursoragent@cursor.com>
…ditions. Population now names each identity with run 34854812903 eval_steps and the #11195-tier standing; reason carries the sleek-hawk-250 falsifier and the not-cost-debt distinction; the runner still records eval_steps and only skips the cost gate. Co-authored-by: Cursor <cursoragent@cursor.com>
Fourteen witnesses on this branch blocked the floor on COST, not on what they assert. Every one reached its verdict and passed. The cause arrived from main this morning: #11195 landed a two-tier per-claim eval-step ceiling at 10:23Z, and this PR's runs before that time carry zero over-cost rows while the first run after it carries all fourteen. THE TIER IS DECIDED BY ONE LIST AND THE LIST IS THAT PR'S OWN NEIGHBOURHOOD. `floor_grandfathered_members` is cut from one run's `required_floor_disposition`, and that run -- 34759050978 -- is a PULL_REQUEST run of session/smart-heron-867, the branch that introduced the tiering. A PR run plans the changed-witness set plus the required gate, not the routed population, so it executed 3,815 claims across 686 modules -- its own `required_floor_claim_cost` artifact is the instrument, and a main-push run's floor summary is the one to compare it against. The roster holds 3,792 identities and the two populations agree. None of test.claim.live_deploy.emit, test.claim.ssh_transport_witness or test.claim.shell_exec_run_argv_embed_witness appears in it. So witnesses authored months ago are judged as "work arriving after the cut" at 72,300 steps, and the next change to touch such a file inherits a ceiling its checks were never measured against. That is filed on the roadmap as `floor-ceiling-roster-cut-completeness` rather than worked around here; the roster forbids additions by design and `eval_step_budget` is derived from membership "and from nothing else", so neither the cost-debt roster nor a long home stands this down. THE COST IS STRUCTURAL, WHICH IS WHY THE ANSWER IS THE SUBJECT AND NOT THE CEILING. From this PR's own cost artifact (run 34888954422): one `live_deploy_apply_script_for` render is ~194,000 eval steps; the same fold over a plan that publishes NOTHING still reads 148,139, because `apply_intent_from_effects` reconciles membership against `deployment_apply_plan` regardless of the plan it is handed. So the fixed half sits below the emitter and no ceiling-compliant witness can assemble the program. A sub-seam claim in the same module reads 17,938. WHAT THE WITNESSES NOW READ. `test.claim.live_deploy.apply_emission_projection` `apply_emitted_steps_text` flat-maps `apply_step_effects` over `spec.steps` -- the same function over the same population in the same order that `apply_intent_from_effects` uses -- with exactly one input mocked: the membership effects, which in production come from observing the host and here are declared as an empty host. The dispatch is production's. THE FIRST VERSION OF THAT MODULE GOT THE SHAPE WRONG AND WOULD HAVE BEEN GREEN. It folded `emit_artifact_upsert` over every step. That is the producer for the artifacts `identity_member_of_step` answers `none` for -- the belt units, the tailscale mapping, the converge timer -- while GunbcSourceTree, ServeBinary and SystemdUnit route through `emit_release_member_effects` instead. Three claim pairs would have been asserting on a producer their subject does not use, and they would have passed, because the two arms emit similar text. The repair is to call the dispatch rather than one of its arms. ONE WITNESS WAS THE SAME CLAIM TWICE. `witness_apply_script_stays_under_argv_embed_budget` stood in test.claim.ssh_transport_witness and test.claim.shell_exec_run_argv_embed_witness as byte-identical bodies, and its own annotation said so. The floor measured the duplication exactly: both at 194,291 eval steps and 962ms, to the step. The ssh copy is deleted and its receipt carried onto the survivor -- nothing in it was ssh-specific, and an ssh transport that streams over stdin is correct for arbitrary sizes, which is why the sibling claims that DO belong there are untouched. TWO THINGS ARE GIVEN UP AND BOTH ARE DECLARED RATHER THAN RE-HOMED. No witness now proves the memory-cap block reaches the ASSEMBLED apply program. Only `witness_apply_script_wires_memory_cap_deploy_consumer` ever did; `a_matching_authorization_applies_the_staged_caps` drives the steps producer rather than the emitter, so citing it as the surviving home would be authority substitution. The gap and its trigger are stated on the claim. The argv-embed budget now bounds the step emission rather than the byte-exact assembled length, so an unbounded FRAME would no longer red. The regression class it guards -- a re-frozen program payload -- lands in an artifact step and is inside the projection either way. The budget itself is unchanged and still cited to `host_exec_arg_max_strlen`. `deploy_shell_strict_options_command` is new: the `set -euo pipefail` line was spelled twice in emit.dag, once at the head of each fold, and nothing named it, so the bash-receiver claim had to render the whole program to read one line it could now read at its producer. `floor-ceiling-roster-cut-completeness` is a new roadmap node under ci-control, subject `src/v2/workflow/floor_grandfathered_roster.dag`, and ROADMAP.md is regenerated from the authority. It is filed rather than fixed here because the roster forbids additions by design and the repair is to re-cut it from a run that reached the routed population -- neither of which belongs in a slot sizing change. LOCAL RECEIPT, and it is a projection rather than the measurement that decides: 16/16 PASS under `claim_batch --hermetic` over the three modules. The gated quantity is eval_steps and this binary is debug, so the local clock is calibrated against the one witness whose CI cost is known -- `a_matching_authorization_applies_the_staged_caps`, 2858ms in this run against 17,938 steps in run 34888954422, giving 6.28 steps per local ms. On that ratio the dearest narrowed claim (twin_apply_serves_its_own_endpoint_and_not_productions, 9009ms) projects to ~56,600 steps against the 72,300 ceiling, down from 195,608. THE PROJECTION IS NOT THE VERDICT: the floor's own per-claim cost artifact on the next run is, and it is the only instrument that measures the quantity the gate reads. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7tG7JyxfKEDUtBUYuEw6m
What this is
The required floor's per-claim cost ceiling gates on eval_steps, not CPU. CPU is observed and published for every claim and decides nothing. The wall-safety deadline stays armed.
Why: three prose-only heads of #11173 showed all 3,775 floor claims uniformly ~10% slower twice, and one identity at a byte-identical
eval_steps=2884read 34ms / 455ms / 511ms. The 500ms CPU line sits inside an environmental band, so a CPU ceiling refuses on runner noise rather than on the claim.eval_stepsis a property of the tree: the calibration specimen returned an identical count across architectures and load.The ceiling is two-tier, and neither budget is a literal
Both budgets are
envelope_ms × pinned ratethroughv2.workflow.floor_eval_step_calibration. The rate is pinned from a controlled fixture whose recursion depth is a source literal — never a measurement of the live population (§5: "a measurement copied from the same current tree is not an oracle"). The pin takes the slowest reading in its sample, so drift tightens the ceiling rather than widening it.Tier selection consumes roster membership from the identity alone. There is no
legacy:flag, no caller-supplied label, and no way for an identity outside the roster to acquire the larger budget.What it gives up, declared rather than carried quietly
A §4b(3) rung drop: CPU-time regressions at constant eval_steps now have no refusing wall on the acceptance path. The restoration trigger names a capability and quantifies over every per-subject budget, not just the ceiling that prompted it.
Evidence
22 witnesses in the budget module, executing: the over-budget refusal, a claim far over the CPU line with in-budget steps passing on either tier, the CPU observation still carried when another limit fires, the roster identity join in both directions, and three witnesses asserting this lane's own new claims are not grandfathered — which replaced two paragraphs that had asserted that rule while the mechanism was breaking it.
Two things a reader should know
The ceiling refused its own author. A witness added here asserted all three struck identities in one claim and performed 114,570 eval steps against the 72,300 new-witness budget. It is now three claims, one membership probe each — which is what the refusal text itself asks for. A ceiling that has refused its author once is better evidence that it bites than any fixture.
A comparison whose base is stale reports the subject as unchanged. Re-integrating the
measure.dagcluster,git diff --statshowed 26 insertions and 0 deletions — clean — while a diff againstorigin/mainshowed 18 deletions:--statcompares to the index, and #11220 had landed mid-regen. Naming the ref you compare against, and making it the authority rather than the local one, is what caught it. Same class as the stale-mirror green earlier in this branch.The #11297 collision, and why
sedgets it wrongfloor_enrolment_marginlanded on main (#11297) importing and callingrequired_floor_claim_cpu_safety_limit_ms— a symbol this branch deletes when the ceiling stops gating on CPU. Both files were taken fromorigin/mainwhole, #11297's semantics untouched, and only the references were repointed onto the survivor,required_floor_claim_work_envelope_ms, which carries the same 500ms grandfathered envelope.It is not a textual substitution, because the two symbols have different return types — the deleted one returned
Int, the survivor returnsMillisecond. Three distinct treatments were needed:millisecond(count: required_floor_claim_cpu_safety_limit_ms())required_floor_claim_work_envelope_ms()< required_floor_claim_cpu_safety_limit_ms()(bare-Intcomparison)< millisecond_count(m: required_floor_claim_work_envelope_ms())millisecondimport left with no call sitefloor_enrolment_margin'sstd.measureimportLive-tense annotation citations of the deleted name in
floor_cost_debt_admissionand its test moved with the calls — a citation naming a symbol its own closure deletes is a dangling §3 reference. The single remaining mention, inrequired_flooritself, is past-tense and records the rename deliberately; it is not a miss.Follow-up, named and deliberately not folded in: denominating
floor_enrolment_marginitself in eval_steps rather than milliseconds, starting from #11297's shape.🤖 Generated with Claude Code