Repository navigation
Restoration trigger for floor_cost_contention_verdict: a deterministic evaluator-step work measure invariant across execution envelopes - #10105
gunbai-bot[bot] wants to merge 11 commits into
Conversation
…lumn in the floor's cost receipt that is not a clock `gunbc.rung_drop` `floor_cost_contention_verdict` names three arms that would restore a claim-owned cost basis, and one of them is "a deterministic work measure such as evaluator steps". This lands that measure as an INSTRUMENT. It does not land it as a BASIS and does not retire the row. `v1.interpreter` now counts one evaluator step per `eval_expr` entry, unconditionally — not under `GUNBC_INTERP_PROFILE`, because a measure available only in an instrumented envelope is not available in the envelopes the row is about, and not through the per-variant `EVAL_COUNTS` array beside it, which pays two `Instant::now()` calls per node. `run_claim_measured` takes the per-claim delta and nets stored shared-artifact fills out of it by exactly the rule the CPU clock is netted by, so the figure is not a function of execution order either. It reaches `PerformanceReceipt.eval_steps`, the `[floor-shared-fill]` and `[over-cost]` lines, and an `eval_steps` column in `required_floor_claim_cost.tsv`. The evidence is executed and discriminating — `evaluator_step_work_measure_tests` in `v1.interpreter`, green on the remote runner: - EXACT equality of the count across two genuinely different envelopes: one arm with the CPU deadline ARMED (a different path through `eval_expr`, taking the stride poll and two clock reads the other arm never executes) under a co-tenant thread spinning for the whole evaluation. Not a tolerance — a work measure that needed one would be a slow clock. - A work control at a different fixture size, so a counter frozen at any constant including zero fails the suite rather than passing every invariance assertion vacuously. - A netting arm where the claim that PAYS a shared fill and the claim that reads it warm carry the SAME marginal count, while their RAW counts are asserted to differ by more than a factor of ten — so the netted equality is not two identical numbers compared. Nothing compares the column against a line, deliberately: a column no verdict reads is a measurement and not a basis, and calling the row retired on the strength of a published column would be the rung inflation §4b(1) forbids. The row is updated to say what landed, and to name the two things still missing — a step-denominated line (unsizable until a run publishes the distribution this column now makes publishable) and the cross-envelope A/B on the shared runner at corpus grain, an identity join of `eval_steps` across two attempts of one identical tree where the cpu column moves and this one must not. Until that second one is measured the invariance claim is grounded at FIXTURE grain and nowhere wider. The CPU deadline is unchanged: still the armed enforcement clock, still cpu-ms, and this column changes no threshold and no verdict. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JaugFkN1vzZmVH6efyrZHR
The generated-artifact merge driver refused docs/design-ledgers.md: GeneratedArtifactConcurrentDivergence. Both sides changed the projection since the merge base — main added a recurring_failure_mode row, this branch edited a rung_drop row — so neither side's bytes project the merged authorities. THIS COMMIT CARRIES THE OURS SIDE, WHICH IS WRONG BYTES, AND SAYS SO. It exists only to give the regenerator a well-formed tree to run against; the next commit replaces these bytes with the regenerator's output. Nothing is pushed until that has happened. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JaugFkN1vzZmVH6efyrZHR
The previous commit deliberately carried the OURS side of the projection so
the regenerator had a well-formed tree to run against, and said so. This
replaces those bytes with the regenerator's output:
gunbc run --source-root dag --source-root src/v2 \
--entry dag/gunbc/instruments/generated_artifact_gate.dag --function main_wet
run from a gunbc built from this tree. NOT resolved by hand and not by
picking a side -- the driver's refusal exists precisely to stop that.
THE INTERMEDIATE COMMIT REACHED ORIGIN AND WENT RED, which is worth
recording rather than quietly fixing: session branches are pushed for me, so
"local until I finish" was never available. required-witnesses-build failed
on 745edf5 naming exactly this projection. That red was correct and it was
mine. It also establishes the gate discriminates, which is what makes its
earlier PASS on this branch's hand-applied projection worth anything.
The regenerated delta is main's liveness_probe_read_as_currency
recurring_failure_mode row. The rung_drop paragraph this branch edits is
untouched by the regeneration -- the prediction recorded before the run, and
the evidence that the hand-applied bytes were byte-correct.
REMOTE REGENERATION IS NOT AVAILABLE FOR THIS ACTUATOR: gunbc refuses on the
BuildBuddy runner with HostBudgetUnreadable -- no cgroup memory limit binds
that process, and it will not plan against a machine-wide reading. That is
the fail-closed arm working, not a breakage; builds stay remote and this
actuator runs locally.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JaugFkN1vzZmVH6efyrZHR
…d not cover it Review 58626 (codex/gpt-5.6-sol) blocked on the new evaluator instrumentation carrying no migration disposition. The narrow half of that is correct and is a defect I introduced. WITNESS_COST_CLOCK_BASIS_NOTE governs PerformanceReceipt and said the two u128 fields "are the ONLY reason this note exists". That stopped being true when eval_steps landed in the same struct, and the consequence is not cosmetic: the note's DISSOLVE-ON is ClockBasis replacing "these two bare fields". eval_steps is a COUNT, not a duration -- it has no clock basis, ClockBasis says nothing about it, and firing that trigger would delete the only disposition the field has while the field survives. A trigger discharging more than it covers is the same shape DESIGN 4b(3) names when it insists a trigger name the capability it retires and nothing wider. So eval_steps now carries its own disposition, with a trigger a reader can evaluate rather than a judgment about how far v2 has got: the roadmap row v1-zero-hand-maintained-rust, whose acceptance condition is that no hand-maintained Rust remains in the seed. The counter counts THIS evaluator's steps, so it lives wherever that evaluator lives. Earlier removal is permitted and expected -- v2 projecting the work measure from a modeled receipt subsumes it -- but that is not the trigger, because a scaffold may always dissolve early and a trigger fixes the point by which it MUST be gone. The evidence does NOT dissolve with it, and the row says so: per DESIGN 4b(4) a climb deletes lower-rung PRODUCTION handling and never the evidence, so evaluator_step_work_measure_tests stays enrolled as the executing proof that the measure is invariant wherever it comes to be computed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JaugFkN1vzZmVH6efyrZHR
…host_compile_phase fold The floor lane refused this branch's previous head with 15 CPU-ceiling preemptions, all in self_host_compile_phase_frontier_witness and self_host_compile_phase_live_gate_witness. main now carries #10038, three cost-shape repairs in that same module family's fold, plus #10031, #10025 and #10043. Integrating is the REAL change rather than a re-roll: the next floor run measures a materially different tree, so it is not another sample of the run that refused. Re-running the same tree until it answers is retry-until-green, which gunbc.rung_drop floor_cost_contention_verdict names as fail-open wearing a fail-closed label. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JaugFkN1vzZmVH6efyrZHR
The driver refused the projection again: this branch adds a rung_drop row and main's ledger moved, so both sides changed the same generated file. That collision is structural rather than unlucky -- two branches adding rung_drop rows cannot both be driver-clean at once, and whichever lands second pays one regeneration. This branch is paying it. REGENERATED BEFORE THE MERGE WAS COMMITTED this time, which is the ordering the previous round got wrong and pushed a red for: gunbc built FROM the merged tree, then gunbc run --entry dag/gunbc/instruments/generated_artifact_gate.dag --function main_wet, then git add. No intermediate commit carrying ours-side bytes ever existed, so no window in which the drift gate could correctly refuse this branch. The regenerated ledger carries BOTH authorities: main's recurring_failure_mode row and this branch's rung_drop paragraph. Worktree is empty after the regen, so nothing else diverged. Also carries main's changes to cli_run.rs and required_floor_runner.rs (#10034's interrupted-bucket mechanism), auto-merged; the eval_steps plumbing and the TSV column survive intact. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JaugFkN1vzZmVH6efyrZHR
Triggered by bright-ram-778 on a merge window. Same ordering as last time: gunbc built FROM the merged tree, then the generated_artifact_gate regen, then git add and one commit -- no intermediate commit carrying ours-side bytes. The invalidating merge was #10007 (a0f03e4), which added a rung_drop row; the other main-side merges since e49ad23 touched neither projection nor its authority. A regen is invalidated by a merge touching the SAME projection or its authority, not by the tip moving. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JaugFkN1vzZmVH6efyrZHR
Fourth regen cycle. Same ordering: gunbc built FROM the merged tree, then the generated_artifact_gate regen, then git add and one commit. The invalidating merge was #10033 (f6d872e), a recurring_failure_mode roster append. Verified the regenerated projection carries both authorities by joining rather than by spot-grep: all 15 rung_drop subjects present, and 54 authority failure-mode rows against 54 ledger bullets with all 54 joining on their authored prose. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JaugFkN1vzZmVH6efyrZHR
…rity text Ruling via bright-ram-778. The instrument is unchanged; what was refused is the declared row's text, and it is corrected here rather than after the merge, because the row is what future readers consume. (1) THE CAUSE WAS STALE. `floor_cost_contention_verdict` asserted contention as the mechanism, and sentences like "may have done nothing but run beside somebody else's build" asserted it again -- an attribution this lane has explicitly refused to make, having killed three candidate causes with evidence. Renamed to `floor_cost_claim_qualification_unavailable`, and renamed rather than reworded because a row identity carrying a refuted attribution gets cited onward as if the attribution were the finding. The reason now states only the measured composition: a CLOSURE-LEVEL component insensitive to the claim's assertion work, and an EXECUTION-POSITION-SENSITIVE component whose cause and bound are not established. No contention, memory pressure or warm-up is named, and "no bound established" is stated as distinct from "unbounded cause". (2) ONE SENTENCE WAS BROADER THAN ITS EVIDENCE. The row claimed the netting made the figure "not a function of EXECUTION ORDER". What was demonstrated is narrower: the net count is not determined by WHICH TESTED CLAIM PAYS THE MODELED SHARED-ARTIFACT FILL -- one modeled path, not independence from arbitrary corpus order. The broad sentence also contradicted the row's own missing-item (b) two sentences later. Narrowed to what was measured. (3) THE RESTORATION TRIGGER WAS DISJUNCTIVE AND IS NOW CONJUNCTIVE. Three alternative arms are refuted by the measured composition: isolation can stabilise the wrong subject, a deterministic measure can count the wrong subject exactly, and calibration can normalise a wrongly allocated charge. All three must hold -- charge subject aligned, basis invariant-or-bounded across position and envelope by exact identity joins, and the policy line grounded over the independently defined full population and consumed at the same subject grain it was derived at. Also records the two controls that would discharge the first two conjuncts (an order-rotation position control, and a same-closure charge-subject control), including that the larger-fixture size control proves the counter is ALIVE and does not prove the steps belong to the claim rather than its closure -- which this row previously leaned on as if it did. The rung is NOT retired and the row STAYS STANDING. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JaugFkN1vzZmVH6efyrZHR
…e, and merge main's receipt into it deep-badger-41 found two citations of the old symbol in dag/gunbc/recurring_failure_mode.dag (non_verdict_disposition_surfaces_as_refusal and recurrence_ledger_scoped_below_the_recurrence), plus main added one in v1_interpreter.rs. Left alone they would point at an identity that no longer exists, and the failure is SILENT AT THE CITING END -- nothing refuses, the prose just stops naming anything. PROVENANCE, because it changes what can recur: these were NOT missed by the original sweep. They were absent at a3a5726 and landed on main afterwards in cfb448d. A rename is a moving target while the branch is open, so any main commit citing the old symbol before this lands re-creates the defect. MAIN ALSO APPENDED ~6.5KB TO THE ROW ITSELF -- a dated 2026-09-02 receipt making the reroll mitigation visible -- and the merge left rung_drop.dag CONFLICTED. That content is preserved: main's receipt is spliced into the renamed row rather than dropped, which is what taking either side would have done. THE GATE CAUGHT AN ERROR I INTRODUCED WHILE DOING THIS, and it is worth recording. My first edit ran over an unresolved conflict, so the file still carried <<<<<<< markers and my "duplicate row" was really the two sides of the conflict block. The regen REFUSED with "unparseable .dag source: expected expression, found Lt" and wrote nothing. I noticed because the ledger came back carrying my corrections but NOT main's receipt -- an inconsistency between two things that must move together. A regen that had silently succeeded here would have shipped a projection missing one side's authority. Verified after the repair: exactly one floor_cost row, the roster references it, the row carries both main's receipt and this branch's corrections, and both projections carry both. The only remaining occurrences of the old spelling are the row's own deliberate record of its former name, kept so the old identity stays findable. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JaugFkN1vzZmVH6efyrZHR
Merge commit, not a rebase -- this branch touches four generated projections and the merge driver is the wrong thing to fight. THE ROSTER CONFLICT IS THE FINDING. main added five floor_cut_* rows to rung_drop_roster beside the OLD spelling of the row this branch renames. Both inputs are individually closed: main's roster is 21-for-21 with floor_cost_contention_verdict correctly declared, and this branch's roster is closed under the new name. The dangling reference exists ONLY IN THE MERGE RESULT -- a textual merge produced a roster line naming a symbol with zero declarations, and no author wrote that line. So the honest statement is not "someone cited a stale symbol" but REFERENTIAL CLOSURE IS NOT PRESERVED UNDER MERGE. A closure check run on either input passes. Only a check on the result can see it. Resolved as the union: main's five new rows plus the renamed identity in the sixth position. Taking either side alone loses content -- five rows dropped, or a dangling name kept. VERIFIED, both directions: 21 roster entries against 21 declarations, nothing rostered-undeclared and nothing declared-unrostered. The second direction is the one that catches a silently dropped row. DESIGN.md and docs/design-ledgers.md regenerated, not hand-resolved, from a gunbc REBUILT FROM THE MERGED TREE. The binary on hand predated the 04_infer changes this merge brought in; regenerating with it would have run a compiler older than the tree it emits, which is the trap the recipe's own step 3 names. Coupled-pair check after regen -- both sides' authority content survives into both projections: the renamed subject, main's five floor_cut rows, and main's dated 2026-09-02 receipt. Presence greps alone cannot see a dropped side; a pair that must move together and did not is what can. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JaugFkN1vzZmVH6efyrZHR
|
Closing: this duplicates #10030, which merged at 18:30:49Z as The auto-opener created this at 18:40 against the same branch. It looks like unlanded work because this repo squash-merges — the branch's commit SHAs never enter main's history, so Verified by content rather than by ancestry:
The remaining difference is the other direction: this branch is 4 commits behind main and lacks ~2103 lines that landed after the merge (the live-deploy Git-native convergence work among them). Every file where this branch appears to "add" content is a file where it simply holds the older version. So there is nothing here to finish and nothing to flip to ready. The follow-up work — a modeled entry point replacing the prototype join script — will open as a fresh branch off post-merge main, not from here. — sent from gentle-wolf-793 |
Auto-opened by session-dashboard for session
gentle-wolf-793.Pushing to
session/gentle-wolf-793advances this PR.Worker attestation
Before flipping this PR to ready for review, confirm each item:
npm test,cargo test) and the result.Closes #Ndirective.Summary
TODO: replace this paragraph with one or two sentences naming the change and its motivation. Reviewers read this first.
Test plan