Repository navigation
The 500ms per-witness CPU budget: one witness family preempts across unrelated diffs — four lanes taxed in a day. Establish whether the budget or the family is wrong, at class grain - #10076
Closed
gunbai-bot[bot] wants to merge 11 commits into
Conversation
…e the floor a cross-claim demand census The 500ms per-claim CPU ceiling has been converting unrelated diffs into merge blocks -- four lanes in one day, every run reporting failed=0. Neither the budget nor the witnesses are wrong. MEASURED (main run 33615659836, required_floor_claim_cost.tsv, 3477 rows, cost_basis=cpu). The corpus p50 is 2ms and 3379 rows sit under 280ms, so the ceiling is not mis-set against the corpus. And the tail family's own work is among the smallest in it: a witness and its DISCRIMINATING RED do materially different work, and they cost 296/294, 333/331 and 384/382 -- within 2ms of each other while both sit against the ceiling. A cost that does not move when the assertion changes is not the assertion's cost. One 34-row module spans 0ms to the ceiling in three bands that track WHICH SHARED PRODUCER a row forces rather than what it asserts. The charge is a per-CLOSURE constant: the floor builds a fresh evaluation frame per claim, so every claim re-derives the pure substrate its import closure reaches, and that constant is billed as the claim's own marginal work. The run's [floor-shared-fill] ledger nets nothing for any of the four families. THE FINDING IS THE BLIND INSTRUMENT. v2.workflow.floor_pure_producer_share already exists to stop exactly this recompute, and its roster is six hand-authored rows because nothing produces its candidates: the interpreter's recompute-trace ledger ranks pure calls with count>=2 WITHIN one InterpContext and is printed and dropped at every claim frame exit, so a producer evaluated exactly once per claim, in 3477 claims, carries count=1 in every ledger and appears in none. The instrument that ranks redundant recompute is structurally blind to the population its own repair roster is enrolled from -- which is why roster coverage is discovered when a budget refusal lands on someone else's PR. WHAT THIS LANDS. absorb_claim_recompute_demand folds each claim's ledger into a run-scoped census keyed on declaration identity plus argument row, before the frame dies; the floor prints the ranked head and writes required_floor_cross_claim_demand.tsv, uploaded as an artifact. Every truncation is disclosed (retention floor, key cap, bounded module sample beside an exact module count), and single-claim rows are retained so the shared population has a control. IT ENROLS NOTHING AND GATES NOTHING. A row is a candidate whose SERVE cost is unmeasured; floor_pure_producer_share records the case where the serve lost to the recompute and two enrolled rows were removed. No budget change, no quarantine, no per-lane accommodation -- the limit is correctly placed against the corpus, so raising it would move a line that is right to accommodate a charge that is wrong. Class filed as recurrence_ledger_scoped_below_the_recurrence. Seed growth justified at gunbc.cross_claim_demand_census_seed_growth (precedent gunbc#9721). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Fjt9vSfVCKwZu9x7d44xZP
…laims would double-fold Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Fjt9vSfVCKwZu9x7d44xZP
Single-claim rows are RETAINED and rank at zero -- they are the control that makes claims>1 mean something. The docstring said they were dropped, which points a reader at the wrong side of the census's one load-bearing claim. The code, the writer and the tests were always right; only the sentence was wrong. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Fjt9vSfVCKwZu9x7d44xZP
REVIEW 58673 (REQUEST_CHANGES), both findings real and both fixed by construction rather than by validation: THE ARGUMENT ROW WAS A 64-BIT HASH. Two distinct argument rows could collide and merge into one row reporting cross-claim demand that never happened -- a fabricated identity, forbidden outright by DESIGN §5, and contradicting this file's own promise of "sound argument identity" ten lines up. The key now carries the canonical argument vector the ledger already holds, so the collision is unrepresentable rather than unlikely. `keyed` and `unkeyed` are now discriminated in the key too: a nullary keyed call and a declaration's composite-argument bucket both carry an empty argument vector, and merging them would sum one identity's cost with all of another's. THE MODULE COUNT WAS COUNTING CLAIMS. The bounded sample and the counter shared a container, so once the eight-name sample filled, every later claim from an unsampled module incremented the count again -- a column promising distinct consumer modules while reporting claim occurrences. The complete set is now kept (interned, one allocation per module name for the whole census) and the count is its length; the cap bounds only how many names a row shows. RULINGS APPLIED. MEANING BELONGS IN .dag. Ranking is a judgment about which demands matter, so the artifact now leaves in deterministic IDENTITY order and the runner's log preview sorts a copy and says in band that it is a preview and not a candidate roster. Grouping stays where the ledger's own key is; carrying facts across a frame boundary only the seed can reach is the seed's warrant. THE INSTRUMENT MEASURES ITSELF. absorb_ms and absorb_max_ms are reported and written: the absorb runs after each claim's measurement returns, so it is outside every charged window and cannot trip the deadline, and that claim is no longer asked to stand on reasoning alone. Also removed a quadratic in it: declaration sites are resolved once per fn pointer per frame instead of by a linear scan per ledger key -- the cost-shape defect §6 always fixes, inside an instrument whose subject is cost. THE DROP IS SPLIT, NOT RETIRED. The class row now says a cross-unit census plus the modeled fold retires blind candidate discovery, incident-as- discovery and the missing carrier -- and NOT environment-independent claim-cost qualification, which no amount of producer naming makes invariant across execution envelopes. Claiming the whole drop would have been the rung inflation this ledger exists to catch. THE NEGATIVE CONTROL IS EXECUTED. The two rust target models were removed from the share roster on a measured serve-versus-recompute experiment, and the first artifact ranks that chain in the top ten while it stays unenrolled and unenrollable from here. ALSO CARRIED, from four lanes' observations: the census explains the LEVEL and not the VARIANCE (a per-closure constant is identical on two runs of one tree, while byte-identical evaluator steps have been measured against 1.31-1.80x cpu), the two compose into the wandering victim, and a cost inside a native builtin called directly from a claim body is not a row here at all -- the same boundary std.evaluation_budget's opaque-host-call note draws for the deadline. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Fjt9vSfVCKwZu9x7d44xZP
Asked to denominate the displaced cost as one number, I computed it and it was wrong by construction: summing cross-claim over the shared rows of the first artifact gives ~1850s against a run whose ENTIRE claim-side CPU was 130s. A 14x overcount, and it is nesting -- the ledger times a producer's whole subtree, so a producer and its callees both appear and overlap. A per-row figure is a valid statement about that producer; no sum of rows is a valid statement about the run. The artifact now carries claim_cpu_total_ms as the ceiling any true total must sit under, says cost_columns=inclusive_of_ callees_do_not_sum in its summary, and names what a real displaceable-cost figure needs: self-time, which this ledger does not carry. The artifact invited the error, so the refusal belongs in the artifact rather than in a reviewer's memory. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Fjt9vSfVCKwZu9x7d44xZP
… underivable rather than unstated The sentence 'this floor recomputes N seconds of pure producer work per run' is what makes the stake legible, and no arithmetic over this artifact produces it: the quantity it needs is not in the column. That is a next-rung trigger, not a caveat. The capability is SELF TIME -- inclusive minus the callees the same pass already counted -- after which the column is additive and the sentence is derivable. It is a TRACKED stall rather than a wish: CrossClaimFillGuard's Drop already computes exactly that netting against the CROSS_CLAIM_FILL_FRAMES child stack, one tier over, for the shared-fill ledger. What is missing is a child stack over the recompute ledger's frames. Deliberately not built here: this bridge is what stops the next lanes paying, and a new measurement tier would put it behind a fresh review cycle on the very fleet condition it exists to explain. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Fjt9vSfVCKwZu9x7d44xZP
… is single-run calm-boar-314 retracted the two-run reading this prose leaned on: a rerun replaces a job's ATTEMPT while the logs API serves the latest attempt for a fixed job id, so the six-row and thirteen-row readings addressed different attempts through one identifier. No error is raised and nothing is visible from the citing end. In the reproducible reading the two rows said to have dropped out are in the set. Replaced with evidence that needs no cross-run delta: two rows of one module measured 502ms and EXACTLY 500ms against the 500ms budget. A row sitting at its budget lands in interrupted-before-verdict or completed-over-cost according to whether the poll fired before or after the work finished -- a fact about the poll, not about the row. Same conclusion, one run, nothing to retract. The class row now carries the retraction as part of the class, because it sharpens the recommendation: a remedy must never be keyed to WHICH rows tipped. A producer census is invariant under that retraction; a ranking of tipped rows would have been built on it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Fjt9vSfVCKwZu9x7d44xZP
…d independently
calm-boar-314 over-retracted and I cut a true finding on it. Attempt logs ARE
recoverable -- /actions/runs/<id>/attempts/<n>/logs -- and I pulled both
attempts of run 33620893203 myself rather than restoring on a second say-so:
attempt 1: interrupted_before_verdict=6, and
..an_empty_receipt_series_leaves_the_live_tree_unmeasured_rather_than_held
is INTERRUPTED-BEFORE-VERDICT at 502ms
attempt 2: interrupted_before_verdict=13, and the SAME row is
COMPLETED-OVER-COST-REQUIREMENT at cpu_ms=500 -- it reached its verdict
One row, one tree, crossing the two populations the 2026-08-19 cut separated
BECAUSE THEY ARE DIFFERENT FACTS WITH DIFFERENT REMEDIES. Which one it lands in
is decided by whether the poll fired before or after the work finished.
STILL WITHDRAWN: near-disjointness (attempt 2's set largely CONTAINS attempt
1's) and the 3-6-13 trend (the 3 was a different head). Only the migration
returns.
CITE THE ATTEMPT, NOT THE RUN, now recorded as part of the class: a bare run id
and a bare job id both resolve to the latest attempt, so a rerun changes what a
stable citation serves with no error and nothing visible from the citing end --
which is why the retraction looked sound. My own figures now carry
run_attempt=1 for runs 33615659836 and 33622954427, checked rather than assumed.
The census reported the same thing through the retraction AND the restoration,
because it keys producers and never which rows tipped.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fjt9vSfVCKwZu9x7d44xZP
I wrote the restored migration as 'INTERRUPTED-BEFORE-VERDICT at 502ms'. The attempt-1 diagnostic says the opposite in three clauses: cost=UNMEASURED, above 500ms with NO UPPER BOUND, and interrupt_point names where the POLL observed the ceiling -- a property of the budget, not of the row. The pair is a BOUND beside a VALUE, not two measurements of one quantity. As 502-against-500 it reads as a row getting two milliseconds cheaper and crossing a line, which is a story about the row; what actually changed between attempts is WHAT THE OBSERVER COULD SAY. That is the same point in its sharpest form, and my number quietly converted it back into the weaker one. It is also the class this census exists to stop: a transcribed instrument property standing where a subject property belongs, inside the diff that files it. Receipt (i) is left as figures because both those rows COMPLETED over budget -- measured values, not preemption bounds -- and the prose now says so. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Fjt9vSfVCKwZu9x7d44xZP
The only authored conflict was gunbc.recurring_failure_mode: main added three class rows (ambient_process_state_read_by_a_concurrent_reader, predicate_vacuously_true_on_an_empty_domain, check_subject_narrower_than_its_declared_claim) at the same append point as recurrence_ledger_scoped_below_the_recurrence. Resolved purely additively -- all four rows and all four roster entries, source order preserved, 55 declarations against 55 roster entries. The three conflicted PROJECTIONS -- witnesses.yml, DESIGN.md and docs/design-ledgers.md -- were not resolved by taking a side. They were REGENERATED over the merged .dag authority through the generated-artifact gate, so the committed bytes are derived rather than composed. That is what the refusing merge driver exists to force: a text merge of a projection can produce bytes no authority would emit. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Fjt9vSfVCKwZu9x7d44xZP
claim_cpu_total_ms was added as a bound against summing a column that is inclusive of callees. It caught the variance term instead: 130335ms on run 33615659836 and 99943ms on run 33631458679, both run_attempt=1, over the same corpus. A per-closure constant cannot produce a 30% swing -- it is by construction identical on two runs of one tree. Unexplained, and deliberately not pursued here: this census's subject is the level, and that boundary is declared in three carriers. But it is a property of the ENVIRONMENT measured from inside the run, which is what a fleet-variance account has lacked, and its only other home was two logs that age out. The two figures are transcribed as a declared exception to naming the instrument rather than copying its output, on the same ground the floor-cost carrier grants it: the subject IS the divergence between two runs, and no producer on this side of the boundary can re-derive it once the logs expire. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Fjt9vSfVCKwZu9x7d44xZP
Contributor
Author
|
Closing as empty: this draft was auto-attached to |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Auto-opened by session-dashboard for session
nimble-lynx-128.Pushing to
session/nimble-lynx-128advances this PR.Worker attestation
Before flipping this PR to ready for review, confirm each item:
npm test,cargo test) and the result.Closes #Ndirective.Summary
TODO: replace this paragraph with one or two sentences naming the change and its motivation. Reviewers read this first.
Test plan