Skip to content

The 500ms per-witness CPU budget: one witness family preempts across unrelated diffs — four lanes taxed in a day. Establish whether the budget or the family is wrong, at class grain - #10076

Closed
gunbai-bot[bot] wants to merge 11 commits into
mainfrom
session/nimble-lynx-128

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor

Auto-opened by session-dashboard for session nimble-lynx-128.
Pushing to session/nimble-lynx-128 advances this PR.

Worker attestation

Before flipping this PR to ready for review, confirm each item:

  • Title describes the change (not the session id or branch).
  • PR body summarises what and why (replace the TODO below).
  • Tests run: name the command (e.g. npm test, cargo test) and the result.
  • If this closes a work item, the body contains a Closes #N directive.
  • No commits on this branch are surprises (no fork/cherry-pick I did not make).
  • No secrets / credentials / large binaries staged.

Summary

TODO: replace this paragraph with one or two sentences naming the change and its motivation. Reviewers read this first.

Test plan

  • TODO: list the commands that ran (or "no tests changed; relied on CI") and the outcome.

gunbc-ci-auto-heal and others added 11 commits September 2, 2026 11:06
…e the floor a cross-claim demand census

The 500ms per-claim CPU ceiling has been converting unrelated diffs into merge
blocks -- four lanes in one day, every run reporting failed=0. Neither the
budget nor the witnesses are wrong.

MEASURED (main run 33615659836, required_floor_claim_cost.tsv, 3477 rows,
cost_basis=cpu). The corpus p50 is 2ms and 3379 rows sit under 280ms, so the
ceiling is not mis-set against the corpus. And the tail family's own work is
among the smallest in it: a witness and its DISCRIMINATING RED do materially
different work, and they cost 296/294, 333/331 and 384/382 -- within 2ms of each
other while both sit against the ceiling. A cost that does not move when the
assertion changes is not the assertion's cost. One 34-row module spans 0ms to
the ceiling in three bands that track WHICH SHARED PRODUCER a row forces rather
than what it asserts.

The charge is a per-CLOSURE constant: the floor builds a fresh evaluation frame
per claim, so every claim re-derives the pure substrate its import closure
reaches, and that constant is billed as the claim's own marginal work. The run's
[floor-shared-fill] ledger nets nothing for any of the four families.

THE FINDING IS THE BLIND INSTRUMENT. v2.workflow.floor_pure_producer_share
already exists to stop exactly this recompute, and its roster is six
hand-authored rows because nothing produces its candidates: the interpreter's
recompute-trace ledger ranks pure calls with count>=2 WITHIN one InterpContext
and is printed and dropped at every claim frame exit, so a producer evaluated
exactly once per claim, in 3477 claims, carries count=1 in every ledger and
appears in none. The instrument that ranks redundant recompute is structurally
blind to the population its own repair roster is enrolled from -- which is why
roster coverage is discovered when a budget refusal lands on someone else's PR.

WHAT THIS LANDS. absorb_claim_recompute_demand folds each claim's ledger into a
run-scoped census keyed on declaration identity plus argument row, before the
frame dies; the floor prints the ranked head and writes
required_floor_cross_claim_demand.tsv, uploaded as an artifact. Every truncation
is disclosed (retention floor, key cap, bounded module sample beside an exact
module count), and single-claim rows are retained so the shared population has a
control.

IT ENROLS NOTHING AND GATES NOTHING. A row is a candidate whose SERVE cost is
unmeasured; floor_pure_producer_share records the case where the serve lost to
the recompute and two enrolled rows were removed. No budget change, no
quarantine, no per-lane accommodation -- the limit is correctly placed against
the corpus, so raising it would move a line that is right to accommodate a
charge that is wrong.

Class filed as recurrence_ledger_scoped_below_the_recurrence. Seed growth
justified at gunbc.cross_claim_demand_census_seed_growth (precedent gunbc#9721).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fjt9vSfVCKwZu9x7d44xZP
…laims would double-fold

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fjt9vSfVCKwZu9x7d44xZP
Single-claim rows are RETAINED and rank at zero -- they are the control that
makes claims>1 mean something. The docstring said they were dropped, which
points a reader at the wrong side of the census's one load-bearing claim. The
code, the writer and the tests were always right; only the sentence was wrong.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fjt9vSfVCKwZu9x7d44xZP
REVIEW 58673 (REQUEST_CHANGES), both findings real and both fixed by
construction rather than by validation:

  THE ARGUMENT ROW WAS A 64-BIT HASH. Two distinct argument rows could collide
  and merge into one row reporting cross-claim demand that never happened --
  a fabricated identity, forbidden outright by DESIGN §5, and contradicting
  this file's own promise of "sound argument identity" ten lines up. The key
  now carries the canonical argument vector the ledger already holds, so the
  collision is unrepresentable rather than unlikely. `keyed` and `unkeyed` are
  now discriminated in the key too: a nullary keyed call and a declaration's
  composite-argument bucket both carry an empty argument vector, and merging
  them would sum one identity's cost with all of another's.

  THE MODULE COUNT WAS COUNTING CLAIMS. The bounded sample and the counter
  shared a container, so once the eight-name sample filled, every later claim
  from an unsampled module incremented the count again -- a column promising
  distinct consumer modules while reporting claim occurrences. The complete set
  is now kept (interned, one allocation per module name for the whole census)
  and the count is its length; the cap bounds only how many names a row shows.

RULINGS APPLIED.

  MEANING BELONGS IN .dag. Ranking is a judgment about which demands matter, so
  the artifact now leaves in deterministic IDENTITY order and the runner's log
  preview sorts a copy and says in band that it is a preview and not a
  candidate roster. Grouping stays where the ledger's own key is; carrying
  facts across a frame boundary only the seed can reach is the seed's warrant.

  THE INSTRUMENT MEASURES ITSELF. absorb_ms and absorb_max_ms are reported and
  written: the absorb runs after each claim's measurement returns, so it is
  outside every charged window and cannot trip the deadline, and that claim is
  no longer asked to stand on reasoning alone. Also removed a quadratic in it:
  declaration sites are resolved once per fn pointer per frame instead of by a
  linear scan per ledger key -- the cost-shape defect §6 always fixes, inside
  an instrument whose subject is cost.

  THE DROP IS SPLIT, NOT RETIRED. The class row now says a cross-unit census
  plus the modeled fold retires blind candidate discovery, incident-as-
  discovery and the missing carrier -- and NOT environment-independent
  claim-cost qualification, which no amount of producer naming makes invariant
  across execution envelopes. Claiming the whole drop would have been the rung
  inflation this ledger exists to catch.

  THE NEGATIVE CONTROL IS EXECUTED. The two rust target models were removed
  from the share roster on a measured serve-versus-recompute experiment, and
  the first artifact ranks that chain in the top ten while it stays unenrolled
  and unenrollable from here.

ALSO CARRIED, from four lanes' observations: the census explains the LEVEL and
not the VARIANCE (a per-closure constant is identical on two runs of one tree,
while byte-identical evaluator steps have been measured against 1.31-1.80x cpu),
the two compose into the wandering victim, and a cost inside a native builtin
called directly from a claim body is not a row here at all -- the same boundary
std.evaluation_budget's opaque-host-call note draws for the deadline.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fjt9vSfVCKwZu9x7d44xZP
Asked to denominate the displaced cost as one number, I computed it and it was
wrong by construction: summing cross-claim over the shared rows of the first
artifact gives ~1850s against a run whose ENTIRE claim-side CPU was 130s. A 14x
overcount, and it is nesting -- the ledger times a producer's whole subtree, so
a producer and its callees both appear and overlap.

A per-row figure is a valid statement about that producer; no sum of rows is a
valid statement about the run. The artifact now carries claim_cpu_total_ms as
the ceiling any true total must sit under, says cost_columns=inclusive_of_
callees_do_not_sum in its summary, and names what a real displaceable-cost
figure needs: self-time, which this ledger does not carry.

The artifact invited the error, so the refusal belongs in the artifact rather
than in a reviewer's memory.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fjt9vSfVCKwZu9x7d44xZP
… underivable rather than unstated

The sentence 'this floor recomputes N seconds of pure producer work per run' is
what makes the stake legible, and no arithmetic over this artifact produces it:
the quantity it needs is not in the column. That is a next-rung trigger, not a
caveat.

The capability is SELF TIME -- inclusive minus the callees the same pass already
counted -- after which the column is additive and the sentence is derivable. It
is a TRACKED stall rather than a wish: CrossClaimFillGuard's Drop already
computes exactly that netting against the CROSS_CLAIM_FILL_FRAMES child stack,
one tier over, for the shared-fill ledger. What is missing is a child stack over
the recompute ledger's frames.

Deliberately not built here: this bridge is what stops the next lanes paying,
and a new measurement tier would put it behind a fresh review cycle on the very
fleet condition it exists to explain.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fjt9vSfVCKwZu9x7d44xZP
… is single-run

calm-boar-314 retracted the two-run reading this prose leaned on: a rerun
replaces a job's ATTEMPT while the logs API serves the latest attempt for a
fixed job id, so the six-row and thirteen-row readings addressed different
attempts through one identifier. No error is raised and nothing is visible from
the citing end. In the reproducible reading the two rows said to have dropped
out are in the set.

Replaced with evidence that needs no cross-run delta: two rows of one module
measured 502ms and EXACTLY 500ms against the 500ms budget. A row sitting at its
budget lands in interrupted-before-verdict or completed-over-cost according to
whether the poll fired before or after the work finished -- a fact about the
poll, not about the row. Same conclusion, one run, nothing to retract.

The class row now carries the retraction as part of the class, because it
sharpens the recommendation: a remedy must never be keyed to WHICH rows tipped.
A producer census is invariant under that retraction; a ranking of tipped rows
would have been built on it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fjt9vSfVCKwZu9x7d44xZP
…d independently

calm-boar-314 over-retracted and I cut a true finding on it. Attempt logs ARE
recoverable -- /actions/runs/<id>/attempts/<n>/logs -- and I pulled both
attempts of run 33620893203 myself rather than restoring on a second say-so:

  attempt 1: interrupted_before_verdict=6, and
    ..an_empty_receipt_series_leaves_the_live_tree_unmeasured_rather_than_held
    is INTERRUPTED-BEFORE-VERDICT at 502ms
  attempt 2: interrupted_before_verdict=13, and the SAME row is
    COMPLETED-OVER-COST-REQUIREMENT at cpu_ms=500 -- it reached its verdict

One row, one tree, crossing the two populations the 2026-08-19 cut separated
BECAUSE THEY ARE DIFFERENT FACTS WITH DIFFERENT REMEDIES. Which one it lands in
is decided by whether the poll fired before or after the work finished.

STILL WITHDRAWN: near-disjointness (attempt 2's set largely CONTAINS attempt
1's) and the 3-6-13 trend (the 3 was a different head). Only the migration
returns.

CITE THE ATTEMPT, NOT THE RUN, now recorded as part of the class: a bare run id
and a bare job id both resolve to the latest attempt, so a rerun changes what a
stable citation serves with no error and nothing visible from the citing end --
which is why the retraction looked sound. My own figures now carry
run_attempt=1 for runs 33615659836 and 33622954427, checked rather than assumed.

The census reported the same thing through the retraction AND the restoration,
because it keys producers and never which rows tipped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fjt9vSfVCKwZu9x7d44xZP
I wrote the restored migration as 'INTERRUPTED-BEFORE-VERDICT at 502ms'. The
attempt-1 diagnostic says the opposite in three clauses: cost=UNMEASURED, above
500ms with NO UPPER BOUND, and interrupt_point names where the POLL observed the
ceiling -- a property of the budget, not of the row.

The pair is a BOUND beside a VALUE, not two measurements of one quantity. As
502-against-500 it reads as a row getting two milliseconds cheaper and crossing
a line, which is a story about the row; what actually changed between attempts
is WHAT THE OBSERVER COULD SAY. That is the same point in its sharpest form, and
my number quietly converted it back into the weaker one.

It is also the class this census exists to stop: a transcribed instrument
property standing where a subject property belongs, inside the diff that files
it. Receipt (i) is left as figures because both those rows COMPLETED over budget
-- measured values, not preemption bounds -- and the prose now says so.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fjt9vSfVCKwZu9x7d44xZP
The only authored conflict was gunbc.recurring_failure_mode: main added three
class rows (ambient_process_state_read_by_a_concurrent_reader,
predicate_vacuously_true_on_an_empty_domain,
check_subject_narrower_than_its_declared_claim) at the same append point as
recurrence_ledger_scoped_below_the_recurrence. Resolved purely additively -- all
four rows and all four roster entries, source order preserved, 55 declarations
against 55 roster entries.

The three conflicted PROJECTIONS -- witnesses.yml, DESIGN.md and
docs/design-ledgers.md -- were not resolved by taking a side. They were
REGENERATED over the merged .dag authority through the generated-artifact gate,
so the committed bytes are derived rather than composed. That is what the
refusing merge driver exists to force: a text merge of a projection can produce
bytes no authority would emit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fjt9vSfVCKwZu9x7d44xZP
claim_cpu_total_ms was added as a bound against summing a column that is
inclusive of callees. It caught the variance term instead: 130335ms on run
33615659836 and 99943ms on run 33631458679, both run_attempt=1, over the same
corpus. A per-closure constant cannot produce a 30% swing -- it is by
construction identical on two runs of one tree.

Unexplained, and deliberately not pursued here: this census's subject is the
level, and that boundary is declared in three carriers. But it is a property of
the ENVIRONMENT measured from inside the run, which is what a fleet-variance
account has lacked, and its only other home was two logs that age out.

The two figures are transcribed as a declared exception to naming the instrument
rather than copying its output, on the same ground the floor-cost carrier grants
it: the subject IS the divergence between two runs, and no producer on this side
of the boundary can re-derive it once the logs expire.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fjt9vSfVCKwZu9x7d44xZP
@gunbai-bot

gunbai-bot Bot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor Author

Closing as empty: this draft was auto-attached to session/nimble-lynx-128 after the lane's work already landed. The content shipped as #10053, squash-merged at 0abc7c3 — dag/gunbc/floor/cross_claim_demand_census_seed_growth.dag and absorb_claim_recompute_demand are both present on main. The branch now carries nothing unlanded and is behind main, so there is nothing here to flip to ready.

@gunbai-bot gunbai-bot Bot closed this Sep 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants