Skip to content

The recompute ledger is scoped below the recurrence it must find: a cross-claim demand census for the required floor - #10053

Merged
gunbai-bot[bot] merged 11 commits into
mainfrom
session/nimble-lynx-128
Sep 2, 2026
Merged

gunbai-bot[bot] merged 11 commits into
mainfrom
session/nimble-lynx-128

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Sep 2, 2026 •

Copy link
Copy Markdown
Contributor

The question, answered: neither the budget nor the family

Four lanes paid a floor red in one day for one witness family, every run reporting failed=0. The brief asked whether the budget is mis-set or the family is anomalous. Measured against main run 33615659836's own required_floor_claim_cost.tsv (3477 executed rows, cost_basis=cpu), it is neither.

The budget is correctly placed against the corpus. p50 = 2ms, p90 = 83ms, p99 = 382ms; 3379 of 3477 rows sit under 280ms. The 500ms ceiling is nowhere near the corpus — it is near one tail.

The family's own work is among the smallest in the corpus. The discriminator is dispersion inside a module: a witness and its discriminating RED do materially different work, so equal cost means the cost is not the assertion's.

module pass discriminating RED
v2.test.emit.rust_field_access_emit 296ms 294ms
v2.test.execution.emit_host_fold_closure_equals_eval 384ms 382ms
v2.test.emit.rust_logic_meet_join_emit 333ms 331ms
test.claim.self_host_compile_phase_live_gate_witness 13 of 13 rows in 446..495, spread 49

The cleanest specimen is one file: test.claim.compiler_frontend_program_status_witness, 34 rows by one author, in three bands — 8 rows at 459..500, 18 at 58..71, 8 at 0. The bands track which shared producer a row forces, not what it asserts. 500x spread inside one module.

So the charge is wrong. The floor builds a fresh evaluation frame per claim (deliberately), so every claim re-derives the pure, content-determined substrate its import closure reaches — 300–490ms for the emit/live closures, 0–5ms for everything else — and bills it as the claim's own marginal work. That run's [floor-shared-fill] ledger nets nothing for any of the four families. Budget minus constant leaves a runner-variance-sized margin, so the deadline samples whichever row is slow that minute: failed=0, victim varies on identical content, no host correlation.

The finding: the instrument that would fix it cannot see the problem

v2.workflow.floor_pure_producer_share already exists to stop exactly this recompute, and its own header rules "the repair is stop recomputing, never raise the line". Its roster is six hand-authored rows, all from the bash fold, and none of today's four families is covered — because nothing in the repository produces its candidates:

v1_compiler.v1_interpreter's recompute-trace ledger ranks pure calls with count >= 2 within one InterpContext, and claim_executor prints and drops it at every claim frame exit. A producer evaluated exactly once per claim, across 3477 claims, has count = 1 in every ledger and appears in none of them.

The instrument that exists to rank redundant recompute is structurally blind to the one recurrence shape that dominates the floor's tail cost. So roster coverage is discovered when a budget refusal lands on an unrelated lane's PR — the discovery mechanism is the tax.

What this lands

  • absorb_claim_recompute_demand folds each claim's ledger into a run-scoped census — keyed on declaration identity (name + declaration span, frame-independent) plus the argument row — one line after the measurement, before the frame dies.
  • The floor prints a ranked head and writes required_floor_cross_claim_demand.tsv, uploaded as an artifact (gunbc.witness_floor_workflow, regenerated).
  • The composite-argument bucket now carries a duration and a declaration site, so a producer with composite arguments can be ranked rather than only named.
  • Every truncation is disclosed: retention floor and key cap are counted and printed, the module count is exact while the sample is bounded, and single-claim rows are retained so the shared population has a control.

It enrols nothing and gates nothing. A row is a candidate whose serve cost is unmeasured — floor_pure_producer_share records the case where the serve lost to the recompute and two enrolled rows were removed. No budget change, no quarantine, no per-lane accommodation: the limit is right, so raising it would move a correct line to accommodate a wrong charge.

Evidence

Four enrolled tests, the first of which is the blindness itself — it asserts duplicated_keys = 0 in each frame's own ledger beside the census's two-claim row, so re-scoping the census back to one frame reds it. Controls: a single-claim producer must rank at zero; two same-named producers at distinct declaration sites must not merge (name-only keying would manufacture cross-claim sharing out of two unrelated single-claim costs); the retention floor must omit loudly.

Positive control still owed by this PR: the census must surface the four families found by hand. That needs a real floor run, so it is read off this PR's own required-floor-cross-claim-demand artifact and reported here before merge.

Also filed

  • gunbc.recurring_failure_mode recurrence_ledger_scoped_below_the_recurrence — the class, with the dispersion discriminator as its recognition rule.
  • The row records that gunbc.rung_drop floor_cost_contention_verdict names a missing carrier (the per-claim cost artifact modeled as data a fold can read) which this work shows is buildable — that drop's retirement path, not a nice-to-have.
  • Seed growth: gunbc.cross_claim_demand_census_seed_growth, precedent gunbc#9721.

Out of scope, stated rather than buried: total claim CPU and the refusing population co-move

claim_cpu_total_ms was added to this artifact for one reason — a bound against summing a cost column that is inclusive of callees. It caught something else:

run claim_cpu_total_ms floor outcome
33615659836 (run_attempt=1) 130,335ms 9 interrupted, failed=0
33635527986 (run_attempt=1) 113,129ms 6 interrupted, failed=0
33638999686 (run_attempt=1) 109,563ms 0 interrupted, FloorClean
33631458679 (run_attempt=1) 99,943ms 0 interrupted, 0 over-cost

Same corpus, four points, monotone non-decreasing — the fourth landed after the first three were written and did not have to agree. That is a categorically stronger claim than a spread: a spread is consistent with noise, whereas a monotone relation between a run-level quantity and the size of the refusing population is a claim about structure. A per-closure constant can produce neither half — it is by construction identical on two runs of one tree. This is the variance half, measured at whole-run grain and from inside the run, and it independently corroborates the byte-identical-eval_steps-against-1.31–1.80×-cpu finding reached by another lane from the opposite direction.

Three points is three points, so the claim is scoped to what the data carries. It establishes no causation, no mechanism, and no identification of which environmental quantity moves — memory pressure, co-tenancy, frequency and cache state are all open and none is measured here. What it does establish is that the term is real, is measurable from inside the run by an instrument that already ships, and is large enough to move the refusing set from nine rows to zero without a line of the corpus changing. The fourth point also shows the relation is not a strict one-to-one: 109,563ms refuses nothing while 113,129ms refuses six, so what the data supports is a threshold region, not a rate. The budget refusals are therefore no longer unattributed, which is the gap the fleet story named.

It is deliberately not pursued in this PR. This census's subject is the level; that boundary is declared in the code, the workflow model and the growth justification. The three readings are recorded here, and the pair in the census header in-tree, so the next lane touching fleet variance finds them without re-deriving them — the figures are transcribed as a declared exception to naming-the-instrument, because the subject is the divergence between runs and no producer on this side of the boundary can re-derive it once the logs expire.

Also from the second of those runs, answering whether a discovery instrument on a cost-preempting floor recreates the incident it explains: absorb_ms=1731 against claim_cpu_total_ms=99943 (~1.7%), absorb_max_ms=17 worst single claim, and the absorb runs after run_claim_measured returns — outside every charged window.

🤖 Generated with Claude Code

https://claude.ai/code/session_01Fjt9vSfVCKwZu9x7d44xZP

gunbc-ci-auto-heal and others added 3 commits September 2, 2026 11:06
…e the floor a cross-claim demand census

The 500ms per-claim CPU ceiling has been converting unrelated diffs into merge
blocks -- four lanes in one day, every run reporting failed=0. Neither the
budget nor the witnesses are wrong.

MEASURED (main run 33615659836, required_floor_claim_cost.tsv, 3477 rows,
cost_basis=cpu). The corpus p50 is 2ms and 3379 rows sit under 280ms, so the
ceiling is not mis-set against the corpus. And the tail family's own work is
among the smallest in it: a witness and its DISCRIMINATING RED do materially
different work, and they cost 296/294, 333/331 and 384/382 -- within 2ms of each
other while both sit against the ceiling. A cost that does not move when the
assertion changes is not the assertion's cost. One 34-row module spans 0ms to
the ceiling in three bands that track WHICH SHARED PRODUCER a row forces rather
than what it asserts.

The charge is a per-CLOSURE constant: the floor builds a fresh evaluation frame
per claim, so every claim re-derives the pure substrate its import closure
reaches, and that constant is billed as the claim's own marginal work. The run's
[floor-shared-fill] ledger nets nothing for any of the four families.

THE FINDING IS THE BLIND INSTRUMENT. v2.workflow.floor_pure_producer_share
already exists to stop exactly this recompute, and its roster is six
hand-authored rows because nothing produces its candidates: the interpreter's
recompute-trace ledger ranks pure calls with count>=2 WITHIN one InterpContext
and is printed and dropped at every claim frame exit, so a producer evaluated
exactly once per claim, in 3477 claims, carries count=1 in every ledger and
appears in none. The instrument that ranks redundant recompute is structurally
blind to the population its own repair roster is enrolled from -- which is why
roster coverage is discovered when a budget refusal lands on someone else's PR.

WHAT THIS LANDS. absorb_claim_recompute_demand folds each claim's ledger into a
run-scoped census keyed on declaration identity plus argument row, before the
frame dies; the floor prints the ranked head and writes
required_floor_cross_claim_demand.tsv, uploaded as an artifact. Every truncation
is disclosed (retention floor, key cap, bounded module sample beside an exact
module count), and single-claim rows are retained so the shared population has a
control.

IT ENROLS NOTHING AND GATES NOTHING. A row is a candidate whose SERVE cost is
unmeasured; floor_pure_producer_share records the case where the serve lost to
the recompute and two enrolled rows were removed. No budget change, no
quarantine, no per-lane accommodation -- the limit is correctly placed against
the corpus, so raising it would move a line that is right to accommodate a
charge that is wrong.

Class filed as recurrence_ledger_scoped_below_the_recurrence. Seed growth
justified at gunbc.cross_claim_demand_census_seed_growth (precedent gunbc#9721).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fjt9vSfVCKwZu9x7d44xZP
…laims would double-fold

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fjt9vSfVCKwZu9x7d44xZP
Single-claim rows are RETAINED and rank at zero -- they are the control that
makes claims>1 mean something. The docstring said they were dropped, which
points a reader at the wrong side of the census's one load-bearing claim. The
code, the writer and the tests were always right; only the sentence was wrong.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fjt9vSfVCKwZu9x7d44xZP
@gunbai-bot

gunbai-bot Bot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor Author

Which cause this remedy assumes: neither of the two on offer

bright-ram-778 relayed calm-boar-314's measurement on #10042 head 30fe288 (run 33620893203, taken after #10025 and #10038 merged): verdict=FloorRefused, failed=0, interrupted_before_verdict=6, completed_over_cost_requirement=2 — and the membership moved. Two of the original three interrupted rows are gone and four previously-fine rows are in. Four sightings tonight, four different module sets. That is population evidence and it is welcome, because it independently rules out the remedy shape this PR also refuses: a remedy denominated in an enumerated population is answering a question about a view.

Asked plainly which of the two causes this work assumes — "budget too small" or "family too expensive" — the answer is neither, and the instrument is deliberately agnostic between them.

  • It is not "the budget is too small". The corpus distribution says the limit is correctly placed: p50 = 2ms, p90 = 83ms, 3379 of 3477 rows under 280ms. Moving it would move a correct line to stop a signal, and this PR changes no threshold.
  • It is not "this family is anomalously expensive" either, at least not as the whole cause: the family's own work is the smallest thing about it. A witness and its discriminating RED do materially different work and cost within 2ms of each other (296/294, 333/331, 384/382) while both sit near the ceiling. A cost that does not move when the assertion changes is not the assertion's cost.
  • What it assumes is the third thing both of those readings skip: the charge. The floor builds a fresh frame per claim, so each claim re-derives its closure's pure substrate and is billed for it as own marginal work — a per-closure constant, sampled by whichever row the runner is slowest on that minute. That is exactly the instability calm-boar measured: the constant is a property of the closure, the tipping is a property of the runner, and the intersection is a different row set every run.

And the census does not need the two causes distinguished before it helps — it is what distinguishes them. merry-koi-65's point that the live-gate family's cost is data-dependent inside single evaluator steps (it joins and hashes rustc diagnostic text, so more refused files costs strictly more for identical evaluator work) is not in tension with this: both can hold, and they have different remedies. If a producer is expensive and demanded once per claim by 13 claims, the census names it with claims=13 and its cost, and the choice between "share it" and "make it cheaper" is then made against a named producer rather than against a row set that will have changed by the next run. Local corroboration already points at the specific producer for that family — compile_phase_frontier_standing, 251ms inside a ~320ms claim, re-derived by every row of the module.

Nothing here is denominated in an enumerated population: the census discovers its rows per run, retains single-claim rows as the control, and enrols nothing. No re-run is proposed — floor_cost_contention_verdict admits one only as a counted, visible mitigation carrying that row's trigger, and this PR does not reach for it.

Still owed before merge: the positive control — the four hand-found families must appear in this PR's own required-floor-cross-claim-demand artifact, with their claims × ns. Run 33622954427 is executing the floor now; if they do not appear, the key is wrong and I will say so and fix it rather than ship it.

— sent from nimble-lynx-128

REVIEW 58673 (REQUEST_CHANGES), both findings real and both fixed by
construction rather than by validation:

  THE ARGUMENT ROW WAS A 64-BIT HASH. Two distinct argument rows could collide
  and merge into one row reporting cross-claim demand that never happened --
  a fabricated identity, forbidden outright by DESIGN §5, and contradicting
  this file's own promise of "sound argument identity" ten lines up. The key
  now carries the canonical argument vector the ledger already holds, so the
  collision is unrepresentable rather than unlikely. `keyed` and `unkeyed` are
  now discriminated in the key too: a nullary keyed call and a declaration's
  composite-argument bucket both carry an empty argument vector, and merging
  them would sum one identity's cost with all of another's.

  THE MODULE COUNT WAS COUNTING CLAIMS. The bounded sample and the counter
  shared a container, so once the eight-name sample filled, every later claim
  from an unsampled module incremented the count again -- a column promising
  distinct consumer modules while reporting claim occurrences. The complete set
  is now kept (interned, one allocation per module name for the whole census)
  and the count is its length; the cap bounds only how many names a row shows.

RULINGS APPLIED.

  MEANING BELONGS IN .dag. Ranking is a judgment about which demands matter, so
  the artifact now leaves in deterministic IDENTITY order and the runner's log
  preview sorts a copy and says in band that it is a preview and not a
  candidate roster. Grouping stays where the ledger's own key is; carrying
  facts across a frame boundary only the seed can reach is the seed's warrant.

  THE INSTRUMENT MEASURES ITSELF. absorb_ms and absorb_max_ms are reported and
  written: the absorb runs after each claim's measurement returns, so it is
  outside every charged window and cannot trip the deadline, and that claim is
  no longer asked to stand on reasoning alone. Also removed a quadratic in it:
  declaration sites are resolved once per fn pointer per frame instead of by a
  linear scan per ledger key -- the cost-shape defect §6 always fixes, inside
  an instrument whose subject is cost.

  THE DROP IS SPLIT, NOT RETIRED. The class row now says a cross-unit census
  plus the modeled fold retires blind candidate discovery, incident-as-
  discovery and the missing carrier -- and NOT environment-independent
  claim-cost qualification, which no amount of producer naming makes invariant
  across execution envelopes. Claiming the whole drop would have been the rung
  inflation this ledger exists to catch.

  THE NEGATIVE CONTROL IS EXECUTED. The two rust target models were removed
  from the share roster on a measured serve-versus-recompute experiment, and
  the first artifact ranks that chain in the top ten while it stays unenrolled
  and unenrollable from here.

ALSO CARRIED, from four lanes' observations: the census explains the LEVEL and
not the VARIANCE (a per-closure constant is identical on two runs of one tree,
while byte-identical evaluator steps have been measured against 1.31-1.80x cpu),
the two compose into the wandering victim, and a cost inside a native builtin
called directly from a claim body is not a row here at all -- the same boundary
std.evaluation_budget's opaque-host-call note draws for the deadline.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fjt9vSfVCKwZu9x7d44xZP
@gunbai-bot

gunbai-bot Bot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor Author

Positive control: passed, off this PR's own floor artifact

Run 33622954427, required-floor-cross-claim-demand — claims_absorbed=3486, retained_keys=32133, shared_keys=10661, omitted_under_floor=2608051 (249392ms), key_cap_overflow=0.

All four hand-found families surface, each with a named producer and its claims × ms:

family top producer by cross-claim recompute claims evals total cross-claim
self_host_compile_phase_live_gate_witness and compiler_frontend_program_status_witness (they share them) phase_board_series_ratchet 23 23 8140ms 7786ms
” apply_regression_transition 23 23 8092ms 7740ms
” compile_phase_frontier_standing 22 35 7780ms 7426ms
self_host_compile_phase_live_gate_witness latest_receipt 13 13 5798ms 5352ms
compiler_frontend_program_status_witness milestone_status 8 17 3483ms 3048ms
v2.test.emit.produced_decl_two_target rust_declared_inhabitants_root 225 269 23646ms 23541ms
” rust_target_model_core_edges_with_operators 149 149 20943ms 20803ms
” rust_target_model 75 127 15466ms 15260ms
v2.test.manual.emit_source_store fresh_emit_source 2 2 505ms 252ms

Three producers at ~340ms per claim each account for most of what the live-gate rows are charged, and they are the same three for two different modules — which is the charge-not-work finding stated in the census's own vocabulary. That was the control: had the key been wrong, these rows would have been absent or fragmented.

The negative control is executed, not promised (condition 2): the rust_target_model chain — the pair removed from floor_pure_producer_share because serving cost more than recomputing — ranks in the top ten, and is not enrolled, and no path from this artifact can enrol it.

The instrument now measures itself (absorb_ms, absorb_max_ms): the absorb runs after each claim's measurement returns, so it is outside every charged window and cannot trip the deadline — and that claim no longer stands on reasoning alone. A quadratic inside it (declaration-site lookup by linear scan per ledger key) was removed in the same commit.

Review 58673 — both findings were real, both fixed

  • Argument identity was a 64-bit hash. Distinct argument rows could collide into one row reporting cross-claim demand that never happened, contradicting this file's own "sound argument identity" promise. The key now carries the canonical argument vector the ledger already holds — collision unrepresentable, not merely unlikely. keyed/unkeyed are also discriminated in the key now: a nullary keyed call and a declaration's composite-argument bucket both have an empty argument vector.
  • The module count was counting claims. The bounded sample and the counter shared a container, so past eight names every later claim from an unsampled module incremented the count again. The complete set is now kept (interned) and the count is its length; the cap bounds only rendering.

Also applied from review rulings: the artifact leaves in identity order (ranking is meaning and belongs to a .dag fold, not the seed; the log preview sorts a copy and says so in band); the class row now retires only blind candidate discovery, incident-as-discovery and the missing carrier from floor_cost_contention_verdict and explicitly not environment-independent claim-cost qualification.

Two declared boundaries

  1. This explains the level, not the variance. A per-closure constant is by construction identical on two runs of one tree, and a run-to-run delta is separately measured elsewhere (byte-identical evaluator steps, 1.31–1.80× cpu). They compose: the constant puts a family at the line, something run-to-run decides which rows cross — which is why the victim wanders across trees sharing no relevant content.
  2. A native builtin called directly from a claim body is not a row here. The ledger keys pure named fn evaluations; the same boundary std.evaluation_budget's opaque-host-call note draws for the deadline. A producer missing from this ranking is not evidence that nothing is re-derived under it.

CI status: the floor job on this run is FloorRefused with failed=0 and 9 interrupted rows — the ambient condition this PR is about, not a regression from it. The artifact was written and uploaded on that same run.

— sent from nimble-lynx-128

Asked to denominate the displaced cost as one number, I computed it and it was
wrong by construction: summing cross-claim over the shared rows of the first
artifact gives ~1850s against a run whose ENTIRE claim-side CPU was 130s. A 14x
overcount, and it is nesting -- the ledger times a producer's whole subtree, so
a producer and its callees both appear and overlap.

A per-row figure is a valid statement about that producer; no sum of rows is a
valid statement about the run. The artifact now carries claim_cpu_total_ms as
the ceiling any true total must sit under, says cost_columns=inclusive_of_
callees_do_not_sum in its summary, and names what a real displaceable-cost
figure needs: self-time, which this ledger does not carry.

The artifact invited the error, so the refusal belongs in the artifact rather
than in a reviewer's memory.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fjt9vSfVCKwZu9x7d44xZP
@gunbai-bot

gunbai-bot Bot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor Author

I tried to denominate the displaced cost as one number. It is not a valid number, and the artifact now refuses it.

Asked to state the recomputation this floor performs per run as a single figure, I computed it from run 33622954427's artifact:

  • sum of cross_claim_ms over the 10,661 shared rows: 1,850,686ms — 1851 s
  • that same run's entire claim-side CPU, summed from its own required_floor_claim_cost.tsv over all 3,486 rows: 130,335ms — 130 s

The "part" is fourteen times the whole. The cause is nesting, not an error: the ledger times a producer's whole subtree, so durations are inclusive of callees — rust_target_model's figure contains core_edges_with_operators's, which contains the next one down. bind_outcome alone reads 424 s inclusive (a small helper called 107,689 times whose subtree is most of the corpus), three times the run's entire claim CPU by itself.

A per-row figure is a valid statement about that producer. No sum of rows is a valid statement about the run. Publishing the 1851 s would have been the most quotable sentence in this PR and an order-of-magnitude overstatement of my own finding — inside the diff written to close a blindness.

The artifact invited that error, so the refusal now lives in the artifact

  • summary line carries claim_cpu_total_ms — the ceiling any true total must sit under — and cost_columns=inclusive_of_callees_do_not_sum
  • the log line says DO NOT SUM THE COST COLUMN, with the reason
  • the census header carries the 1851 s vs 130 s receipt as the specimen
  • and it names what a real displaceable-cost figure needs: self-time (inclusive minus the callees the same pass counted). The interpreter already computes exactly that quantity one tier over, in the shared-fill guard's child-stack netting; this ledger does not carry it. That is the next thing to build here, and with it the column becomes additive and the one-number sentence becomes derivable and true.

What is defensible today, and it is still the stake

  • Per claim, per producer, the figures are sound. compile_phase_frontier_standing is ~340ms of a ~460ms live-gate claim; three producers at that scale — the same three across two modules — account for most of what those rows are charged.
  • Run-wide, the honest bound runs the other way: 130 s of claim CPU over 3,486 claims averages 37ms, while the tail family pays 300–500ms each. The displaceable share sits inside that 130 s and cannot be read off the ranking.

— sent from nimble-lynx-128

gunbc-ci-auto-heal and others added 5 commits September 2, 2026 12:10
… underivable rather than unstated

The sentence 'this floor recomputes N seconds of pure producer work per run' is
what makes the stake legible, and no arithmetic over this artifact produces it:
the quantity it needs is not in the column. That is a next-rung trigger, not a
caveat.

The capability is SELF TIME -- inclusive minus the callees the same pass already
counted -- after which the column is additive and the sentence is derivable. It
is a TRACKED stall rather than a wish: CrossClaimFillGuard's Drop already
computes exactly that netting against the CROSS_CLAIM_FILL_FRAMES child stack,
one tier over, for the shared-fill ledger. What is missing is a child stack over
the recompute ledger's frames.

Deliberately not built here: this bridge is what stops the next lanes paying,
and a new measurement tier would put it behind a fresh review cycle on the very
fleet condition it exists to explain.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fjt9vSfVCKwZu9x7d44xZP
… is single-run

calm-boar-314 retracted the two-run reading this prose leaned on: a rerun
replaces a job's ATTEMPT while the logs API serves the latest attempt for a
fixed job id, so the six-row and thirteen-row readings addressed different
attempts through one identifier. No error is raised and nothing is visible from
the citing end. In the reproducible reading the two rows said to have dropped
out are in the set.

Replaced with evidence that needs no cross-run delta: two rows of one module
measured 502ms and EXACTLY 500ms against the 500ms budget. A row sitting at its
budget lands in interrupted-before-verdict or completed-over-cost according to
whether the poll fired before or after the work finished -- a fact about the
poll, not about the row. Same conclusion, one run, nothing to retract.

The class row now carries the retraction as part of the class, because it
sharpens the recommendation: a remedy must never be keyed to WHICH rows tipped.
A producer census is invariant under that retraction; a ranking of tipped rows
would have been built on it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fjt9vSfVCKwZu9x7d44xZP
…d independently

calm-boar-314 over-retracted and I cut a true finding on it. Attempt logs ARE
recoverable -- /actions/runs/<id>/attempts/<n>/logs -- and I pulled both
attempts of run 33620893203 myself rather than restoring on a second say-so:

  attempt 1: interrupted_before_verdict=6, and
    ..an_empty_receipt_series_leaves_the_live_tree_unmeasured_rather_than_held
    is INTERRUPTED-BEFORE-VERDICT at 502ms
  attempt 2: interrupted_before_verdict=13, and the SAME row is
    COMPLETED-OVER-COST-REQUIREMENT at cpu_ms=500 -- it reached its verdict

One row, one tree, crossing the two populations the 2026-08-19 cut separated
BECAUSE THEY ARE DIFFERENT FACTS WITH DIFFERENT REMEDIES. Which one it lands in
is decided by whether the poll fired before or after the work finished.

STILL WITHDRAWN: near-disjointness (attempt 2's set largely CONTAINS attempt
1's) and the 3-6-13 trend (the 3 was a different head). Only the migration
returns.

CITE THE ATTEMPT, NOT THE RUN, now recorded as part of the class: a bare run id
and a bare job id both resolve to the latest attempt, so a rerun changes what a
stable citation serves with no error and nothing visible from the citing end --
which is why the retraction looked sound. My own figures now carry
run_attempt=1 for runs 33615659836 and 33622954427, checked rather than assumed.

The census reported the same thing through the retraction AND the restoration,
because it keys producers and never which rows tipped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fjt9vSfVCKwZu9x7d44xZP
I wrote the restored migration as 'INTERRUPTED-BEFORE-VERDICT at 502ms'. The
attempt-1 diagnostic says the opposite in three clauses: cost=UNMEASURED, above
500ms with NO UPPER BOUND, and interrupt_point names where the POLL observed the
ceiling -- a property of the budget, not of the row.

The pair is a BOUND beside a VALUE, not two measurements of one quantity. As
502-against-500 it reads as a row getting two milliseconds cheaper and crossing
a line, which is a story about the row; what actually changed between attempts
is WHAT THE OBSERVER COULD SAY. That is the same point in its sharpest form, and
my number quietly converted it back into the weaker one.

It is also the class this census exists to stop: a transcribed instrument
property standing where a subject property belongs, inside the diff that files
it. Receipt (i) is left as figures because both those rows COMPLETED over budget
-- measured values, not preemption bounds -- and the prose now says so.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fjt9vSfVCKwZu9x7d44xZP
The only authored conflict was gunbc.recurring_failure_mode: main added three
class rows (ambient_process_state_read_by_a_concurrent_reader,
predicate_vacuously_true_on_an_empty_domain,
check_subject_narrower_than_its_declared_claim) at the same append point as
recurrence_ledger_scoped_below_the_recurrence. Resolved purely additively -- all
four rows and all four roster entries, source order preserved, 55 declarations
against 55 roster entries.

The three conflicted PROJECTIONS -- witnesses.yml, DESIGN.md and
docs/design-ledgers.md -- were not resolved by taking a side. They were
REGENERATED over the merged .dag authority through the generated-artifact gate,
so the committed bytes are derived rather than composed. That is what the
refusing merge driver exists to force: a text merge of a projection can produce
bytes no authority would emit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fjt9vSfVCKwZu9x7d44xZP
@gunbai-bot

gunbai-bot Bot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor Author

Ruling 3 answered with a measurement, and the floor ran green on this branch

Run 33631458679 (head eebba445e, run_attempt=1, attempt-pinned) — required-witnesses-floor: success:

required-floor: planned=3486 executed=3486 passed=3411 failed=0
                interrupted_before_verdict=0 completed_over_cost_requirement=0

Zero refusals on the same corpus that produced 9 interrupted rows two runs earlier, on a tree carrying this change. That is the account restated by the lane itself: membership in the refusing set is not a property of the rows.

The instrument's own cost, measured rather than argued

[cross-claim-demand] claims_absorbed=3486 retained_keys=26550 shared_keys=7932
                     omitted_under_floor=2756555 omitted_under_floor_ms=260864
                     key_cap_overflow=0
                     absorb_ms=1731 absorb_max_ms=17 claim_cpu_total_ms=99943
  • 1731ms total against a run whose entire claim-side CPU is 99,943ms — ~1.7%.
  • 17ms worst single claim, and it runs after run_claim_measured returns, so it is outside every claim's charged window and cannot enter a cost figure or trip the deadline.

A discovery instrument on a cost-preempting floor had to answer whether it recreates the incident it explains. It does not, and now the artifact says so with numbers instead of reasoning.

One more datum for the variance half, from my own ceiling column

claim_cpu_total_ms was 130,335ms on the earlier red run and 99,943ms here — the same corpus, a 30% swing in total claim CPU between runs. The per-closure constant this PR is about cannot produce that; it is the second variable, showing up in the very figure I added as a bound. Consistent with the byte-identical-eval_steps-against-1.31–1.80×-cpu measurement from the other lane, and with this PR's declared boundary that the census explains the level and not the variance.

— sent from nimble-lynx-128

claim_cpu_total_ms was added as a bound against summing a column that is
inclusive of callees. It caught the variance term instead: 130335ms on run
33615659836 and 99943ms on run 33631458679, both run_attempt=1, over the same
corpus. A per-closure constant cannot produce a 30% swing -- it is by
construction identical on two runs of one tree.

Unexplained, and deliberately not pursued here: this census's subject is the
level, and that boundary is declared in three carriers. But it is a property of
the ENVIRONMENT measured from inside the run, which is what a fleet-variance
account has lacked, and its only other home was two logs that age out.

The two figures are transcribed as a declared exception to naming the instrument
rather than copying its output, on the same ground the floor-cost carrier grants
it: the subject IS the divergence between two runs, and no producer on this side
of the boundary can re-derive it once the logs expire.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fjt9vSfVCKwZu9x7d44xZP
@gunbai-bot
gunbai-bot Bot merged commit 0abc7c3 into main Sep 2, 2026
6 checks passed
@gunbai-bot
gunbai-bot Bot deleted the session/nimble-lynx-128 branch September 2, 2026 14:56
@briansrls
briansrls restored the session/nimble-lynx-128 branch September 2, 2026 15:00
gunbai-bot Bot pushed a commit that referenced this pull request Sep 2, 2026
…lision)

AUTHORITY ONLY, same disposition as the previous union. Three more classes landed
on main since the last one -- #10006 (6764d17), #10044 (7cbadb5) and
#10053 (0abc7c3) -- each appending to dag/gunbc/recurring_failure_mode.dag,
so both the declaration block and the roster list conflicted again. Both regions
resolved by keeping BOTH sides, main's first and this branch's row last.

Checked as an IDENTITY JOIN rather than by counting the names this branch added:
58 declarations, 58 roster entries, 58 unique each, empty in both directions --
no roster entry without a declaration, no declaration without a roster entry. A
count equality would have passed even if one of each had drifted apart.

NO PROJECTION WAS HAND-RESOLVED. DESIGN.md and docs/design-ledgers.md are
byte-identical to origin/main here, verified rather than assumed. The probe
returned an unmerged AUTHORITY at stages 1, 2 and 3, not a
GeneratedArtifactConcurrentDivergence row: the driver refuses generated bytes so
that no human adjudicates them, and does not refuse source.

STILL DELIBERATELY INCOMPLETE. stable_citation_mutable_referent appears 0 times in
either projection, so the drift gate still refuses this branch, correctly. The
regen runs once, on the tip that will carry it.

Module compiles: 0 blocking errors, 95 advisories (all pre-existing
where-refinement rows in std.decl_ref).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FdxzwWekWhHR2FCTTf8a1b
gunbai-bot Bot pushed a commit that referenced this pull request Sep 2, 2026
…derived

Second additive roster merge on this branch in one session -- #10053 landed
`recurrence_ledger_scoped_below_the_recurrence` while this PR was waiting on its
own checks, so the merge failed on ORDERING rather than on readiness. Both sources
had agreed the PR was ready; main simply moved between the readiness call and the
land.

Only the roster LIST conflicted this time; the declarations auto-merged. Kept both
sides, main's row first.

VERIFIED BY IDENTITY JOIN AGAINST A POST-CONDITION STATED FIRST: main declared 57,
so 59 was owed. Measured declared 59, identity-fields 59, rostered 59, with
declared-not-rostered, rostered-not-declared and name-disagrees-with-its-own-
identity all empty. Equal counts alone would not have caught a severed unit, which
is the failure `merge_region_excludes_shared_tail` records for this exact carrier.

DESIGN.md and docs/design-ledgers.md were NOT hand-resolved. Both are derived; the
markers were discarded and both re-emitted by main_wet over the merged .dag, then
confirmed at the fixed point by a second run leaving nothing unstaged. All three
new rows -- main's one and my two -- verified present in the projection by content
rather than by the merge reporting success. Taking the ours side of a generated
file has silently dropped authority-derived bytes twice on this repo.

Merged, not rebased: the dashboard notice asks for a rebase and merge policy
forbids it on a squash-merge repo. Same divergence as the previous merge, flagged
for the same reason.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VSP89XiSm2YMnUvSwSR1ct
gunbai-bot Bot pushed a commit that referenced this pull request Sep 2, 2026
Main moved from 7cbadb5 to 0abc7c3 while the previous merge was being
resolved (#10053 landed), so this second merge picks up the row that briefly
looked like one this branch had dropped. It had not: the previous resolution
was correct relative to the MERGE_HEAD it actually merged, and the apparent
loss was main moving underneath it.

The authority auto-merged cleanly this time -- #10053 appends a new row rather
than editing the one both lanes had been appending to -- so only the generated
projection needed work, and it was regenerated rather than hand-resolved.

Verified on CONTENT, not exit code, because a previous round of this exact
regeneration exited non-zero while leaving a plausible file, and the tell was
the row count rather than the status: 57 declared, recurrence_ledger_scoped_below_the_recurrence
present, and both specimens on the shared row intact. The projection renders 40
bulleted rows, which matches main's own 40-against-57 ratio exactly, so the
standing 17-row rendering gap is unchanged and not introduced here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XMY7pX8yLX44MtbRPpFeuf
gunbai-bot Bot added a commit that referenced this pull request Sep 2, 2026
…e two sentences it makes false (#10078)

Part 2 of the operator-approved item, and the promotion is now literally one row
because #10036 made the lane axis a roster. Three edits, only one of which is the
promotion:
  one `RequiredLane` in `required_lanes_roster`
  one `required_lanes_gate_unit_var` beside its BUILD/FLOOR siblings, so the
    variable is a declared row rather than a bare literal at the use site
  the aggregate step's name, because "Both required lanes must have succeeded" is
    false with three

WHY THIS MATTERS AT ALL. The 774 `#[test]`s under src/v1/stage0 ran on no CI path
before that job existed, then ran on every push and pull request while GATING
NOTHING -- and the gap produced exactly the harm it predicts: #9886 landed two
failing tests on main with every required check green. That is the whole reason
this item exists.

THE EMITTED DELTA IS THE RECEIPT, and it is the same six surfaces PR A's forward
control predicted before any of this was written:
  needs: [...build, ...floor, rust-unit-tests]
  a third `|| [ "$UNIT" != success ]` conjunct in the unestablished fold
  a third `|| [ "$UNIT" = failure ]` conjunct in the red fold
  ` unit=$UNIT` in the receipt line and in BOTH refusal messages
  UNIT: ${{ needs['rust-unit-tests'].result }} in the step env
  the step name
A `needs`-only edit would have produced the first and last of those and NONE of
the middle -- the lane would have been waited on and still unable to fail the
gate. That failure mode is why PR A landed first.

THE ANNOTATION IS REWRITTEN, NOT APPENDED TO. It said "It is not a `needs` of the
aggregate, so it does not gate a merge yet", which this commit makes false, and
§4c forbids an annotation restating what the declaration no longer says. Verified
the rewrite added ZERO emitted bytes -- annotations are erased before emission, so
the YAML delta is unchanged by it and carries none of its text.

ALL THREE PRECONDITIONS DISCHARGED, NOT ARGUED AWAY, and the annotation now states
them as a RULE rather than as history:
  MAIN GREEN -- and this was not hypothetical. While the promotion was held,
  #10036 was blocked by shell_service_unmodeled_output_key_refuses, main's own
  defect, whose fix its author had already landed under another number. A
  promotion whose first act blocks every open PR on an already-fixed defect is a
  self-inflicted outage.
  FLEET HEALTHY -- a lane that cannot be delivered its admitted memory or
  toolchain produces reds carrying no information about the diff.
  COST ACCOUNTING UNDERSTOOD (#10053) -- the same argument one layer down.
The generalisation is in the annotation because a later reader will be tempted to
drop the third: A LANE MAY BE PROMOTED ONLY WHEN A RED IN IT DISCRIMINATES. Wall
clock is the cheap question; whether the lane's failures are ABOUT THE DIFF is the
load-bearing one.

COST ON THE CRITICAL PATH IS ZERO, by comparison rather than by bound: the lanes
run in parallel with the aggregate only waiting, and measured across 120 witnesses
runs this job sits BELOW required-witnesses-floor at every quantile. The
annotation names the producer to re-derive it and deliberately does NOT carry the
figures -- a timeout-headroom argument would have been the wrong one, since
headroom says nothing about what the aggregate waits for.


Claude-Session: https://claude.ai/code/session_01VSP89XiSm2YMnUvSwSR1ct

Co-authored-by: gunbc-ci-auto-heal <gunbc-ci-auto-heal@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants