Repository navigation
The emit near-ceiling family is not a defect: file the negative result, and the two derivations it rests on - #10181
Merged
Conversation
…t, and the two derivations it rests on Three v2.test.emit identities refused the required floor and the question routed here was whether the family carries a repairable cost defect. It does not, and the useful part of that answer is the four plays that look obviously right from outside and are each already measured and dead: serving the shared producer across claim frames (three candidate rows, every one made its consumers WORSE), finding the one dominant shared producer (refuted twice, independently), per-producer repair (8ms against a question 87 rows deep), and reading an interrupted row's printed cost as a margin (it is the ceiling that interrupted it, not a cost). TWO DERIVATIONS, NOT PROSE, because the memo must cite a producer rather than carry figures. identity_inflation_permille_at_equal_work restricts the existing per-identity pairing to rows whose eval_steps are EQUAL across the two runs. That removes the work confound by construction instead of modelling it: the evaluator executed the same number of steps, so the only thing that could have changed did not, and the remaining cpu movement is inflation. It needs no control population, which is what distinguishes it from every normalisation this lane ran in gunbc#10094 -- two of which had to be withdrawn. Its discriminating red is the row whose steps MOVED: admitting it would report a work change as contention. crossing_recurrence and recurring_crossers join the per-run crossing sets across runs, which is the question that separates a property of the ROW from a property of the RUN. An identity crossing in many runs is expensive and tightening the line reaches it; identities that each cross once are being selected by the run. The empty answer is the finding, so it lands with a positive control -- without one, "no identity recurs" is indistinguishable from a join that never matches. FloorCostRow gains eval_steps, parsed from its own column, with a witness for that column specifically: the artifact carries wall, cpu, steps and a diagnostic threshold side by side and reading the wrong one produces a plausible number rather than an error. WHAT THE MEMO REFUSES TO CLAIM is as load-bearing as what it does. The eleven runs span tree vintages, so "no identity crossed twice" is confounded by repair. Equal steps is equal EVALUATOR work, not equal host work. And a 500/p90 safe-cost figure this lane circulated is recorded as WITHDRAWN with both of its errors named -- a p90 answers a per-row question while the ceiling refuses the whole run, and the stratum quoted stopped one bucket short of the population that can reach 500, where inflation rises again. Both ran in the unsafe direction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QWLoiyrKq3zNTiDNtrs9gC
… a positional read produced a plausible number, not an error MEASURED, NOT IMAGINED. Extending this parser to carry eval_steps at index 6, I ran the microseconds-per-step figure across eleven sampled runs and three of them came back near 2500 instead of near 2 -- three orders of magnitude off and perfectly well-formed. The cause is that `eval_steps` was INSERTED into this artifact between vintages: run 33661252708's header is `identity module outcome verdict_reached wall_ms cpu_ms cost_line_ms` and run 33699540412's is the same with `eval_steps` before `cost_line_ms`. A parser keyed on index 6 reads the older artifact's `cost_line_ms` -- the constant 100 -- as a step count. That is DESIGN section 3's standing rule applied to a TSV: a position is a second naming scheme for something the header already names, and any insertion above the cited index silently invalidates it. So the header is now the schema. `claim_cost_columns` resolves identity, module, cpu_ms and eval_steps by NAME and REFUSES when one is absent, and `parse_claim_cost_line` takes the resolved layout rather than assuming one. An artifact that does not carry a needed column yields no rows, which is the whole population refusing rather than a silently wrong one. THREE WITNESSES, and the third is the discriminating one. The old-vintage header must yield zero rows; the same shape WITH the column must parse, so the empty answer cannot be read as a broken parser; and a REORDERED header must be followed, which under the previous parser returns a row with cpu and steps swapped -- two plausible numbers rather than an error, which is the silent-wrongness shape this whole change is about. The old vintage's absence of the column is also why the eleven-run figures in the memo are the ones that do not need it: the run-factor and crossing-recurrence results read cpu and outcome, both of which sit at stable indices in both layouts and are now resolved by name regardless. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QWLoiyrKq3zNTiDNtrs9gC
…lways fixed, not priced per site Review 59079 flagged the nested filter as an efficiency observation and explicitly declined to request a change, because the population is per-run crossers and n is small. DESIGN section 6's standing rule says the opposite and is the authority here: a proven cost-shape defect -- the rule names a quadratic fold and a copied accumulator specifically -- is ALWAYS fixed, because "n is small here" is not a time-stable fact and pricing per-site exceptions is itself redundant work. The population is small today because the floor is healthy; a bad week is exactly when this runs on a large one. Both defects were present: the fold re-filtered the whole crossing list for every unseen id, and it appended with `concat`, which copies the accumulator each time. The ids are now sorted once and scanned once, runs of equal ids counted in a single pass, and the accumulator prepended and reversed at the end. The output order changes from first-appearance to identity order, which is the stronger property: a census whose order depends on which run happened to be sampled first cannot be diffed against another. Both recurrence witnesses and the two pre-existing crossing witnesses pass unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QWLoiyrKq3zNTiDNtrs9gC
…ose above the type Second grain of the same rule that refused this session's other PR. DESIGN section 4c models annotations at MODULE-ITEM grain only, so a `//` line on a record field refuses with `source annotation sits inside a declaration body` -- six parse errors, and the floor refuses the whole run on them. The prose is unchanged and now leads `type FloorCostRow`, naming eval_steps in its first clause so it still reads as being about that field rather than about the record. MY PRE-PUSH SCANNER MISSED IT BECAUSE IT ONLY KNEW THE FIRST GRAIN. It walked forward from each comment looking for a following item and reported clean, which is exactly true and exactly insufficient: a field comment HAS a following item. The check now tracks brace depth, because "attached to something" and "attached to a MODULE ITEM" are different questions and only the second is the rule. Five witnesses pass locally. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QWLoiyrKq3zNTiDNtrs9gC
gunbai-bot Bot
pushed a commit
that referenced
this pull request
Sep 3, 2026
…l calls the withdrawn rows measured-out Recomposed onto current main (db3caed), which moved through #10166's parser/occurrence-binding change and #10175's witness/heal machinery since this PR's last receipt. Compiler and resolution movement is material even with zero path overlap, so the prior exact-head CI no longer speaks for this tree. The recomposition also surfaces a contradiction that would have landed silently. #10181 landed a memo whose plays table says serving the shared producer across claim frames is "measured and refused", citing the same three rows this PR reclassifies as UNMEASURED. After a merge, main would carry both sentences about one subject -- a §3 meaning fork, and the more dangerous half is the memo, because a negative result is exactly the artifact a later lane cites to decide NOT to try something. Both statements were true of different serves, which is the whole point, so the repair is to say which serve each priced rather than to delete either: measured and refused AGAINST #10094's O(size) serve -- historical, still valid as that; this PR fires their re-enrol trigger, so their CURRENT state is UNMEASURED, neither admitted nor measured-out, pending a controlled present/absent pair nobody has run. That keeps the memo's conclusion (the emit near-ceiling family is not a defect) intact -- it does not rest on those three rows staying excluded -- while removing the sentence that would have let a future lane read a stale exclusion as a standing measurement. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019LhF5WCbZqrZHPqsnjpkYu
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A negative result, filed as a typed carrier plus the memo that cites it, so the next lane skips
four plays that do not work rather than re-deriving them.
The finding
Three
v2.test.emitidentities refused the required floor on run33695697323, routed to thislane as the charge-subject owner. The emit family is not a defect. Four rows, two clusters,
292–350ms off the floor, none individually remarkable; what put two of them over 500ms is the run
they landed in.
Two of the three were
budget_interrupted, so their cost was unbounded above — the artifactrecords 505/503 and 508/504, which is the interrupt point rather than a measurement. Bounding them
came first, because a row with no upper bound cannot be ranked, compared, or shown to have improved
by any repair. Off the floor, all four complete: 346 / 294 / 292 / 350ms.
They then split into two mechanisms rather than one family:
named_addis contention — 294ms against its sibling's 292ms off the floor, while on thefloor the sibling completed and it was interrupted. Three instruments that share no mechanism
agree: µs/step from the artifact, off-floor sibling parity, and the required multiple (~1.70x,
above the measured p90, and exactly one of the pair crossed).
two_targetsis bigger in work and ordinary in time — 188,416 steps at interrupt exceedevery sibling's completed total, yet it lands at 346ms beside a sibling at 350ms. A work bound is
not a cost property, and inferring one from the other is the error this row documents.
The two derivations
Prose was not an acceptable vessel for this, so the analysis lands in
gunbc.floor_cost_distribution— the existing authority for floor cost analysis — rather than asecond one.
identity_inflation_permille_at_equal_workrestricts the existing per-identity pairing to rowswhose
eval_stepsare equal across the two runs. That removes the work confound by constructioninstead of modelling it: the evaluator executed the same number of steps, so the only thing that
could have changed did not, and the remaining cpu movement is inflation. It needs no control
population — which is what distinguishes it from every normalisation this lane ran in #10094, two
of which had to be withdrawn. Its discriminating RED is the row whose steps moved: admitting it
would report a work change as contention.
crossing_recurrence/recurring_crossersjoin the per-run crossing sets across runs, whichis the question that separates a property of the ROW from a property of the RUN. An identity
crossing in many runs is expensive and tightening the line reaches it; identities that each cross
once are being selected by the run. The empty answer is the finding, so it lands with a positive
control — without one, "no identity recurs" is indistinguishable from a join that never matches.
FloorCostRowgainseval_stepsfrom its own column, with a witness for that column specifically:the artifact carries wall, cpu, steps and a diagnostic threshold side by side, and reading the wrong
one produces a plausible number rather than an error.
What the eleven-run join found
crossing_recurrenceover eleven floor runs: no identity crossed twice, and crossings clusterby family within a run (ten crossers in one run, all ten in one module). The per-run median factor
over 296 identities present in all eleven runs spans 0.858–1.154 — 1.35x best to worst on the
median row. So the run factor decides whether anything crosses; baseline cost decides which
family is exposed. The line adjudicates a product of two things and attributes it to one.
What the memo refuses to claim
As load-bearing as what it does claim:
confounded by repair — one crossing family was fixed between runs. Not a pure contention result.
reports few steps in both runs, and
required_floorcarries a declared rung drop for that region.500 / p90safe-cost figure this lane circulated is recorded as WITHDRAWN, with both errorsnamed: a p90 answers a per-row question while the ceiling refuses the whole run, and the stratum
quoted stopped one bucket short of the population that can actually reach 500 — where inflation
rises again rather than continuing to fall. Both ran in the unsafe direction.
Checks
Eight witnesses in
test.claim.floor_cost_distribution_witnesspass locally, including the four newones and the four pre-existing parser and pairing witnesses that the
eval_stepsfield change couldhave broken. Independent of #10141; no shared files.
🤖 Generated with Claude Code
https://claude.ai/code/session_01QWLoiyrKq3zNTiDNtrs9gC