Repository navigation
Root-cause and repair the 57s/1.22GB qualified-spelling claim: a whole-tree policy resolve, not name resolution - #9044
Conversation
…name resolution The floor line that opened this (qualified_spelling_takes_the_shared_layer, wall_ms=57337, +1.22GB, bare arm free) reads as "qualified-name resolution is expensive". It is not. Qualified resolution is a map_get. Measured, four orderings plus a discriminating control: a qualified PATTERN HEAD costs nothing, and the premium lands on whichever claim first compiles a qualified TYPE ANNOTATION -- once per process, then 5ms forever after. The mechanism is severity classification, not resolution. A qualified annotation's authored name misses env.source_visible_names (which carries the bare imported names), so the masked type-ref arm emits an advisory UnlistedImportUse. Classifying that one diagnostic calls compile_clean_unlisted_import_use_blocks_from_policy, which resolves and typechecks a whole separate entry closure over default_source_roots() -- the whole tree -- to evaluate one nullary Bool. A bare annotation emits no such diagnostic and skips it; a qualified pattern head never enters that arm. This is gunbc.ci_spec's already-documented cost B from 2026-07-25, whose dissolve-on (scope the policy resolve to the policy module's own import closure) was never discharged. That note measured 57.8s for a green compile paying this tail against the floor's 57337ms. Shape, answering the brief: constant per process, independent of the subject compiled, and corpus-denominated. Same probe, roots varied: 216ms warm against 44886ms when from_policy's whole-tree root set diverges from the caller's and pays a cold load -- 44.9s locally against 57.3s on the floor. Diagnosis only; no repair, and the 1.22GB is reported as consistent-with rather than measured, since the witness is quarantined and the floor line was not re-run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ng rather than widening Discharges the cost-B half of gunbc.ci_spec gunbc_ci_floor_batch_clamp_note's dissolve-on, which has named this exact repair since 2026-07-25 and was never landed. No second row is filed beside it: the obligation is discharged in place, because landing the fix while leaving its obligation open would be two authorities for one fact. compile_clean_unlisted_import_use_blocks_from_policy resolved and typechecked an entry closure over default_source_roots() -- the whole tree -- to evaluate one nullary Bool. That is not merely oversized (the policy module has three imports); it is a function answering a question about the caller's world by consulting a different one, which is why it presents as cost but is correctness-shaped, and why "n is small here" was never available. It now assembles the policy entry's own import closure and resolves that explicit source set. Every arm that cannot produce the exact closure REFUSES with a located message naming the module and the path. There is deliberately no whole-tree fallback: that arm would restore today's cost, zero the deficit's frequency by construction, and make the widening unrankable ever after (DESIGN section 5). It is a new closure builder rather than a reuse of resolve_virtual_source_with_imports because that BFS silently SKIPS an unresolvable import -- a silent skip here would answer the policy question from a graph missing the module the answer depends on. MEASURED, same probe and orderings, all probes PASS so the narrowed closure still returns the policy Bool: roots dag only C first qualified 44886ms -> 109ms roots dag + src/v2 C first qualified 216ms -> 151ms The qualified/bare asymmetry is gone rather than reduced: on the cold root set the qualified arm is now CHEAPER than the bare one. RESIDUE, reported rather than absorbed: the bare arm's 12967ms on the dag-only root set did not move (12881ms). The bare arm pays it too, so it was never part of the qualified asymmetry and this repair does not touch it. NOT UPGRADED: the floor's 1.22GB is still unattributed by execution. If the memory line survives this repair that is a second defect to find, not one to absorb here. v1 freeze admission: PURPOSE test (operator ruling 2026-08-20) -- a defect repair on a path the required floor executes every run, not growth on a v1 surface for v1's own sake. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
CI status: the The red is not mine. Run Main's own latest run ( What that costs, stated rather than glossed: the floor refused before the witness fold, so it never executed a single witness on this head. So CI has not exercised the changed path at all, and a green from it would not have been available to claim either way. The evidence for the repair is the remote probe runs recorded in the PR body — which do drive the real Re-running is pointless until #9031 lands; per the sequencing ruling it merges first, and the un-quarantine is a deliberate follow-up on its row rather than anything this PR should absorb. — sent from lively-bat-222 |
|
CI is green on the merge head. The earlier red was a timing artifact, now resolved at the root: #9031 merged at 20:26:49Z and my failing run had started at 20:20 — six minutes too early, against pre-merge main. Main has been merged in (clean, no conflicts) and the run is green.
What this green does and does not establish. It does establish the repair at floor scale: 10693 claims executed against the changed path — It does not re-measure the 57s/1.22GB line, and the tempting reading here is wrong. The summary shows Probe re-run against the shipped tree, since the pre-merge numbers were measured on a different tree than this PR now ships — reproduces to within 2ms, all PASS:
— sent from lively-bat-222 |
…he floor measures it again (#9065) #9044 landed the repair for the cost defect this witness exposed, so the quarantine row's dissolution condition is met on its first half. The row said it dissolves when the defect 'is root-caused and repaired, at which point the arm measures under gunbc_ci_fast_lane_eval_budget_ms' -- and deliberately NOT on a cost envelope, because an envelope large enough to admit 1.22GB and 57s would have admitted the defect rather than measured it. THE SECOND HALF IS UNMEASURED, AND THIS PR IS THE MEASUREMENT. The 412x figure (44886ms -> 109ms) comes from local instrumented probes against varied source roots, not from the floor: the witness has been excluded since before the repair existed, so no floor run has ever executed the repaired path. Restoring it is the only way to take that reading, and a red here is a legitimate outcome rather than a failure of this change -- it would mean the floor's subject differs from the probe's in a way that matters, which is worth knowing before the row disappears. Two things stay open on purpose and must not be read as settled by a green: The 1.22GB RSS was never attributed to the repaired term. Three agreeing wall figures and a plausible shape is consistent-with, not measured. If the memory line survives this run, that is a SECOND defect and it should stay visible rather than be credited to #9044. And a residue did not move: the bare arm measured 12967ms before the repair and 12881ms after, on the dag-only root set. It was never part of the qualified-annotation asymmetry -- a separate cold-index term for a root set that does not match the process's warm index. Restores the file to dag/test/claim/, its module path, and the provider fixture's pointer, which named the long/ home. Co-authored-by: Brian Searls <briansearls1@gmail.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ed overrun are not one state (#9073) * Split the budget outcome into two arms: an interruption and a completed overrun are not one state The floor's ledger described one row three inconsistent ways -- BUDGET-REFUSED on the per-row line, outcome=timed_out on the over-cost line, and "cost exactly 57193ms" on the gating line. Those cannot all be true of one claim. They came from one habit. `ClaimOutcome::TimedOut` carried the discriminator as a field (`completion: Interrupted | CompletedOverBudget`), and five projections over it wrote `TimedOut { .. }`, giving the interrupted row's label to a claim that had reached its verdict. Four were in the seed; the fifth was the .dag AUTHORITY -- `v2.workflow.floor_terminal_ledger` `claim_disposition` wildcarding `ClaimSafetyOutcome` and answering `BudgetRefusedBeforeVerdict`, on a type whose own model already separates `SafetyInterrupted` from `CompletedPastSafetyLimit`. THE MODEL HAD THE DISTINCTION; EVERY PROJECTION OVER IT DROPPED IT. Its wire renderer dropped it a sixth time. The axis is now the arm: `BudgetInterrupted { elapsed_at_least_ms, .. }` and `CompletedOverBudget { elapsed_ms, .. }`. A labelling consumer must name both or fail to compile, and the field names carry the reading so a bound cannot be spelled as a cost. `BudgetCompletion` is deliberately NOT deleted -- it is the axis as data where a value has to travel (`BudgetRefusal`, the wire verdict, the `budget_event` projection), and deleting it there would force a second spelling of the same two states beside it. This is the seed comment's own declared next-rung trigger firing: it kept the field, called the wildcard hazard "paid for by review", and named THE NEXT CONSUMING SITE ADDED THAT DROPS THE AXIS as the evidence that review is not the mechanism. Five sites did. The prior decision is not overruled -- it argued against splitting on RAISE MECHANISM, which no consumer should act on; this splits on PASSED-VERSUS-INTERRUPTED, which that same comment calls a real distinction because it determines the remedy. Evidence, enrolled permanently rather than as scaffolding (DESIGN 4b(4)): `the_two_budget_arms_project_apart_in_every_consumer` asserts all four seed projections disagree across the two arms -- each pair was EQUAL before the split -- and the wire round-trip's every-shape population grows 9 -> 12 with the three shapes a wildcarded `safety` could not tell apart. This fixes the INSTRUMENT, not the cost. The 57s row is #9044's, already merged; keeping them separate is what lets a reader tell which change moved which number. * The contradictory-pairing arm refuses the ledger instead of riding inside a green one Operator ruling, 2026-08-24, on the one call I flagged as mine alone. KEEP THE NAME. Folding `BudgetRefused { safety: CompletedWithinSafetyLimits }` into either real arm publishes a refusal or a completion that nothing observed -- the fabricated plausible output, committed by the function whose purpose is to stop committing it. DO NOT MAKE IT UNCONSTRUCTIBLE EITHER, and the reason is not cost. `ClaimSafetyOutcome` is owned by `v2.workflow.required_floor` and reused here deliberately; a narrowed two-arm twin declared beside it is one concept forked by rigor -- §3 nicknaming -- and a fork gets consolidated later at someone else's cost. So the class sits at MITIGATABLE with its next-rung trigger written on the carrier (§4b(2), no untracked stall): the safety axis declared as a family in its OWNING module, at which point this arm, its refusal and their evidence dissolve together. WHAT WAS ACTUALLY BROKEN: nothing happened when the arm fired. It appeared twice -- declaration and produce site -- and no consumer read it, so `ledger_green_admission` judged publication, reconciliation and footer count and admitted a ledger carrying it as GREEN. That is a name that describes a contradiction and permits it: the deficit's frequency is unobservable by construction, which is the absorbing fallback's quiet half. It now REFUSES, and the refusal is LOCATED rather than merely typed -- `EvidenceCarriesIncoherentRow` carries the offending identities, so a reader gets rows to look at instead of a count to hunt with. Checked ahead of completeness and the footer, because those ask whether the right rows are present and this asks whether a row says anything at all; calling a population complete while a member carries no readable outcome answers a question the count cannot reach. The check folds over the SAME `claim_disposition` every other consumer reads, so it cannot become a second authority over the fact the disposition already answers. Evidence: `a_row_carrying_a_contradictory_budget_terminal_refuses_the_ledger`, with the positive control beside it -- the identical ledger with a coherent row IS complete, without which an admission that refused everything would also pass. All 20 tests across both ledger modules return true by execution. * Hoist the annotations out of declaration bodies: only module-item grain is modeled CI refused the floor at `phase=strict-preparation` with 78 diagnostics, one per line, all of them mine: source annotation sits inside a declaration body. Only module-item grain is modeled; move it above the declaration it describes. DESIGN §4c admits only STANDALONE LEADING `//` blocks attached to MODULE-SCOPE declarations. I put prose where it read best -- inside `type` coproduct arms, inside `fn` bodies, inside `match` arms -- rather than where the grammar admits it. Every block is hoisted above the `type`/`fn` it describes and re-worded to NAME its subject ("the `BudgetRefused` arm below", "two of the arms below"), because a hoisted annotation is no longer adjacent to the thing it explains and a reader cannot rely on position to tell them what it is about. Nothing is dropped: the §4b rung and next-rung trigger, the §3 argument for why narrowing `BudgetRefused`'s carrier is refused rather than deferred, the wire-tag finding, and the refuse-rather-than-count reasoning all survive at module scope. WHY THIS REACHED CI AT ALL, which is the part worth recording. `gunbc run` does NOT enforce annotation grain, and neither does `gunbc compile`: the check lives on the MODULE-INDEX path (`via_index_source_annotation_diagnostics`), so a tree full of body-grain annotations evaluates witness-by-witness perfectly green -- I reported 20/20 true and a mutation receipt on exactly this tree -- and then refuses the whole floor before a single witness runs. Green by execution on the lenient path was not evidence about the strict one, and I presented it as though it were. The first verification of this fix was thrown away because it failed its own control: `gunbc compile` over a fixture carrying a deliberate in-body annotation reported ZERO refusals, so its zero over the real tree measured nothing. An absence claim whose known-positive is dead is not evidence. --------- Co-authored-by: Brian Searls <briansearls1@gmail.com>
…9177) `docs/probes/qualified_name_resolution_cost_2026-08-23.md` is cited by src/v1/stage0/src/cli_run.rs and dag/gunbc/ci_spec.dag. It is absent from main and has never been present. It was authored in #9044, whose squash produced an EMPTY COMMIT because the code had already landed via #9025 five minutes earlier; the document was the only part of that PR not carried by the earlier one, so it was dropped in silence. This is the DESIGN section 3 stale-citation class in its worst form, and the reason it is worth its own change rather than a passing fix: a rotted line offset at least pointed somewhere once. This path never resolved for any reader, at any commit. No authoring-time care would have caught it, and the cited-symbol census could not either -- that lens checks SYMBOLS and this is a PATH. The cli_run.rs citer is load-bearing prose. It backs the 216ms/44886ms figures that justify the fail-closed scoped-closure design over a whole-tree resolve, so a reader following it to check the measurement found nothing. REPAIRED BY REMOVING THE DEPENDENCY, NOT BY RESTORING THE DOCUMENT. Both sites keep the figures they already carried inline, state plainly that the method document never landed, and point at refs/pull/9044/head, which exists and is stable. That leaves nothing in-tree that can rot. The document is deliberately NOT restored. docs/probes was BANKRUPTED on 2026-08-24 by d3bebd0, which deleted every transcription and left the directory at two files; re-landing a probe document there would re-open a corpus this repository closed the day before, and would strand 200 lines nothing cites. Verified disjoint from the live doc_graph_roots cleanup: this path is not in that population, and the two citers here are code, not documents. Comment- and note-only. No semantics, no generated-artifact drift (gate run, only these two files modified). The cli_run.rs half is a v1-seed edit and is admitted under the seed's maintenance arm as a defect repair -- it changes prose that was false, not behaviour. Co-authored-by: gunbc-ci-auto-heal <gunbc-ci-auto-heal@users.noreply.github.com>
Root-causes the cost defect behind the floor line
Now diagnosis AND repair (operator ruling, this session — the diagnosis landed first and the
repair was ruled admissible after). The witness is not touched. Document:
docs/probes/qualified_name_resolution_cost_2026-08-23.md.v1 freeze admission test being passed: PURPOSE, not shape (operator ruling 2026-08-20, "anything
in support of v2 self host is safe", alongside the standing "it is pretty frozen but fixing issues is
fine"). This is a defect repair on a path the required floor executes on every run — the gate the
self-host program is measured through — not growth on a v1 surface for v1's own sake, which is what
the freeze actually closes.
Prior authority discharged, not duplicated:
gunbc.ci_specgunbc_ci_floor_batch_clamp_notehas named this exact repair in its
dissolve-onsince 2026-07-25. That obligation is discharged inplace rather than filed as a second row beside it.
The headline
The name says "qualified-name resolution cost". Qualified-name resolution is a
map_get, and it isnot the cost. The qualified pattern head is exonerated by execution, and the real payer is the
severity classification of an advisory diagnostic that only the qualified spelling produces.
Why the built-in control was misleading
The qualified and bare arms differ in two ways at once — the head spelling and the type
annotation spelling — so neither witness could tell which axis pays. A five-probe factorial over one
fixture (identical imports; head and annotation varied independently) separates them, run in four
orderings because the first run showed the premium on whichever probe ran first:
cpu ms per claim; bold = premium payer. D (qualified head) is 5ms whenever it is not first and
44ms when it is — the same as trivial E costs when it leads, i.e. generic warmup. Running D first
does not discharge the premium; C still pays 201ms after it.
The mechanism
A whole-tree resolve + typecheck of a separate entry closure, to compute one nullary
Bool, cachedper thread — so paid once per process by whichever claim first emits
UnlistedImportUse, andindependent of what was compiled.
Why only the qualified spelling.
UnlistedImportUseis emitted by the masked type-ref arm ofresolve_node_boundedwhen the authored name missesenv.source_visible_names. Importscontribute bare names, so
test.fixture.qualpat_provider.QualpatResultmisses and a bareannotation hits. A qualified pattern head never enters that arm — it goes through
lookup_variant_in_type/symbol_index_lookup. That is the whole asymmetry.Discriminating control: order B,D (both bare annotations, D carrying the qualified head) — 72ms
/ 8ms, and
compile_clean_diagnostic_policy.dagis absent fromspan_nanos_by_entry. In bothqualified-annotation orderings it appears exactly once, at 259ms / 269ms — the premium to within
~10ms. The axis switches the resolve on and off.
It is already documented, and never fixed
gunbc.ci_specgunbc_ci_floor_batch_clamp_note, 2026-07-25, cost B: this exact function,default_source_roots()whole tree, "~34-42s, once per process, and INDEPENDENT of what wascompiled", with a green compile measured at 57.8s total against this floor line's 57337ms.
Its
dissolve-onalready names the repair — scope the policy resolve to the policy module's ownimport closure — and was never discharged. The defect never left.
Shape (the brief's question 2)
Not superlinear in the witness: constant per process, unaffected by subject size; a second qualified
compile costs 5ms. But corpus-denominated, because the root set is the whole tree. Same probe,
same process, only the roots varied:
dag+src/v2dagonly~208×, because
from_policyasks fordefault_source_roots()regardless of what the process wasgiven, so its resolve is a different key from the warm shared index and pays a cold whole-tree
load. 44.9s locally against 57.3s on the floor — same order, same mechanism. That closes the
magnitude gap.
Scheduled for deletion? (question 3)
Partly, and it moves ownership rather than the verdict.
cli_run.rsdies wholesale withintegration/cli-run-cut, but no bounded event retires it: v1 is semantics-frozen /maintenance-active with no cutover date, and the required floor runs this path every run today.
The repair
compile_clean_unlisted_import_use_blocks_from_policyno longer callsdefault_source_roots(). Itassembles the policy entry's own import closure and resolves that explicit source set.
Every failure arm refuses; none widens. An import naming no module, or an unreadable file,
returns a located
Errnaming module and path. No whole-tree fallback — that arm would restoretoday's cost, zero the deficit's frequency by construction, and make the widening unrankable ever
after. It is a new closure builder rather than a reuse of
resolve_virtual_source_with_importsbecause that BFS silently skips an unresolvable import, which here would answer the policy
question from a graph missing the module the answer depends on.
Re-measured (all probes PASS — the narrowed closure still returns the policy
Bool)dagonlydagonlydag+src/v2The asymmetry is gone rather than reduced: on the cold root set the qualified arm is now
cheaper than the bare one.
Residue, reported rather than absorbed
Bdid not move (12967 → 12881ms). The bare arm pays it too, so it was never part of thequalified asymmetry and this repair does not touch it — a separate cold-index term for a root set
that does not match the process's warm index. Named here rather than folded into this result.
Un-quarantine: deliberately NOT in this PR
Superseded, and corrected in place rather than left standing. An earlier revision of this
section said the quarantine was not on main. That was true when written and is now false: #9031
merged at 20:26:49Z carrying it, and the witness now sits at
dag/test/claim/long/qualified_spelling_identity_witness_test.dag.The un-quarantine is still not done here, by ruling: the row's deletion should cite a merged
fix rather than a concurrent one, so it is a follow-up on that row after this PR lands — keeping the
repair and the row separate.
What is NOT claimed
That the 1.22GB is entirely this term. The floor line was not re-run — the witness is quarantined
as of #9031 and the offline recipe is what ran here. Wall figures agree three ways (44.9s local
cold, 57.3s floor, 57.8s in the 2026-07-25 note for exactly this tail) and the RSS shape is
consistent with a whole-tree entry resolve, but consistent-with is not measured and the document
does not upgrade it.
The probe module is embedded in the document with its re-run recipe rather than committed: it has
no consumer, so as a tracked
.dagit would be experimental residue.🤖 Generated with Claude Code