Repository navigation
Whole-tree census that resolves per module: a third driver mode at occurrence grain (XL-2's third capability) - #11738
Conversation
…currence grain XL-2's rehearsal owes a failure roster captured AS the census, and on the v2 route a post-strip roster is an UNDERCOUNT BY CONSTRUCTION: the driver's census mode runs tokenize/parse/normalize per file and never calls resolve, so an absence of refusals establishes nothing about affected dependents. This lands the missing half. THREE FOLDS IN v2.compiler.compile, and one of them is a de-fork rather than an addition. native_module_resolve_verdict is factored OUT of native_lane_module_resolution so the adjudicating lane and the census share one authority for "did this module resolve in this context" -- a second function reaching native_test_resolve_module behind its own file-refusal precheck would be the DESIGN section 3 fork in the place it matters most, where the two spellings could disagree about whether a file the context refused is a module that resolved. The verdict carries NonEmptyDiagnostics UNFOLDED, because native_test_fatal_reason collapses to the LAST reason -- right for a member row whose grain is the declaration, wrong for a census whose grain is the occurrence. Beside it, native_census_modules rosters every ingested module (read from the source, not the context index: a file the context refused never reached absorb, so an index-derived roster would drop exactly the files a census exists to report), and native_census_residual_rows fans a refusal out to one row per diagnostic carrying std.diagnostic Diagnostic itself rather than a re-coined reason/locus pair. A THIRD MODE, NOT A FLAG ON census. That mode has a live consumer -- the malformed control, which the host runner spawns over a ONE-FILE scratch root -- and teaching it to resolve would make that control pay a whole-tree resolve to learn what the front end already answered. Two modes because they are two subjects. It is also not the adjudicate loop widened: that loop walks universe.modules over a context folded from the import CLOSURE of the v2.test.* universe (2215 of 6128 modules), and a residual over the closure of the test corpus is not a residual over the tree. WHERE IT LANDS, STATED RATHER THAN DEFAULTED. The mode is hand Rust in v1.compiler.emit_rust emit_source_root_eval_driver_main_rs. Moving the orchestration into .dag is that declaration's OWN dissolution trigger and a partial move does not retire it, so taking it here would have been the tempting smaller artifact the seed-growth row names. The enrolled SeedGrowthJustification reason is updated from two modes to three, since that row is the standing enumeration of what is retained and a mode absent from it is hand Rust no row accounts for. The decision surface did not grow: the roster, the verdict and the rows are .dag folds, and the mode is the same loop-and-print shape its two siblings already are. The resolved tree is dropped rather than kept -- NativeCensusModuleOutcome carries diagnostics on refusal and nothing on acceptance -- so the census holds one tree at a time instead of accumulating one per module across the corpus. Both new folds prepend and reverse once; appending per item is quadratic at 6128 elements, which DESIGN section 6 makes unconditional. EVIDENCE. Five claims over three supplied specimens run through the production folds -- native_test_context_from_ingest, then the verdict -- and the set is self-discriminating: if everything refused, the clean-module claim fails; if nothing refused, the unbound claim fails; and the poison specimen must come back file_refused and NOT resolve_refused, which is the control that stops a green on a cause never observed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Cost, measured — and the headline is that the new half is nearly freeThree subject sizes through the production route (the emitted binary,
Resolve — the half this PR adds — is ~0.68 ms/module, projecting to about 4 seconds over the whole tree. The cost is essentially all in the context fold, which is the tokenize/parse/normalize that A correction to my own arithmetic. From the two closely-spaced points I predicted ~21 min for the corpus. The 1423-file point falsifies that: per-file context cost rises with n (261 → 278 → 348 ms/file). A two-term fit gives a ≈ 0.269 s/file and b ≈ 5.5e-5 s/file², predicting ~3728 s (62 min) of context and ~4 s of resolve for 6130 files. The corpus run in flight is the adversary for that prediction and I will post what it actually returns, including if it misses. The located cost-shape observation, stated as an observation and not a diagnosis: the super-linear term is in the context fold, not in anything this PR adds. The driver template's own comment already flags Affordability read, which is a separate question from the number: for XL-2's rehearsal — a one-shot census over a stripped throwaway tree — an hour of operator-invoked wall is affordable, and it is the same hour a front-end-only census would have cost. It would not be affordable as anything recurring, and nothing here proposes that: the required-v2-native lane is deleted ( Memory is not a constraint: 1.3 GB at 1423 files, and the corpus run is sitting at 1.5 GB — far below the 14.8 GB the adjudicate baseline records. |
Whole-corpus measurement — and my own prediction falsified
The prediction I posted before this run was wrong and the run is what corrected it. I projected 3728 s (62 min) of context from a two-term fit; actual is 1757 s — 2.1× lower. The resolve projection (4.2 s vs 3.836 s actual) held. The error was methodological: I fitted a curve across three different subjects and read the variation as curvature in n, when it is composition — The headline stands and is stronger than expected: resolve is 3.836 s of a 1766 s run — 0.22% of the cost. The half this PR adds is free at whole-tree scale. Everything else is the tokenize/parse/normalize that What the census actually reports todayOnly 232 of 6130 modules resolve clean. That is the undercount the brief predicted this instrument would expose, and it is dominated by unbound symbols — consistent with the declaration-grafting gap #11574 is closing (nothing rehomes a Two findings for the consumers1. Occurrence ids are mixed: 1777 2. Affordability~29 minutes of operator-invoked wall and 1.73 GB for a one-shot rehearsal census over a stripped tree. Affordable, and materially cheaper than the 4h06m / 14.8 GB the adjudicate baseline records — because adjudicate's cost is infer plus per-declaration eval, which a census does not pay. Not affordable as anything recurring, and nothing here proposes that: the lane is deleted, so this runs only when a person asks. |
… paid at preparation
The required floor refused the four context-consuming claims of this file at
197116-200054 evaluator steps against the 72300-step new-witness budget (run
35464948266, job 105955463881, cause=completed_over_cost_requirement). The gate
was right and the defect was mine: each claim called a helper that re-entered
native_test_context_from_ingest, so every claim rebuilt the language model, the
grammar preparation and the parse of all three specimens in order to inspect ONE
module's outcome. That is DESIGN section 3's own tell -- a claim whose evaluation
is dominated by one call whose result shape is all it inspects -- and the cost
was never what the claims assert.
THE REPAIR IS TO STOP RECOMPUTING, NOT TO RAISE THE LINE, which is
v2.workflow.floor_pure_producer_share's own sentence. census_probe_outcomes is
nullary and pure, so it is preparation-forceable: enrolled as a WARM row the
floor evaluates it once during strict preparation and bills it there, and each
claim then reads a decided value. The precedent is exact --
v2.test.native_decl_selection drove this same producer over one synthetic ingest,
was refused on the same ground, and was repaired by the warm row
collision_resolved rather than by moving or excusing it.
THE SHARE POINT IS THE DECIDED OUTCOMES, NOT THE CONTEXT. NativeTestContext
carries a resolution context, which is the origin-bound shape that roster refuses
at publication; CensusProbeOutcomes carries three NativeCensusModuleOutcome
variants and an Int, and those carry only diagnostics -- symbols, loci and nodes
-- so the stored value is positively portable for the same reason
collision_resolved's ResolvedTree is.
VERIFIED BEFORE LANDING, because a warm row that fails to store STOPS THE LINE
for every lane and the roster's own header says claim_batch cannot verify one:
the warm path is forced at strict preparation, which a plain --entry run does not
execute. Confirmed by running the lane itself (claim_executor --required-ci
--required-lane witnesses) in-session, which is the instrument that roster names:
phase=pure-producer-share-warm producer=...census_probe_outcomes
disposition=Stored cpu_ms=598 provenance=built-by-preparation
[floor-shared-fill] cache=cross_claim_pure_share key=census_probe_outcomes
fill_ms=600 paid_by=<outside-fold> consumer_claims=4 disposition=outside-fold
and zero BLOCKING lines across the run, against eight on the refusing CI run.
WHAT WAS CONSIDERED AND REJECTED: moving the claims to a long home. Long modules
are DECLINED from the floor, so the discriminating control that stops this
instrument going green on a cause it never observed would stop being observed by
any gate -- silence, in DESIGN's own word for it. The claims stay on the floor and
still execute the real route; sharing moves WHERE the front end is paid, never
WHETHER it runs, so deleting the integration still reds every claim here.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The floor caught a real defect in my claims, and the fix is verified before landing
That is DESIGN §3's own tell: "a claim whose evaluation is dominated by ONE call whose RESULT SHAPE is all it inspects." Worth noting against the approving review above, which saw this and dismissed it under §2's "one undeclared pure demand may recompute" — correct about correctness, but the armed cost gate measures something the reviewer wasn't measuring, and it was right. The repair is the modeled one: stop recomputing, don't raise the line. Verified before landing, because this one can hurt other people. That roster says a warm row which fails to store stops the required-floor line for every lane, and its own header says Stored (so it reifies, no line stop), billed outside the fold, serving all four claims — and zero What I considered and rejected: moving the claims to a long home. That is the gate's other suggestion, and it is wrong here — long modules are declined from the floor, so the discriminating control that stops this instrument going green on a cause it never observed would stop being observed by any gate. DESIGN's word for that is silence. The claims stay on the floor and still execute the real route: sharing moves where the front end is paid, never whether it runs, so deleting the integration still reds every claim. |
One departure from the brief I had not stated, recorded nowRe-reading the request against the diff rather than against the approvals, I found an instruction I followed differently without saying so. §3b is explicit that the middle value — diverges with a stated reason — is fine, and that an unstated divergence is the only red, so this closes it. The brief said: report the residual "reusing the What I did: minted a new Why. The two are not the same subject, and the reuse became impossible once the grain question was answered. So this is one subject getting one surface, not two surfaces for one subject: The same reasoning is why this is a third driver mode rather than a flag on — sent from vivid-lark-741 |
The third of the three capabilities XL-2's rehearsal residual is blocked on, and the one with no owner. Shape is wise-pike-696's from the investigation behind #11735; two of its three stated constraints turned out to be false and are corrected below rather than inherited.
What was missing
XL-2 owes the cut executed in a throwaway tree with its failure roster captured AS the census. The strip half exists (
std.import strip_import_statements). The census half did not: the driver'scensusmode runs tokenize/parse/normalize per file and never calls resolve, so a post-strip roster of front-end refusals is an undercount by construction — an absence of refusals cannot establish completeness of affected dependents.What this lands
Three folds in
v2.compiler.compile, one of which is a de-fork rather than an addition:native_module_resolve_verdict— factored out ofnative_lane_module_resolutionso the adjudicating lane and the census share one authority for "did this module resolve in this context". A second function reachingnative_test_resolve_modulebehind its own file-refusal precheck would be the §3 fork in the place it matters most: the two spellings could disagree about whether a context-refused file is a module that resolved. It carriesNonEmptyDiagnosticsunfolded —native_test_fatal_reasoncollapses to the last reason, which is right for a member row at declaration grain and wrong for a census at occurrence grain.native_census_modules— every ingested module, read from the source rather than the context index. A file the context refused never reachedabsorb, so an index-derived roster would silently drop exactly the files the census exists to report.native_census_residual_rows— one row per diagnostic, carryingstd.diagnostic Diagnosticitself rather than a re-coined reason/locus pair.Plus a third driver mode,
census-resolve.A third mode, not a flag on
census. That mode has a live consumer — the malformed control, spawned by the host runner over a one-file scratch root — and teaching it to resolve would make that control pay a whole-tree resolve to learn what the front end already answered. It is also not theadjudicateloop widened: that loop walksuniverse.modulesover a context folded from the import closure ofv2.test.*(2215 of 6128 modules), and a residual over the closure of the test corpus is not a residual over the tree.Where it lands, stated rather than defaulted (the brief asked for this explicitly). The mode is hand Rust in
emit_source_root_eval_driver_main_rs. Moving the orchestration into.dagis that declaration's own dissolution trigger, and a partial move does not retire it — so taking it here would have been precisely the tempting smaller artifact the seed-growth row names. The enrolledSeedGrowthJustificationreason is updated from two modes to three. The decision surface did not grow: roster, verdict and rows are.dagfolds.Evidence
Five claims over three supplied specimens run through the production folds. The set is self-discriminating: if everything refused,
census_clean_module_resolvesfails; if nothing refused,census_unbound_module_is_resolve_refusedfails; and the poison specimen must come backfile_refusedand notresolve_refused— the control that stops a green on a cause never observed. All five pass.The emitted closure builds clean under
RUSTFLAGS=-D warnings(exit 0, 0 warnings) and the mode is live in the binary; the stage0 mirror regen reported exactly the expected drift (v1_compiler_emit_rust.rs) and is adopted here.Two constraints in the brief that were false
gunbc.rung_drop v2_native_route_off_the_merge_pathdeleted the job (DeletedWithoutReplacement). Nothing here runs on a merge candidate. Accepted by the brief's author and by the lane manager.source_root_for_storage_pathclassifies by a textualsrc/prefix, so an absolute scratch path makes every moduleDagTree, andcross_tree_edge_decisionthen returnsEdgeNormalResolutionwhere it would have denied. The requirement is production-relative spellings with cwd inside the tree. Recorded on XL-2's carrier by wise-pike-696.What is NOT established, stated as such
A three-arm differential (real tree / path-faithful copy / absolute-path copy) over a straddling subtree produced byte-identical rows in all three arms, so it did not discriminate and establishes nothing about path-faithfulness. The reason is precise:
cross_tree_edge_decisionbypasses toEdgeNormalResolutionwheneversource_root_index_lookupisSourceRootAbsent, and on a restricted root set most modules never reachabsorb. Arm A bypassed via absent-index and arm C via same-root collapse — two different bypasses, one output. A prediction I stated before the run (thatresolve_cross_tree_import_deniedwould appear in arm A) was falsified: it appears zero times.Occurrence ids in the residual are mixed — 69
OccurrenceMintedto 140OccurrenceSyntheticon that subject — which answers the caution wise-pike-696 attached to the grain: the residual is not uniformly joinable by occurrence id.Whole-corpus cost is being measured now and will be posted here with an affordability read; it is deliberately not estimated.
🤖 Generated with Claude Code