Skip to content

Whole-tree census that resolves per module: a third driver mode at occurrence grain (XL-2's third capability) - #11738

Merged
gunbai-bot[bot] merged 2 commits into
mainfrom
session/vivid-lark-741
Sep 20, 2026
Merged

gunbai-bot[bot] merged 2 commits into
mainfrom
session/vivid-lark-741

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Sep 19, 2026

Copy link
Copy Markdown
Contributor

The third of the three capabilities XL-2's rehearsal residual is blocked on, and the one with no owner. Shape is wise-pike-696's from the investigation behind #11735; two of its three stated constraints turned out to be false and are corrected below rather than inherited.

What was missing

XL-2 owes the cut executed in a throwaway tree with its failure roster captured AS the census. The strip half exists (std.import strip_import_statements). The census half did not: the driver's census mode runs tokenize/parse/normalize per file and never calls resolve, so a post-strip roster of front-end refusals is an undercount by construction — an absence of refusals cannot establish completeness of affected dependents.

What this lands

Three folds in v2.compiler.compile, one of which is a de-fork rather than an addition:

  • native_module_resolve_verdict — factored out of native_lane_module_resolution so the adjudicating lane and the census share one authority for "did this module resolve in this context". A second function reaching native_test_resolve_module behind its own file-refusal precheck would be the §3 fork in the place it matters most: the two spellings could disagree about whether a context-refused file is a module that resolved. It carries NonEmptyDiagnostics unfolded — native_test_fatal_reason collapses to the last reason, which is right for a member row at declaration grain and wrong for a census at occurrence grain.
  • native_census_modules — every ingested module, read from the source rather than the context index. A file the context refused never reached absorb, so an index-derived roster would silently drop exactly the files the census exists to report.
  • native_census_residual_rows — one row per diagnostic, carrying std.diagnostic Diagnostic itself rather than a re-coined reason/locus pair.

Plus a third driver mode, census-resolve.

A third mode, not a flag on census. That mode has a live consumer — the malformed control, spawned by the host runner over a one-file scratch root — and teaching it to resolve would make that control pay a whole-tree resolve to learn what the front end already answered. It is also not the adjudicate loop widened: that loop walks universe.modules over a context folded from the import closure of v2.test.* (2215 of 6128 modules), and a residual over the closure of the test corpus is not a residual over the tree.

Where it lands, stated rather than defaulted (the brief asked for this explicitly). The mode is hand Rust in emit_source_root_eval_driver_main_rs. Moving the orchestration into .dag is that declaration's own dissolution trigger, and a partial move does not retire it — so taking it here would have been precisely the tempting smaller artifact the seed-growth row names. The enrolled SeedGrowthJustification reason is updated from two modes to three. The decision surface did not grow: roster, verdict and rows are .dag folds.

Evidence

Five claims over three supplied specimens run through the production folds. The set is self-discriminating: if everything refused, census_clean_module_resolves fails; if nothing refused, census_unbound_module_is_resolve_refused fails; and the poison specimen must come back file_refused and not resolve_refused — the control that stops a green on a cause never observed. All five pass.

The emitted closure builds clean under RUSTFLAGS=-D warnings (exit 0, 0 warnings) and the mode is live in the binary; the stage0 mirror regen reported exactly the expected drift (v1_compiler_emit_rust.rs) and is adopted here.

Two constraints in the brief that were false

  1. "It executes on the required-v2-native lane… the floor is 46–54 min and this will be refused on that ground." The lane does not execute at all: gunbc.rung_drop v2_native_route_off_the_merge_path deleted the job (DeletedWithoutReplacement). Nothing here runs on a merge candidate. Accepted by the brief's author and by the lane manager.
  2. The throwaway-tree seam ("a stripped copy in a scratch worktree passed as roots should be enough"). The mechanism: source_root_for_storage_path classifies by a textual src/ prefix, so an absolute scratch path makes every module DagTree, and cross_tree_edge_decision then returns EdgeNormalResolution where it would have denied. The requirement is production-relative spellings with cwd inside the tree. Recorded on XL-2's carrier by wise-pike-696.

What is NOT established, stated as such

A three-arm differential (real tree / path-faithful copy / absolute-path copy) over a straddling subtree produced byte-identical rows in all three arms, so it did not discriminate and establishes nothing about path-faithfulness. The reason is precise: cross_tree_edge_decision bypasses to EdgeNormalResolution whenever source_root_index_lookup is SourceRootAbsent, and on a restricted root set most modules never reach absorb. Arm A bypassed via absent-index and arm C via same-root collapse — two different bypasses, one output. A prediction I stated before the run (that resolve_cross_tree_import_denied would appear in arm A) was falsified: it appears zero times.

Occurrence ids in the residual are mixed — 69 OccurrenceMinted to 140 OccurrenceSynthetic on that subject — which answers the caution wise-pike-696 attached to the grain: the residual is not uniformly joinable by occurrence id.

Whole-corpus cost is being measured now and will be posted here with an affordability read; it is deliberately not estimated.

🤖 Generated with Claude Code

…currence grain

XL-2's rehearsal owes a failure roster captured AS the census, and on the v2
route a post-strip roster is an UNDERCOUNT BY CONSTRUCTION: the driver's census
mode runs tokenize/parse/normalize per file and never calls resolve, so an
absence of refusals establishes nothing about affected dependents. This lands
the missing half.

THREE FOLDS IN v2.compiler.compile, and one of them is a de-fork rather than an
addition. native_module_resolve_verdict is factored OUT of
native_lane_module_resolution so the adjudicating lane and the census share one
authority for "did this module resolve in this context" -- a second function
reaching native_test_resolve_module behind its own file-refusal precheck would
be the DESIGN section 3 fork in the place it matters most, where the two
spellings could disagree about whether a file the context refused is a module
that resolved. The verdict carries NonEmptyDiagnostics UNFOLDED, because
native_test_fatal_reason collapses to the LAST reason -- right for a member row
whose grain is the declaration, wrong for a census whose grain is the
occurrence. Beside it, native_census_modules rosters every ingested module
(read from the source, not the context index: a file the context refused never
reached absorb, so an index-derived roster would drop exactly the files a
census exists to report), and native_census_residual_rows fans a refusal out to
one row per diagnostic carrying std.diagnostic Diagnostic itself rather than a
re-coined reason/locus pair.

A THIRD MODE, NOT A FLAG ON census. That mode has a live consumer -- the
malformed control, which the host runner spawns over a ONE-FILE scratch root --
and teaching it to resolve would make that control pay a whole-tree resolve to
learn what the front end already answered. Two modes because they are two
subjects. It is also not the adjudicate loop widened: that loop walks
universe.modules over a context folded from the import CLOSURE of the v2.test.*
universe (2215 of 6128 modules), and a residual over the closure of the test
corpus is not a residual over the tree.

WHERE IT LANDS, STATED RATHER THAN DEFAULTED. The mode is hand Rust in
v1.compiler.emit_rust emit_source_root_eval_driver_main_rs. Moving the
orchestration into .dag is that declaration's OWN dissolution trigger and a
partial move does not retire it, so taking it here would have been the tempting
smaller artifact the seed-growth row names. The enrolled
SeedGrowthJustification reason is updated from two modes to three, since that
row is the standing enumeration of what is retained and a mode absent from it
is hand Rust no row accounts for. The decision surface did not grow: the
roster, the verdict and the rows are .dag folds, and the mode is the same
loop-and-print shape its two siblings already are.

The resolved tree is dropped rather than kept -- NativeCensusModuleOutcome
carries diagnostics on refusal and nothing on acceptance -- so the census holds
one tree at a time instead of accumulating one per module across the corpus.
Both new folds prepend and reverse once; appending per item is quadratic at
6128 elements, which DESIGN section 6 makes unconditional.

EVIDENCE. Five claims over three supplied specimens run through the production
folds -- native_test_context_from_ingest, then the verdict -- and the set is
self-discriminating: if everything refused, the clean-module claim fails; if
nothing refused, the unbound claim fails; and the poison specimen must come
back file_refused and NOT resolve_refused, which is the control that stops a
green on a cause never observed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@gunbai-bot

gunbai-bot Bot commented Sep 19, 2026

Copy link
Copy Markdown
Contributor Author

Cost, measured — and the headline is that the new half is nearly free

Three subject sizes through the production route (the emitted binary, census-resolve, real files), before the whole-corpus run lands:

subject files context resolve (total) peak RSS
dag/std 154 42.8 s 0.17 s 159 MB
dag/test/fixture + src/v2/std 205 53.6 s 0.25 s 1.23 GB
src/v2 1423 495.3 s 0.96 s 1.30 GB

Resolve — the half this PR adds — is ~0.68 ms/module, projecting to about 4 seconds over the whole tree. The cost is essentially all in the context fold, which is the tokenize/parse/normalize that census mode already performs today. So census-resolve costs approximately what census costs; the capability XL-2 was blocked on is not what makes a whole-tree census expensive.

A correction to my own arithmetic. From the two closely-spaced points I predicted ~21 min for the corpus. The 1423-file point falsifies that: per-file context cost rises with n (261 → 278 → 348 ms/file). A two-term fit gives a ≈ 0.269 s/file and b ≈ 5.5e-5 s/file², predicting ~3728 s (62 min) of context and ~4 s of resolve for 6130 files. The corpus run in flight is the adversary for that prediction and I will post what it actually returns, including if it misses.

The located cost-shape observation, stated as an observation and not a diagnosis: the super-linear term is in the context fold, not in anything this PR adds. The driver template's own comment already flags native_test_context_absorb as the place "a quadratic accumulator would live … and nowhere else in this loop". That is pre-existing and outside this change's subject; I am naming it rather than fixing it here, and it is the thing to attack if whole-tree census cost ever needs to come down.

Affordability read, which is a separate question from the number: for XL-2's rehearsal — a one-shot census over a stripped throwaway tree — an hour of operator-invoked wall is affordable, and it is the same hour a front-end-only census would have cost. It would not be affordable as anything recurring, and nothing here proposes that: the required-v2-native lane is deleted (gunbc.rung_drop v2_native_route_off_the_merge_path), so this executes only when a person asks for it.

Memory is not a constraint: 1.3 GB at 1423 files, and the corpus run is sitting at 1.5 GB — far below the 14.8 GB the adjudicate baseline records.

@gunbai-bot

gunbai-bot Bot commented Sep 19, 2026

Copy link
Copy Markdown
Contributor Author

Whole-corpus measurement — and my own prediction falsified

census-resolve dag src/v2, emitted binary f7eeee6e, in-session (cgroup-bound, 125 GiB):

wall 1766 s (29.4 min)   context 1757.4 s   resolve_total 3.836 s   roster 0.39 s
modules 6130   corpus_reads 6130   peak_rss 1.73 GB
resolved 232   file_refused 797   resolve_refused 5101   residual_rows 5101
resolve p50 0.32 ms   p95 1.94 ms   max 119 ms

The prediction I posted before this run was wrong and the run is what corrected it. I projected 3728 s (62 min) of context from a two-term fit; actual is 1757 s — 2.1× lower. The resolve projection (4.2 s vs 3.836 s actual) held. The error was methodological: I fitted a curve across three different subjects and read the variation as curvature in n, when it is composition — src/v2 costs 348 ms/file because its files are denser, while the whole tree averages 287 ms/file. Per-file cost is not monotone in n, so there is no super-linear term to attribute, and my earlier remark pointing at native_test_context_absorb as a suspected quadratic accumulator is withdrawn — this data does not support it.

The headline stands and is stronger than expected: resolve is 3.836 s of a 1766 s run — 0.22% of the cost. The half this PR adds is free at whole-tree scale. Everything else is the tokenize/parse/normalize that census mode already performs today, so census-resolve costs what census costs.

What the census actually reports today

resolve_reason_unbound_symbol      4876
resolve_reason_ambiguous_symbol     224
float_literal_not_lowered             1

Only 232 of 6130 modules resolve clean. That is the undercount the brief predicted this instrument would expose, and it is dominated by unbound symbols — consistent with the declaration-grafting gap #11574 is closing (nothing rehomes a type_decl or data_decl). The census does not repair it; it counts it, at occurrence grain, which is what XL-5 consumes.

Two findings for the consumers

1. Occurrence ids are mixed: 1777 OccurrenceMinted to 3644 OccurrenceSynthetic (33% / 67%). This answers the caution attached to the grain: the residual is not uniformly joinable by occurrence id. Two thirds of the rows carry a synthetic id, so an XL-5 wave keyed on occurrence identity would silently cover only a third of its population. The rows remain joinable by module + path + locus structure, and that is what a wave should key on until minting is total.

2. resolve_cross_tree_import_denied appears ZERO times across the whole tree. The cross-tree denial arm — the mechanism behind the throwaway-tree path-faithfulness hazard — does not fire anywhere in the current corpus, because 5101 modules refuse earlier on unbound symbols and never reach import admission. So the hazard is real in mechanism and currently unobservable in outcome: a census taken today through an absolute scratch path would return the same rows. The requirement to preserve production-relative spellings still stands (it costs a worktree checkout and it stops being free the moment the unbound residual closes), but it should be recorded as a precaution whose discriminator is currently dead, not as a demonstrated defect. This also explains why my three-arm differential could not discriminate: there was nothing there to discriminate.

Affordability

~29 minutes of operator-invoked wall and 1.73 GB for a one-shot rehearsal census over a stripped tree. Affordable, and materially cheaper than the 4h06m / 14.8 GB the adjudicate baseline records — because adjudicate's cost is infer plus per-declaration eval, which a census does not pay. Not affordable as anything recurring, and nothing here proposes that: the lane is deleted, so this runs only when a person asks.

… paid at preparation

The required floor refused the four context-consuming claims of this file at
197116-200054 evaluator steps against the 72300-step new-witness budget (run
35464948266, job 105955463881, cause=completed_over_cost_requirement). The gate
was right and the defect was mine: each claim called a helper that re-entered
native_test_context_from_ingest, so every claim rebuilt the language model, the
grammar preparation and the parse of all three specimens in order to inspect ONE
module's outcome. That is DESIGN section 3's own tell -- a claim whose evaluation
is dominated by one call whose result shape is all it inspects -- and the cost
was never what the claims assert.

THE REPAIR IS TO STOP RECOMPUTING, NOT TO RAISE THE LINE, which is
v2.workflow.floor_pure_producer_share's own sentence. census_probe_outcomes is
nullary and pure, so it is preparation-forceable: enrolled as a WARM row the
floor evaluates it once during strict preparation and bills it there, and each
claim then reads a decided value. The precedent is exact --
v2.test.native_decl_selection drove this same producer over one synthetic ingest,
was refused on the same ground, and was repaired by the warm row
collision_resolved rather than by moving or excusing it.

THE SHARE POINT IS THE DECIDED OUTCOMES, NOT THE CONTEXT. NativeTestContext
carries a resolution context, which is the origin-bound shape that roster refuses
at publication; CensusProbeOutcomes carries three NativeCensusModuleOutcome
variants and an Int, and those carry only diagnostics -- symbols, loci and nodes
-- so the stored value is positively portable for the same reason
collision_resolved's ResolvedTree is.

VERIFIED BEFORE LANDING, because a warm row that fails to store STOPS THE LINE
for every lane and the roster's own header says claim_batch cannot verify one:
the warm path is forced at strict preparation, which a plain --entry run does not
execute. Confirmed by running the lane itself (claim_executor --required-ci
--required-lane witnesses) in-session, which is the instrument that roster names:

  phase=pure-producer-share-warm producer=...census_probe_outcomes
    disposition=Stored cpu_ms=598 provenance=built-by-preparation
  [floor-shared-fill] cache=cross_claim_pure_share key=census_probe_outcomes
    fill_ms=600 paid_by=<outside-fold> consumer_claims=4 disposition=outside-fold

and zero BLOCKING lines across the run, against eight on the refusing CI run.

WHAT WAS CONSIDERED AND REJECTED: moving the claims to a long home. Long modules
are DECLINED from the floor, so the discriminating control that stops this
instrument going green on a cause it never observed would stop being observed by
any gate -- silence, in DESIGN's own word for it. The claims stay on the floor and
still execute the real route; sharing moves WHERE the front end is paid, never
WHETHER it runs, so deleting the integration still reds every claim here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@gunbai-bot

gunbai-bot Bot commented Sep 19, 2026

Copy link
Copy Markdown
Contributor Author

The floor caught a real defect in my claims, and the fix is verified before landing

required-witnesses-floor refused this PR: the four context-consuming claims ran 197,116–200,054 evaluator steps against a 72,300-step new-witness budget (run 35464948266, job 105955463881, cause=completed_over_cost_requirement). The gate was right and the defect was mine — each claim called a helper that re-entered native_test_context_from_ingest, so every one of them rebuilt the language model, the grammar preparation and the parse of all three specimens in order to inspect one module's outcome.

That is DESIGN §3's own tell: "a claim whose evaluation is dominated by ONE call whose RESULT SHAPE is all it inspects." Worth noting against the approving review above, which saw this and dismissed it under §2's "one undeclared pure demand may recompute" — correct about correctness, but the armed cost gate measures something the reviewer wasn't measuring, and it was right.

The repair is the modeled one: stop recomputing, don't raise the line. census_probe_outcomes is now nullary and pure, enrolled as a WARM row in v2.workflow.floor_pure_producer_share, so the floor forces it once at strict preparation and bills it there. The precedent is exact — v2.test.native_decl_selection drove this same producer over one synthetic ingest, was refused on the same ground, and was repaired the same way.

Verified before landing, because this one can hurt other people. That roster says a warm row which fails to store stops the required-floor line for every lane, and its own header says claim_batch cannot verify one (the warm path is forced at strict preparation, which a plain --entry run skips — I confirmed that empirically: after the restructure all four claims still billed ~276k steps locally). Pushing an unverifiable change to shared infrastructure whose failure lands on others is externalization with me as the beneficiary, so I ran the lane itself in-session instead — the instrument the roster names:

phase=pure-producer-share-warm producer=…whole_tree_census_resolve.census_probe_outcomes
  disposition=Stored cpu_ms=598 provenance=built-by-preparation

[floor-shared-fill] cache=cross_claim_pure_share key=census_probe_outcomes
  fill_ms=600 paid_by=<outside-fold> consumer_claims=4 disposition=outside-fold

Stored (so it reifies, no line stop), billed outside the fold, serving all four claims — and zero BLOCKING lines across the run, against eight on the refusing CI run. The local run does end red, on an unrelated environmental cause (a local-repo wet lane involving self_host_logic_seed_unavailable_witness) that CI's floor did not report.

What I considered and rejected: moving the claims to a long home. That is the gate's other suggestion, and it is wrong here — long modules are declined from the floor, so the discriminating control that stops this instrument going green on a cause it never observed would stop being observed by any gate. DESIGN's word for that is silence. The claims stay on the floor and still execute the real route: sharing moves where the front end is paid, never whether it runs, so deleting the integration still reds every claim.

@gunbai-bot

gunbai-bot Bot commented Sep 19, 2026

Copy link
Copy Markdown
Contributor Author

One departure from the brief I had not stated, recorded now

Re-reading the request against the diff rather than against the approvals, I found an instruction I followed differently without saying so. §3b is explicit that the middle value — diverges with a stated reason — is fine, and that an unstated divergence is the only red, so this closes it.

The brief said: report the residual "reusing the [native-prepare-split] line it already prints rather than minting a second reporting surface."

What I did: minted a new [native-census-resolve] line.

Why. The two are not the same subject, and the reuse became impossible once the grain question was answered. [native-prepare-split] carries resolve_nanos, infer_nanos, prepare_nanos, decls and an outcome label — a per-module timing row for a loop that resolves and infers. A census never infers, so emitting that line would mean writing infer_nanos=0 and prepare_nanos=resolve_nanos for every module: a fabricated zero on a receipt, which the driver template's own annotations treat as fatal and §5 forbids outright. It also cannot carry what the row now has to carry — wise-pike-696 settled the grain as occurrence, one row per diagnostic with the resolver's own Diagnostic and locus, because XL-5 consumes the residual as a rewrite population and a folded reason cannot drive a rewrite. A per-module timing line has nowhere to put a diagnostic list.

So this is one subject getting one surface, not two surfaces for one subject: [native-prepare-split] stays the adjudicating loop's timing row and is untouched, [native-census-resolve] is the census's per-module outcome row, and the residual rows themselves go to stdout as census_residual objects. The brief's concern — don't mint a second way to say the same thing — is honoured; what changed is that the thing being said is different.

The same reasoning is why this is a third driver mode rather than a flag on census, which wise-pike-696 agreed with when I raised it.

— sent from vivid-lark-741

@gunbai-bot
gunbai-bot Bot added this pull request to the merge queue Sep 20, 2026
Merged via the queue into main with commit c6b4e7e Sep 20, 2026
4 checks passed
@gunbai-bot
gunbai-bot Bot deleted the session/vivid-lark-741 branch September 20, 2026 00:08
@briansrls
briansrls restored the session/vivid-lark-741 branch September 20, 2026 00:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants