Skip to content

Prepare the frontend world once for the whole census - #8193

Merged
briansrls merged 4 commits into
mainfrom
session/crisp-lynx-832-census-exec
Aug 12, 2026
Merged

briansrls merged 4 commits into
mainfrom
session/crisp-lynx-832-census-exec

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Aug 12, 2026 •

Copy link
Copy Markdown
Contributor

The affordability half of the executing 15-member census. Deliberately not the census itself —
see What does NOT land.

What lands

One prepared world. The census rebuilt the world fifteen times: classify_source called
dag_lex_rules() and dag_grammar() itself, so every member revalidated the grammar and
recomputed the FIRST/nullable analysis for a grammar that does not vary between members.
take_frontend_census prepares once and classifies every rostered member through that one world.

The seam already existed and this reuses it rather than inventing one: 02_parse's
prepared_grammar_carrier_note records that prepare_grammar exists so a K-module walk validates
once instead of K times, with parse_module_prepared as the per-module half. parse_module is
the composition of the two, which is right for a single call and wrong for a census.

Three receipt states, not a list. FrontendCensusReceipt = CensusTaken | CensusRefusedAtWorld | CensusRefusedAtReconciliation. A refused world is a fact about the world rather than about any
member — reporting it as fifteen ParseRefused stages would say fifteen members failed to parse
when none was ever examined. CensusTaken is reachable only through the exact-reconciliation arm,
so an unreconciled population cannot present as a census.

classify_source stays lazy and returns an Outcome. The laziness is a budget fact: a lex
refusal is decided before any grammar is consulted, which is the only reason the LEX claim in the
durable witness executes inside the fast lane. The Outcome is because preparation can refuse and
no FrontendStage honestly represents that. Both paths share one pipeline authority
(classify_after_lex), so there is no second copy of the stage chain.

Roster identity is a structural hash fold over the paths in order, parameterized by the rows
so discrimination is testable — an identity that can only see the live roster is unfalsifiable.
Three controls: stable across calls, separates different paths, separates reordered rows.

What does NOT land

The census still does not execute as a required gate, and no claim here runs
take_frontend_census. Doing so reads fifteen live files and builds the grammar once, and one
grammar construction alone was previously measured over the five-second budget.

This increment removes a fifteen-times multiplier from a cost that may still exceed the budget by
itself. I have not measured the result — this container has no claim_batch binary, so any
figure stated here would be invented rather than observed. The enrollment decision is therefore
deferred to a measurement, not to an opinion. If the prepared census is still over budget the
answer is realization or materialization, not a long/ directory and not silence.

Known limit, stated rather than solved

Roster identity does not cover the labels. The only Symbol-to-text surface in the tree is
v2.std.compilers.lexing symbol_lexeme, whose .dag body is a self-recursive stub reachable
only through host interception — the shape DESIGN records as deleted elsewhere for that reason.
Grounding a provenance claim on it would be worse than the gap. Labels are separately constrained
by verify_roster, which refuses duplicates; the one case this identity does not separate is two
rosters differing only by a permutation of labels over identical paths in identical order.

…ember

The census rebuilt the world fifteen times. classify_source called dag_lex_rules() and
dag_grammar() itself, so every member revalidated the grammar and recomputed the FIRST/nullable
analysis for a grammar that does not vary between members.

The seam already existed and this reuses it rather than inventing one: 02_parse's
prepared_grammar_carrier_note records that prepare_grammar exists so a K-module walk validates
once instead of K times, with parse_module_prepared as the per-module half. take_frontend_census
prepares once, classifies every rostered member through that one world, and reconciles the
population before anything may read it as a census.

THREE STATES, NOT ONE LIST. FrontendCensusReceipt is CensusTaken | CensusRefusedAtWorld |
CensusRefusedAtReconciliation. A refused world is a fact about the world, not about any member —
reporting it as fifteen ParseRefused stages would say fifteen members failed to parse when none
was ever examined. CensusTaken is reachable only through the exact-reconciliation arm, so an
unreconciled population cannot present as a census.

classify_source keeps preparing LAZILY and now returns an Outcome. The laziness is a budget fact:
a lex refusal is decided before any grammar is consulted, which is the only reason the LEX claim
executes in the fast lane at all. The Outcome is because preparation can refuse and no
FrontendStage honestly represents that.

Roster identity is a structural hash fold over the paths in order, parameterized by the rows so
discrimination is testable — an identity function that can only see the live roster is
unfalsifiable. It does NOT cover labels: the only Symbol-to-text surface is symbol_lexeme, whose
.dag body is a self-recursive host-intercepted stub, and grounding a provenance claim on that
would be worse than the gap. Recorded as a known limit with the one case it cannot separate.

NOT DELIVERED: the census still does not execute as a required gate. This removes a fifteen-times
multiplier from a cost that may still exceed the budget alone, and I have not measured the result
— there is no claim_batch binary in this container, so any figure would be invented. The
enrollment decision is deferred to a measurement, not to an opinion.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DBBegUkJygQyr1zMHiv2eK
@gunbai-bot gunbai-bot Bot changed the title v2 self host Prepare the frontend world once for the whole census Aug 12, 2026
@gunbai-bot
gunbai-bot Bot marked this pull request as ready for review August 12, 2026 15:55
gunbc-ci-auto-heal and others added 3 commits August 12, 2026 16:00
build and ci both refused with `expected expression, found EqEq`. The three roster-identity claims
wrapped their comparison so that `== ContentHashEqual` began a line, and the .dag grammar does not
continue an expression across that break. A line-leading `&&` does continue — the existing
classify_normalize_diagnostics relies on it — so the shape is specific to the equality operator
rather than to operators generally.

Also merges origin/main, which had moved four commits ahead (#8180, #8164, #8185, #8189).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DBBegUkJygQyr1zMHiv2eK
… 51409

REQUEST_CHANGES was correct and the finding is exact: classify_source now returns
Outcome<FrontendStage>, and the six p_* wall-attribution probes still fed that straight into
frontend_stage_label, which takes a bare FrontendStage. The witness was updated; the probes were
not. Same six probes, same omission, second increment running — I changed a callee's type and
did not check its callers.

render_classified matches both arms. A refused world renders WORLD_PREPARATION_REFUSED rather
than any stage word, because no stage was reached and borrowing one would be the conflation the
receipt's three states exist to prevent, repeated at the presentation boundary.

I then swept every symbol this PR added, renamed or retyped for surviving references. Two stale
prose mentions of observe_roster now name take_frontend_census. The remaining mentions of
count_stage and classify_normalize_reasons are deliberate — they are the notes recording why
those were removed, and they should outlive the code.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DBBegUkJygQyr1zMHiv2eK
@briansrls
briansrls merged commit 8a3cad7 into main Aug 12, 2026
5 checks passed
@briansrls
briansrls deleted the session/crisp-lynx-832-census-exec branch August 12, 2026 17:50
briansrls pushed a commit that referenced this pull request Aug 13, 2026
…n alone (#8197)

Three merged notes say the five-second figure is the cost of building dag_grammar(). That is not
established at that grain, and I wrote it three times.

What was actually timed is a claim calling classify_source on a source that REACHES the parser.
That path covers lexing, grammar-value construction, grammar-to-Node conversion, well_formed, the
five validation checks, the nullable fixed point, FIRST-set derivation, ambiguity residue
construction, the prepared parse, and normalize. prepare_grammar's own carrier note enumerates
most of those as its contents, so the aggregate was never a measurement of construction.

No nested span isolated any stage, so no stage can be named as the owner.

This matters beyond phrasing. If PREPARATION rather than construction dominates, #8193 already
hoisted preparation out of the per-member path and the remaining cost is a different repair
entirely — and "grammar preparation naturally costs five seconds" would have hardened into an
assumption that stops anyone asking. Under the 500ms migration policy a five-second fixed cost is
a defect to locate and remove, not a lane-placement decision, and locating it requires exclusive
spans this figure does not have.

The notes now state the aggregate and say explicitly that the per-stage attribution is unknown.


Claude-Session: https://claude.ai/code/session_01DBBegUkJygQyr1zMHiv2eK

Co-authored-by: gunbc-ci-auto-heal <gunbc-ci-auto-heal@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant