Repository navigation
Prepare the frontend world once for the whole census - #8193
Merged
Merged
Conversation
…ember The census rebuilt the world fifteen times. classify_source called dag_lex_rules() and dag_grammar() itself, so every member revalidated the grammar and recomputed the FIRST/nullable analysis for a grammar that does not vary between members. The seam already existed and this reuses it rather than inventing one: 02_parse's prepared_grammar_carrier_note records that prepare_grammar exists so a K-module walk validates once instead of K times, with parse_module_prepared as the per-module half. take_frontend_census prepares once, classifies every rostered member through that one world, and reconciles the population before anything may read it as a census. THREE STATES, NOT ONE LIST. FrontendCensusReceipt is CensusTaken | CensusRefusedAtWorld | CensusRefusedAtReconciliation. A refused world is a fact about the world, not about any member — reporting it as fifteen ParseRefused stages would say fifteen members failed to parse when none was ever examined. CensusTaken is reachable only through the exact-reconciliation arm, so an unreconciled population cannot present as a census. classify_source keeps preparing LAZILY and now returns an Outcome. The laziness is a budget fact: a lex refusal is decided before any grammar is consulted, which is the only reason the LEX claim executes in the fast lane at all. The Outcome is because preparation can refuse and no FrontendStage honestly represents that. Roster identity is a structural hash fold over the paths in order, parameterized by the rows so discrimination is testable — an identity function that can only see the live roster is unfalsifiable. It does NOT cover labels: the only Symbol-to-text surface is symbol_lexeme, whose .dag body is a self-recursive host-intercepted stub, and grounding a provenance claim on that would be worse than the gap. Recorded as a known limit with the one case it cannot separate. NOT DELIVERED: the census still does not execute as a required gate. This removes a fifteen-times multiplier from a cost that may still exceed the budget alone, and I have not measured the result — there is no claim_batch binary in this container, so any figure would be invented. The enrollment decision is deferred to a measurement, not to an opinion. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DBBegUkJygQyr1zMHiv2eK
build and ci both refused with `expected expression, found EqEq`. The three roster-identity claims wrapped their comparison so that `== ContentHashEqual` began a line, and the .dag grammar does not continue an expression across that break. A line-leading `&&` does continue — the existing classify_normalize_diagnostics relies on it — so the shape is specific to the equality operator rather than to operators generally. Also merges origin/main, which had moved four commits ahead (#8180, #8164, #8185, #8189). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DBBegUkJygQyr1zMHiv2eK
… 51409 REQUEST_CHANGES was correct and the finding is exact: classify_source now returns Outcome<FrontendStage>, and the six p_* wall-attribution probes still fed that straight into frontend_stage_label, which takes a bare FrontendStage. The witness was updated; the probes were not. Same six probes, same omission, second increment running — I changed a callee's type and did not check its callers. render_classified matches both arms. A refused world renders WORLD_PREPARATION_REFUSED rather than any stage word, because no stage was reached and borrowing one would be the conflation the receipt's three states exist to prevent, repeated at the presentation boundary. I then swept every symbol this PR added, renamed or retyped for surviving references. Two stale prose mentions of observe_roster now name take_frontend_census. The remaining mentions of count_stage and classify_normalize_reasons are deliberate — they are the notes recording why those were removed, and they should outlive the code. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DBBegUkJygQyr1zMHiv2eK
briansrls
pushed a commit
that referenced
this pull request
Aug 13, 2026
…n alone (#8197) Three merged notes say the five-second figure is the cost of building dag_grammar(). That is not established at that grain, and I wrote it three times. What was actually timed is a claim calling classify_source on a source that REACHES the parser. That path covers lexing, grammar-value construction, grammar-to-Node conversion, well_formed, the five validation checks, the nullable fixed point, FIRST-set derivation, ambiguity residue construction, the prepared parse, and normalize. prepare_grammar's own carrier note enumerates most of those as its contents, so the aggregate was never a measurement of construction. No nested span isolated any stage, so no stage can be named as the owner. This matters beyond phrasing. If PREPARATION rather than construction dominates, #8193 already hoisted preparation out of the per-member path and the remaining cost is a different repair entirely — and "grammar preparation naturally costs five seconds" would have hardened into an assumption that stops anyone asking. Under the 500ms migration policy a five-second fixed cost is a defect to locate and remove, not a lane-placement decision, and locating it requires exclusive spans this figure does not have. The notes now state the aggregate and say explicitly that the per-stage attribution is unknown. Claude-Session: https://claude.ai/code/session_01DBBegUkJygQyr1zMHiv2eK Co-authored-by: gunbc-ci-auto-heal <gunbc-ci-auto-heal@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The affordability half of the executing 15-member census. Deliberately not the census itself —
see What does NOT land.
What lands
One prepared world. The census rebuilt the world fifteen times:
classify_sourcecalleddag_lex_rules()anddag_grammar()itself, so every member revalidated the grammar andrecomputed the FIRST/nullable analysis for a grammar that does not vary between members.
take_frontend_censusprepares once and classifies every rostered member through that one world.The seam already existed and this reuses it rather than inventing one:
02_parse'sprepared_grammar_carrier_noterecords thatprepare_grammarexists so a K-module walk validatesonce instead of K times, with
parse_module_preparedas the per-module half.parse_moduleisthe composition of the two, which is right for a single call and wrong for a census.
Three receipt states, not a list.
FrontendCensusReceipt = CensusTaken | CensusRefusedAtWorld | CensusRefusedAtReconciliation. A refused world is a fact about the world rather than about anymember — reporting it as fifteen
ParseRefusedstages would say fifteen members failed to parsewhen none was ever examined.
CensusTakenis reachable only through the exact-reconciliation arm,so an unreconciled population cannot present as a census.
classify_sourcestays lazy and returns anOutcome. The laziness is a budget fact: a lexrefusal is decided before any grammar is consulted, which is the only reason the LEX claim in the
durable witness executes inside the fast lane. The
Outcomeis because preparation can refuse andno
FrontendStagehonestly represents that. Both paths share one pipeline authority(
classify_after_lex), so there is no second copy of the stage chain.Roster identity is a structural hash fold over the paths in order, parameterized by the rows
so discrimination is testable — an identity that can only see the live roster is unfalsifiable.
Three controls: stable across calls, separates different paths, separates reordered rows.
What does NOT land
The census still does not execute as a required gate, and no claim here runs
take_frontend_census. Doing so reads fifteen live files and builds the grammar once, and onegrammar construction alone was previously measured over the five-second budget.
This increment removes a fifteen-times multiplier from a cost that may still exceed the budget by
itself. I have not measured the result — this container has no
claim_batchbinary, so anyfigure stated here would be invented rather than observed. The enrollment decision is therefore
deferred to a measurement, not to an opinion. If the prepared census is still over budget the
answer is realization or materialization, not a
long/directory and not silence.Known limit, stated rather than solved
Roster identity does not cover the labels. The only
Symbol-to-text surface in the tree isv2.std.compilers.lexingsymbol_lexeme, whose.dagbody is a self-recursive stub reachableonly through host interception — the shape DESIGN records as deleted elsewhere for that reason.
Grounding a provenance claim on it would be worse than the gap. Labels are separately constrained
by
verify_roster, which refuses duplicates; the one case this identity does not separate is tworosters differing only by a permutation of labels over identical paths in identical order.