Repository navigation
02_parse: read the entry production off the prepared grammar (~65k fixed eval steps per parse) - #13244
Merged
Conversation
…prepare it per parse parse_production_prepared routed the entry production through parse_expr, which re-ran prepare_grammar_expr (choice-plan compilation plus the authored-choice overlap residue, then discarded) and looked the production up with an O(P) non-short-circuiting fold, on every parse. Both facts are already on the prepared analysis (prepared_exprs, productions_by_name), which every nested nonterminal already reads. 'x: 1' at dag_production_arg: 89,998 -> 25,383 claim eval steps (parse step ~72.9k -> ~8.3k). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
gunbai-bot Bot
pushed a commit
that referenced
this pull request
Oct 4, 2026
… production); stacks #13244 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…otation (DESIGN §6) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This was referenced Oct 4, 2026
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The defect (DESIGN §6b reading)
v2.compiler.parseparse_production_prepared_measured_seededreceived a prepared grammar, but handled its entry production differently from every nested one:grammar_lookup_production, an O(P)fold_listover all productions that does not short-circuit (P≈204 on the dag grammar). The analysis already carriesproductions_by_name.GrammarExprtoparse_expr, which callsprepare_grammar_expron it. That recompiles every choice plan and recomputesauthored_choice_site_residue(the overlap-roster lookahead analysis), and then keeps only.tree. The analysis already carries that tree asprepared_exprs[production], built once inprepare_grammar. Nested nonterminals already read it from there (parse_nonterminal_memoized→parse_table_prepared_expr).So a property of the grammar was re-derived on every parse. The earliest unjustified link is the entry dispatch. The fact's owner, the prepared analysis, was already correct, so no carrier changes. The fix deletes the re-derivation and reads the owned fact. Nothing is cached beside it.
Prediction stated before measuring: the parse step collapses to roughly the cost of the token walk; parse trees are unchanged (
prepare_grammar_expr'sproduction/productionsarguments reach only the discarded residue, never the tree); the per-token slope is unchanged.Measurements (claim_batch, eval steps, same binary, before vs after)
Probe claims (not committed):
dag_prepared_grammar→tokenize→parse_production_preparedatdag_production_arg.x: 1x: a + b + cx: a + b + c + d + e + f + gThe parse step drops from about 72.9k to about 8.3k (target was under 20k). The fixed term falls by about 64.6k. The slope is unchanged at about 2.6–3k steps per token (fit over the three rows).
The whole
src/v2/test/claim/parsedirectory (361 claims, explicit--entry/--functionsgroups) gives 354 pass / 7 fail both before and after, the same seven claims. They already fail on main; none comes from this change. Largest per-claim deltas: else-less-if routes −67.6k each,a_plain_typed_param_takes_the_bare_route_holds−40.0k,field_patterns_take_pun_and_suffix_routes_holds−38.6k. Claims entering atdag_production_modulesave only about 1.4k (the O(P) fold), because the module production's own preparation is small.Equivalence controls (probe, run against the fix)
prepared_exprs[p]equals the per-parse re-preparation the old route built (samenullable_set/first_rows/token_dispatch): PASS on both grammars.==andcontent_hash, span index, diagnostics). Dag fragments covered:x: 1,x: a + b + c,x:(refusal),requires none|Filesystem, a fn decl, a full module, and an undefined production (same refusal). Python: the extendeddef choosefragment is accepted by both routes; the all-production equality above covers the tree.compute_grammar_first_analysis_given, andcont_the_prepared_grammar_carries_residue_holdshas the same verdict before and after.Why the per-token slope stands (measured, accepted by the program manager)
The slope is honest work for this grammar's tree contract, not a re-derivation. Probe: the token
aparsed entering at each level of the chain, eval steps:+ bcontinuationMemo counters on 3/7/15 tokens: 5/12/24 lookups, 0 hits, every lookup a miss (~1.75 memoized levels per token). Cost per operation: table rebuild 7, indexed lookup 6.5, memo miss 24, insert 58, empty repeat stamp 14. Memo bookkeeping is therefore about 110 steps of a 400–900-step level. The bulk is that every operand walks all seven precedence levels and attempts an empty continuation at each, stamping it as a parse-tree node (repeat/optional nodes carrying minted occurrence ids). Reducing that means changing the tree contract (e.g. precedence climbing), which is a design question, not a defect. Consequence: claims that enter at
dag_production_modulestay slope-bound and need narrower supplied entries (DESIGN §3), not a parse fix. Follow-up, separate PR: repeat progress is decided bylength(rest) != length(remaining), which is O(n) per iteration in native time.🤖 Generated with Claude Code