Skip to content

02_parse: read the entry production off the prepared grammar (~65k fixed eval steps per parse) - #13244

Merged
gunbai-bot[bot] merged 2 commits into
mainfrom
session/wise-ant-549
Oct 4, 2026
Merged

gunbai-bot[bot] merged 2 commits into
mainfrom
session/wise-ant-549

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

The defect (DESIGN §6b reading)

v2.compiler.parse parse_production_prepared_measured_seeded received a prepared grammar, but handled its entry production differently from every nested one:

  • Production lookup. It used grammar_lookup_production, an O(P) fold_list over all productions that does not short-circuit (P≈204 on the dag grammar). The analysis already carries productions_by_name.
  • Entry expression. It passed the raw GrammarExpr to parse_expr, which calls prepare_grammar_expr on it. That recompiles every choice plan and recomputes authored_choice_site_residue (the overlap-roster lookahead analysis), and then keeps only .tree. The analysis already carries that tree as prepared_exprs[production], built once in prepare_grammar. Nested nonterminals already read it from there (parse_nonterminal_memoized → parse_table_prepared_expr).

So a property of the grammar was re-derived on every parse. The earliest unjustified link is the entry dispatch. The fact's owner, the prepared analysis, was already correct, so no carrier changes. The fix deletes the re-derivation and reads the owned fact. Nothing is cached beside it.

Prediction stated before measuring: the parse step collapses to roughly the cost of the token walk; parse trees are unchanged (prepare_grammar_expr's production/productions arguments reach only the discarded residue, never the tree); the per-token slope is unchanged.

Measurements (claim_batch, eval steps, same binary, before vs after)

Probe claims (not committed): dag_prepared_grammar → tokenize → parse_production_prepared at dag_production_arg.

fragment tokens before after
tokenize only — 17,088 17,088
x: 1 3 90,000 25,385
x: a + b + c 7 102,050 37,435
x: a + b + c + d + e + f + g 15 123,032 58,417

The parse step drops from about 72.9k to about 8.3k (target was under 20k). The fixed term falls by about 64.6k. The slope is unchanged at about 2.6–3k steps per token (fit over the three rows).

The whole src/v2/test/claim/parse directory (361 claims, explicit --entry/--functions groups) gives 354 pass / 7 fail both before and after, the same seven claims. They already fail on main; none comes from this change. Largest per-claim deltas: else-less-if routes −67.6k each, a_plain_typed_param_takes_the_bare_route_holds −40.0k, field_patterns_take_pun_and_suffix_routes_holds −38.6k. Claims entering at dag_production_module save only about 1.4k (the O(P) fold), because the module production's own preparation is small.

Equivalence controls (probe, run against the fix)

  • For every production of the dag grammar and of the python grammar, prepared_exprs[p] equals the per-parse re-preparation the old route built (same nullable_set/first_rows/token_dispatch): PASS on both grammars.
  • Trees from the old route (verbatim pre-change body) and the new route are equal (== and content_hash, span index, diagnostics). Dag fragments covered: x: 1, x: a + b + c, x: (refusal), requires none|Filesystem, a fn decl, a full module, and an undefined production (same refusal). Python: the extended def choose fragment is accepted by both routes; the all-production equality above covers the tree.
  • The roster is unaffected. Residue is still computed once in compute_grammar_first_analysis_given, and cont_the_prepared_grammar_carries_residue_holds has the same verdict before and after.

Why the per-token slope stands (measured, accepted by the program manager)

The slope is honest work for this grammar's tree contract, not a re-derivation. Probe: the token a parsed entering at each level of the chain, eval steps:

level added steps
primary (incl. entry overhead) 1,348
+ postfix 860
+ unary 493
+ binary (7 inlined precedence helpers: pipe/or/and/eq/cmp/add/mul) 3,080 (~440 each)
+ expr (guarded choice) 1,665
+ arg (named-head guard) 1,263
+ b continuation 5,859 for 2 tokens

Memo counters on 3/7/15 tokens: 5/12/24 lookups, 0 hits, every lookup a miss (~1.75 memoized levels per token). Cost per operation: table rebuild 7, indexed lookup 6.5, memo miss 24, insert 58, empty repeat stamp 14. Memo bookkeeping is therefore about 110 steps of a 400–900-step level. The bulk is that every operand walks all seven precedence levels and attempts an empty continuation at each, stamping it as a parse-tree node (repeat/optional nodes carrying minted occurrence ids). Reducing that means changing the tree contract (e.g. precedence climbing), which is a design question, not a defect. Consequence: claims that enter at dag_production_module stay slope-bound and need narrower supplied entries (DESIGN §3), not a parse fix. Follow-up, separate PR: repeat progress is decided by length(rest) != length(remaining), which is O(n) per iteration in native time.

🤖 Generated with Claude Code

…prepare it per parse

parse_production_prepared routed the entry production through parse_expr, which
re-ran prepare_grammar_expr (choice-plan compilation plus the authored-choice
overlap residue, then discarded) and looked the production up with an O(P)
non-short-circuiting fold, on every parse. Both facts are already on the
prepared analysis (prepared_exprs, productions_by_name), which every nested
nonterminal already reads. 'x: 1' at dag_production_arg: 89,998 -> 25,383
claim eval steps (parse step ~72.9k -> ~8.3k).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
gunbai-bot Bot pushed a commit that referenced this pull request Oct 4, 2026
… production); stacks #13244

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…otation (DESIGN §6)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@gunbai-bot
gunbai-bot Bot added this pull request to the merge queue Oct 4, 2026
Merged via the queue into main with commit 47f457f Oct 4, 2026
4 checks passed
@gunbai-bot
gunbai-bot Bot deleted the session/wise-ant-549 branch October 4, 2026 13:50
@briansrls
briansrls restored the session/wise-ant-549 branch October 4, 2026 18:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants