Skip to content

01_tokenize: lex matchers as pattern data, dag_prepared_lex tier-served (stacked on #13287) - #13294

Merged
gunbai-bot[bot] merged 7 commits into
mainfrom
session/wise-ant-549-lex-data
Oct 5, 2026
Merged

gunbai-bot[bot] merged 7 commits into
mainfrom
session/wise-ant-549-lex-data

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Stacked on #13287, and the closure-free half of the program manager's ruling. The compiled lex rules become data, so the prepared rule set is portable and dag_prepared_lex can be served across claims like dag_prepared_grammar.

The defect (DESIGN §6b reading)

v2.compiler.tokenize lex_rule_thunk folded each LexPattern into a tree of LexMatchThunk { apply: fn(LexCursor) -> … } closures, and CompiledLexRule stored that closure. The cross-claim tier refuses to publish a value that contains closures. The floor's cross-claim demand census showed the result: lex_compile_rules was evaluated once per claim (16 of 16 on #13279's run) and never served, so every claim that tokenized .dag text first recompiled the language's whole rule set.

Fix, at the owning link. The matcher is the pattern, read:

  • Matcher: lex_match_pattern(pattern, source) is one structural walk over the LexPattern data. Each arm answers exactly what the corresponding closure answered: the consumed source as the lexeme, a keyword refused when an identifier character follows, a choice trying its left alternative first, a repeat stopping at a rejection or an empty match, and a delimited body running until its close matches.
  • Loops: lex_repeat_step and lex_delimited_step take patterns instead of thunks.
  • Compiled rule: CompiledLexRule keeps the rule, its index and its FIRST set; LexMatchThunk and lex_rule_thunk are deleted.
  • Shared producer: v2.compiler.program_assembly gains dag_prepared_lex() beside dag_prepared_grammar(), enrolled warm in v2.workflow.floor_pure_producer_share with its ledger reason.
  • Consumer: the declaration-choice route claims (route_parse_at) read it through tokenize_prepared.

Controls

  • Token streams byte-identical. A probe returned the full token streams for five fixtures under each lexer: two dag modules covering keywords, comments, strings with escapes, hex, float, operators, symbols and blocks; a dag refusal (@); and two python functions. The printed values match: 14,494 bytes each, the same sha256.
  • Population. src/v2/test/claim/tokenize + src/v2/test/claim/parse: 389 pass / 7 fail. The seven are the same claims that fail on main; none is introduced here.
  • Cost of the prepare itself (claim_batch): 7.3k eval steps for data matchers, against 9.0k for the closure version.

Per-claim measurement on the floor

Base run 37216639425 (this PR before the helper switch) vs run 37219911955 (this head).

  • Serving. Base: dag_prepared_lex landed at warm preparation with disposition Stored (fill 54 ms, so the value is portable), but no claim read it yet. The cross-claim demand census still listed prepare_lex_rules with claims=15 evals=15: all 15 claims in test.claim.parse_test_fn_decl_return_clause recompiled the rules through tokenize(rules: dag_lex()). This head: their helper nfbcp_parsed reads tokenize_prepared(dag_prepared_lex()). The fill row shows consumer_claims=15, and prepare_lex_rules no longer appears in the per-claim demand census.
  • Steps. Each of the 15 claims fell by the same ~7.1k eval steps, the per-claim compile. For example, nfbcp_eq_body_specimen_parses_holds went from 44,364 to 37,228, nfbcp_control_ordinary_fn_return_type_holds from 51,802 to 44,663, and nfbcp_repeated_declaration_refuses_at_the_second_declaration_holds from 66,882 to 59,737.
  • CPU. Down 25–85 ms per claim. The remaining ~400 ms is the per-context argument hash of the prepared grammar, which is royal-deer-478's tier follow-up.

Other tokenize(rules: dag_lex()) callers in claims can switch to the served value the same way; this PR switches the two helpers the floor plans.

🤖 Generated with Claude Code

gunbc-ci-auto-heal and others added 2 commits October 4, 2026 15:27
…kenize

prepare_lex_rules compiles the rule set (matchers, first-character dispatch, one dispatch per
mode) and checks its mode declarations once into PreparedLexRules; tokenize_prepared reads it.
tokenize(rules) stays the one-text door. The per-file folds (program assembly, closure ingest,
reference conservation, the native front end) prepare once beside prepare_grammar and hand the
prepared rules to program_assembly_phase_tokenize, instead of recompiling per file and relying on
the interpreter's in-context call memo to dedupe it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…_lex served like dag_prepared_grammar

lex_rule_thunk built one LexMatchThunk closure per pattern node, so a compiled rule set held
closures, which the cross-claim tier refuses to publish: every claim recompiled the .dag rules.
lex_match_pattern now reads the LexPattern data directly, arm for arm what each closure did, so
CompiledLexRule carries the rule, its index and its FIRST set only and the prepared rule set is
portable. dag_prepared_lex (beside dag_prepared_grammar) prepares it once and is enrolled warm in
floor_pure_producer_share; the declaration-choice route claims read it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@gunbai-bot
gunbai-bot Bot changed the base branch from session/wise-ant-549-lex-prepared to main October 4, 2026 16:18
gunbc-ci-auto-heal and others added 2 commits October 4, 2026 16:24
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ed_lex

The module's 15 claims each recompiled the .dag lex rules through tokenize(rules: dag_lex());
they now read the prepared rules served across claims.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@gunbai-bot
gunbai-bot Bot added this pull request to the merge queue Oct 4, 2026
@gunbai-bot
gunbai-bot Bot removed this pull request from the merge queue due to a manual request Oct 4, 2026
claude added 2 commits October 4, 2026 22:14
Both conflicts were #13287's own changes, already on main: the generated v1_compiler_emit_rust.rs
mirror and 05_emit_rust.dag now equal main's exactly, so #13294 no longer touches v1 and main's
regen fixed point covers it; tokenize_prepared takes main's unicode_scalar_unfold ingress (#13151).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… hand warm row

#13043 deleted the hand-maintained floor_cross_claim_pure_producers_warm roster: cross-claim sharing
is now derived from planned claims' call-site demand. dag_prepared_lex is a nullary producer of a
portable value demanded by many claims, so the derivation shares it; the row and its comment go.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…iew 76232)

Structural questions about a pattern go through fold_lex_pattern; matching a pattern against input
threads a cursor through its children, which a catamorphism can only produce as closures, and #13294
exists to keep compiled matchers as data. The departure and its reason are now stated at the rule in
v2.std.compilers.lexing and at the walker; a new LexPattern variant refuses the walker's exhaustive
match at compile time, so it cannot be skipped silently.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@gunbai-bot

gunbai-bot Bot commented Oct 5, 2026

Copy link
Copy Markdown
Contributor Author

Re review 76232 (lex_match_pattern walks LexPattern by hand instead of through fold_lex_pattern): addressed with your option (b) in 2778f01.

Why not (a): fold_lex_pattern is a catamorphism (children are folded before the parent, with no input). Matching depends on input: a sequence's right part starts where its left part stopped, a choice tries its right part only when the left refuses, and a repeat re-runs its element at each new cursor. A catamorphism can produce that only as a cursor -> result function per sub-pattern, i.e. closures, which this PR removes so compiled matchers are portable data that dag_prepared_lex can serve across claims. A fold whose algebra receives the unevaluated children and a recursion callback would be this same walk under another name.

What changed: the rule at v2.std.compilers.lexing (beside lex_pattern_is_keyword) now scopes the fold to structural questions and names v2.compiler.tokenize lex_match_pattern as the one sanctioned direct walk, with that reason. The walker carries the matching annotation. On the silent-miss risk: the walk is a match over the closed LexPattern coproduct, so a new variant refuses it at compile time as non-exhaustive; it can't be skipped.

— sent from stern-bear-500

@gunbai-bot
gunbai-bot Bot added this pull request to the merge queue Oct 5, 2026
Merged via the queue into main with commit fc2877e Oct 5, 2026
4 checks passed
@gunbai-bot
gunbai-bot Bot deleted the session/wise-ant-549-lex-data branch October 5, 2026 10:47
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant