Skip to content

LiteralTerminal in std.grammar; key choice dispatch on (class, word) - #11710

Merged
briansrls merged 5 commits into
mainfrom
session/calm-eagle-42-keyword-literal
Sep 20, 2026
Merged

briansrls merged 5 commits into
mainfrom
session/calm-eagle-42-keyword-literal

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Adds v2.std.grammar LiteralTerminal { token_class, lexeme }, a terminal that matches one exact word of a token class, and keys the parser's choice dispatch on (class, word) so alternatives led by different words are disjoint. This is infrastructure for gunbc#11622 (the G0 service family), which is rewritten on top of it, and for any future marker word.

Why this change and not a narrower one. #11622 first admitted the service marker words (input, output, exit, readonly, …) as 12 keyword token classes. Measured against #11700's floor run 35440523086, that raised eval steps on 192 of 4,114 common witnesses, all up and none down, +409k steps, the worst +38.6%, because every identifier paid a 13-way name-slot choice plus 12 extra lex rules. The alternative, matching the words as plain idents and refusing wrong ones at lowering (the test fn model), erases five located parse refusals, because services have no lowering arm. What both approximate is missing expressiveness: Terminal matches a token class and has no way to say "this word". And the class prefixes can't recover it. Service alternatives are all ident-led, and input { and output { share every class prefix ([ident, lbrace, ident, colon, …]), so no lookahead depth distinguishes them unless the word is part of the key.

Shape.

  • LiteralTerminal is a new GrammarExpr variant, not a field on Terminal, so the roughly 100 existing terminal rows are untouched. Every exhaustive match and every GrammarExprFold must handle the new arm. literal_terminal was added to the fold, and no wildcard arm was added anywhere.
  • Matching is not stamping. A literal terminal checks the word, then captures exactly what a StampClass terminal of its class captures: the class identity, never its own text (which is StampLexeme, a different fact).
  • The FIRST element becomes FirstKey = FirstAnyOfClass { class } | FirstExactWord { class, lexeme }. Equality (first_key_eq), overlap (first_key_overlaps) and token admission (first_key_admits_word) are each defined once, and every consumer goes through them.
  • ChoiceByToken still dispatches by class. Inside a class bucket that contains exact words, ChoiceByWord dispatches on the token's text. Each word's arm holds every alternative that can start on that word, both the exact ones and any-of-class ones, in authored order, so ordered-choice semantics are preserved. A bucket with no exact word compiles exactly the plan it compiled before.
  • Ambiguity rows keep their class-list fields. Overlap is computed through first_key_overlaps and projected to classes.
  • Emit (the backward direction) is improved, not just unaffected. No emit-side code consumes GrammarExpr terminals (06_translate renders from target forms via lex_rules_literal by token class, and target_model's mentions of GrammarExpr are in comments). And LiteralTerminal carries its word in the row itself, so a backward read has the text directly, without the lex_rules_literal_for_class lookup a keyword class forces. The row is a truer inverse than a keyword-class row.

Scope. The first estimate was a few hundred lines across 14 modules that name the FIRST and plan types. The two-arm FirstKey let existing class-only rows wrap instead of being rewritten, so the change is four files: 02_parse +249/-32, std.grammar +20, namespace_graft +1, choice_plan_test +77.

The discriminating evidence (v2.test.parse.choice_plan):

  • literal_words_sharing_a_class_prefix_validate_without_overlap_residue: the pair alpha . pcp_a / beta . pcp_a, which is identical once the leading word is erased, validates with no overlap residue.
  • the_same_pair_keyed_by_class_alone_leaves_overlap_residue: the control, pcp_word . pcp_a / pcp_word . pcp_a. It must be accepted with the named ^parse_grammar_choice_overlap_residue diagnostic, so an unrelated rejection fails it; it is not the negation of a clean acceptance. Together the two show that different exact words eliminate a FIRST-overlap residue that class keys retain. That is an authored-overlap fact, distinct from whether deeper class lookahead could separate some other pair.
  • the_class_bucket_of_a_word_pair_dispatches_on_the_word: the plan is ChoiceByWord.
  • the_second_word_parses_through_its_own_alternative: on a separate unequal-suffix pair (alpha . pcp_a / beta . pcp_b), it consumes all tokens and agrees with the ordered oracle. This is the execution control that catches a dispatch that always takes the first bucket.
  • a_literal_terminal_rejects_a_different_word_of_its_class
  • a_literal_terminal_publishes_the_class_identity_not_its_word: its capture equals a StampClass capture and differs from a StampLexeme capture.

Prediction, stated before the confirming floor run (DESIGN §6b)

This PR adds no keyword and no literal terminal to any real grammar, and a bucket without exact words compiles the plan it compiled before. If the reading is right, the floor run's required-floor-claim-cost artifact, joined against #11700's run 35440523086:

  1. No population moves up beyond runner noise in eval steps. Specifically, v2.test.tokenize.lex_rule_dispatch.lex_rule_dispatch_agrees_with_trying_every_rule stays at about 67,026 steps, and the 192 witnesses that Admit G0 service family so the native door clears leftover service #11622's keyword approach moved stay at baseline.
  2. The only new cost is the six new choice_plan rows.
  3. FirstKey wrapping costs nothing measurable in eval steps on grammars without literals. If it does, it will show here as a uniform small rise across parse-heavy witnesses, which would falsify this reading.

The service-family savings are not claimed here. They are #11622's to establish when it is rewritten on this.

Result (floor run 35446547321 on 359d182, FloorClean; the prediction held. The later heads change only test fixtures, not the parser)

Measured against #11700's run 35440523086 (baseline) and #11622's keyword run 35440687934, using required-floor-claim-cost eval_steps:

  • The full join shows 224 of 4,108 witnesses moved (181 up, 43 down, net +46,791). That population is confounded: the baseline run sits on an older main (aa3e61f against this branch's cb5950a), and the top movers are roster-counting witnesses (witness_admission, witness_exclusion_reconciliation, local_repo_wet_terminal, about +12%), whose cost tracks roster size as rows land on main, not parse cost.
  • The controlled comparison is the 111 witnesses the keyword approach moved by more than 5% and that are present in all three runs. None now differs from baseline by more than 1% either way, and the mean delta is +0.0009%. Examples, as base / keywords / literal:
    • body_lowering_fixture_resolves: 67,077 / 92,974 / 67,118
    • native_decl_selection x2: 35,339 / 47,357 / 35,339
    • arrow_body_form_eval_value_vertical_holds: 91,179 / 117,076 / 91,220
  • lex_rule_dispatch_agrees_with_trying_every_rule: 67,026 exactly (keywords: 82,691).
  • The six new choice_plan rows cost 295 to 7,095 steps each.

An exact return to baseline is what the mechanism predicts (no literal terminal in any real grammar, so the same plan is compiled). A partial improvement would have been consistent with other explanations.

Head 2e23ec3250 is a merge of main into cdc4407942. The heal workflow on main runs dag/gunbc/heal_candidate.dag, which this branch predated. cdc4407942 is a strict ancestor of the head, so this PR's own contribution is carried whole and unchanged. The merge brings main's commits and nothing else. Those touch many files, including namespace_graft.dag, which main changed when #11694 landed its producer marker; that movement is main's, not this PR's. Floor run 35461610687 on this head is FloorClean, and it is the first run of the literal terminal alongside main's SH-2 grammar.

Status at 2e23ec3250. Source is signed off at this exact head: both held findings on the evidence are closed, and the substance was re-verified on the merge head. All four checks pass and mergeStateStatus is CLEAN. The only outstanding requirement is a completed review bound to this head. That is an infrastructure gap, not a question about the change: two review processes on this PR have crashed without a verdict (68443 on 1cc101e29d, 68584 on this head), and an approval on an ancestor head does not count as an exact-head review.

🤖 Generated with Claude Code

… word).

A literal terminal matches one exact word of a token class and publishes
the class identity, so a marker word can be a grammar row without minting
a keyword token class. The parser's FIRST element becomes FirstKey
(any-of-class | exact word) with equality, overlap and admission defined
once; ChoiceByWord dispatches on the token's word inside a class bucket, and
a bucket with no exact word compiles exactly the plan it did before.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
gunbc-ci-auto-heal and others added 2 commits September 19, 2026 13:42
…tors.

The floor's resolve refused six annotations inside declaration bodies
(DESIGN 4c: only leading module-scope blocks are modeled) and an if whose
arms disagreed because a fold started from an untyped Empty.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Review 68453: the constructor had no call site; the claims built the
variant directly. They now use the constructor a grammar author would.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@gunbai-bot

gunbai-bot Bot commented Sep 19, 2026

Copy link
Copy Markdown
Contributor Author

Review 68453: fixed in 8e2b941. grammar_expr_literal_terminal now has a call site: the choice_plan claims' lit helper builds the variant through it, as a grammar author will, and the test no longer names LiteralTerminal directly.

— sent from calm-eagle-42

@briansrls
briansrls added this pull request to the merge queue Sep 19, 2026
… residue positively.

Side-chat hold on 8e2b941: the compared pair also differed in its suffix
class, so class lookahead could separate it; and the control negated a
helper that is false on ANY rejection. The compared pair is now
word . pcp_a / word . pcp_a, and the control requires acceptance WITH the
^parse_grammar_choice_overlap_residue diagnostic, so an unrelated rejection
fails it. The unequal-suffix pair stays as the second-word execution control.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@gunbai-bot
gunbai-bot Bot removed this pull request from the merge queue due to a manual request Sep 19, 2026
@briansrls
briansrls added this pull request to the merge queue Sep 20, 2026
Merged via the queue into main with commit 877b733 Sep 20, 2026
4 checks passed
@briansrls
briansrls deleted the session/calm-eagle-42-keyword-literal branch September 20, 2026 16:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant