Repository navigation
LiteralTerminal in std.grammar; key choice dispatch on (class, word) - #11710
Merged
Merged
Conversation
… word). A literal terminal matches one exact word of a token class and publishes the class identity, so a marker word can be a grammar row without minting a keyword token class. The parser's FIRST element becomes FirstKey (any-of-class | exact word) with equality, overlap and admission defined once; ChoiceByWord dispatches on the token's word inside a class bucket, and a bucket with no exact word compiles exactly the plan it did before. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2 of 3 tasks
…tors. The floor's resolve refused six annotations inside declaration bodies (DESIGN 4c: only leading module-scope blocks are modeled) and an if whose arms disagreed because a fold started from an untyped Empty. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Review 68453: the constructor had no call site; the claims built the variant directly. They now use the constructor a grammar author would. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Contributor
Author
|
Review 68453: fixed in 8e2b941. — sent from calm-eagle-42 |
briansrls
added this pull request to the merge queue
Sep 19, 2026
… residue positively. Side-chat hold on 8e2b941: the compared pair also differed in its suffix class, so class lookahead could separate it; and the control negated a helper that is false on ANY rejection. The compared pair is now word . pcp_a / word . pcp_a, and the control requires acceptance WITH the ^parse_grammar_choice_overlap_residue diagnostic, so an unrelated rejection fails it. The unequal-suffix pair stays as the second-word execution control. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds
v2.std.grammarLiteralTerminal { token_class, lexeme }, a terminal that matches one exact word of a token class, and keys the parser's choice dispatch on(class, word)so alternatives led by different words are disjoint. This is infrastructure for gunbc#11622 (the G0 service family), which is rewritten on top of it, and for any future marker word.Why this change and not a narrower one. #11622 first admitted the service marker words (
input,output,exit,readonly, …) as 12 keyword token classes. Measured against #11700's floor run 35440523086, that raised eval steps on 192 of 4,114 common witnesses, all up and none down, +409k steps, the worst +38.6%, because every identifier paid a 13-way name-slot choice plus 12 extra lex rules. The alternative, matching the words as plain idents and refusing wrong ones at lowering (thetest fnmodel), erases five located parse refusals, because services have no lowering arm. What both approximate is missing expressiveness:Terminalmatches a token class and has no way to say "this word". And the class prefixes can't recover it. Service alternatives are all ident-led, andinput {andoutput {share every class prefix ([ident, lbrace, ident, colon, …]), so no lookahead depth distinguishes them unless the word is part of the key.Shape.
LiteralTerminalis a newGrammarExprvariant, not a field onTerminal, so the roughly 100 existing terminal rows are untouched. Every exhaustive match and everyGrammarExprFoldmust handle the new arm.literal_terminalwas added to the fold, and no wildcard arm was added anywhere.StampClassterminal of its class captures: the class identity, never its own text (which isStampLexeme, a different fact).FirstKey = FirstAnyOfClass { class } | FirstExactWord { class, lexeme }. Equality (first_key_eq), overlap (first_key_overlaps) and token admission (first_key_admits_word) are each defined once, and every consumer goes through them.ChoiceByTokenstill dispatches by class. Inside a class bucket that contains exact words,ChoiceByWorddispatches on the token's text. Each word's arm holds every alternative that can start on that word, both the exact ones and any-of-class ones, in authored order, so ordered-choice semantics are preserved. A bucket with no exact word compiles exactly the plan it compiled before.first_key_overlapsand projected to classes.GrammarExprterminals (06_translaterenders from target forms vialex_rules_literalby token class, andtarget_model's mentions ofGrammarExprare in comments). AndLiteralTerminalcarries its word in the row itself, so a backward read has the text directly, without thelex_rules_literal_for_classlookup a keyword class forces. The row is a truer inverse than a keyword-class row.Scope. The first estimate was a few hundred lines across 14 modules that name the FIRST and plan types. The two-arm
FirstKeylet existing class-only rows wrap instead of being rewritten, so the change is four files:02_parse+249/-32,std.grammar+20,namespace_graft+1,choice_plan_test+77.The discriminating evidence (
v2.test.parse.choice_plan):literal_words_sharing_a_class_prefix_validate_without_overlap_residue: the pairalpha . pcp_a/beta . pcp_a, which is identical once the leading word is erased, validates with no overlap residue.the_same_pair_keyed_by_class_alone_leaves_overlap_residue: the control,pcp_word . pcp_a/pcp_word . pcp_a. It must be accepted with the named^parse_grammar_choice_overlap_residuediagnostic, so an unrelated rejection fails it; it is not the negation of a clean acceptance. Together the two show that different exact words eliminate a FIRST-overlap residue that class keys retain. That is an authored-overlap fact, distinct from whether deeper class lookahead could separate some other pair.the_class_bucket_of_a_word_pair_dispatches_on_the_word: the plan isChoiceByWord.the_second_word_parses_through_its_own_alternative: on a separate unequal-suffix pair (alpha . pcp_a/beta . pcp_b), it consumes all tokens and agrees with the ordered oracle. This is the execution control that catches a dispatch that always takes the first bucket.a_literal_terminal_rejects_a_different_word_of_its_classa_literal_terminal_publishes_the_class_identity_not_its_word: its capture equals aStampClasscapture and differs from aStampLexemecapture.Prediction, stated before the confirming floor run (DESIGN §6b)
This PR adds no keyword and no literal terminal to any real grammar, and a bucket without exact words compiles the plan it compiled before. If the reading is right, the floor run's
required-floor-claim-costartifact, joined against #11700's run 35440523086:v2.test.tokenize.lex_rule_dispatch.lex_rule_dispatch_agrees_with_trying_every_rulestays at about 67,026 steps, and the 192 witnesses that Admit G0 service family so the native door clears leftoverservice#11622's keyword approach moved stay at baseline.choice_planrows.FirstKeywrapping costs nothing measurable in eval steps on grammars without literals. If it does, it will show here as a uniform small rise across parse-heavy witnesses, which would falsify this reading.The service-family savings are not claimed here. They are #11622's to establish when it is rewritten on this.
Result (floor run 35446547321 on 359d182, FloorClean; the prediction held. The later heads change only test fixtures, not the parser)
Measured against #11700's run 35440523086 (baseline) and #11622's keyword run 35440687934, using
required-floor-claim-costeval_steps:aa3e61fagainst this branch'scb5950a), and the top movers are roster-counting witnesses (witness_admission,witness_exclusion_reconciliation,local_repo_wet_terminal, about +12%), whose cost tracks roster size as rows land on main, not parse cost.body_lowering_fixture_resolves: 67,077 / 92,974 / 67,118native_decl_selectionx2: 35,339 / 47,357 / 35,339arrow_body_form_eval_value_vertical_holds: 91,179 / 117,076 / 91,220lex_rule_dispatch_agrees_with_trying_every_rule: 67,026 exactly (keywords: 82,691).choice_planrows cost 295 to 7,095 steps each.An exact return to baseline is what the mechanism predicts (no literal terminal in any real grammar, so the same plan is compiled). A partial improvement would have been consistent with other explanations.
Head
2e23ec3250is a merge ofmainintocdc4407942. The heal workflow on main runsdag/gunbc/heal_candidate.dag, which this branch predated.cdc4407942is a strict ancestor of the head, so this PR's own contribution is carried whole and unchanged. The merge brings main's commits and nothing else. Those touch many files, includingnamespace_graft.dag, which main changed when #11694 landed its producer marker; that movement is main's, not this PR's. Floor run 35461610687 on this head is FloorClean, and it is the first run of the literal terminal alongside main's SH-2 grammar.Status at
2e23ec3250. Source is signed off at this exact head: both held findings on the evidence are closed, and the substance was re-verified on the merge head. All four checks pass and mergeStateStatus is CLEAN. The only outstanding requirement is a completed review bound to this head. That is an infrastructure gap, not a question about the change: two review processes on this PR have crashed without a verdict (68443 on1cc101e29d, 68584 on this head), and an approval on an ancestor head does not count as an exact-head review.🤖 Generated with Claude Code