Skip to content

Reduce clean parser memo overhead - #117

Merged
tinovyatkin merged 3 commits into
mainfrom
codex/reduce-parser-memo-overhead
Jul 18, 2026
Merged

Reduce clean parser memo overhead#117
tinovyatkin merged 3 commits into
mainfrom
codex/reduce-parser-memo-overhead

Conversation

@tinovyatkin

@tinovyatkin tinovyatkin commented Jul 18, 2026

Copy link
Copy Markdown
Contributor

Summary

  • retain grammar-static empty-cycle reachability across parser resets, invalidating it when the ATN identity changes
  • classify clean fast-recognizer memo traffic from actual coordinate reuse, then skip the full memo table for sparse parses
  • periodically reopen the bounded probe so a repeat-heavy region that begins later in the token stream can promote memoization
  • keep recovery-pass memoization unchanged and cover retention, invalidation, sparse failures, promotion, and re-probing with focused tests

Root cause

The no-tree MySQL path resets and reuses one parser for every statement. Each reset discarded empty-cycle reachability even though it depends only on the grammar ATN, forcing the same graph analysis to run again.

The clean recognizer also paid for millions of hash-table lookups and inserts even when coordinates were effectively one-shot. The existing single-outcome probe could not observe a memo hit because the hit returned before its repeat counter ran, and once a sparse classification is selected it needs a bounded way to notice a later repeat-heavy phase.

The new policy observes keys before lookup. It promotes after eight repeated coordinates, enters sparse mode after a 4,096-key one-shot probe, and reopens that probe every 262,144 sparse visits. Recovery still uses the full memo because cached failures carry diagnostics.

Performance

Same-machine Apple M3 Pro, current merged PR #115 binary versus this head. Each binary ran the benchmark's 1 cold + 5 warm iterations and averaged the three fastest warm parse samples per fixture.

Run statements.txt bitrix_queries_cut.sql sakila-data.sql Parse total
merged head A 154 ms 131 ms 692 ms 977 ms
merged head B 156 ms 133 ms 697 ms 986 ms
this head A 57 ms 53 ms 564 ms 674 ms
this head B 56 ms 52 ms 555 ms 663 ms

The two-run mean falls from 981.5 ms to 668.5 ms, a 31.9% parse reduction over merged PR #115. A permanently sparse experiment measured 654 ms; bounded re-probing retains nearly all of that gain while protecting late repeat-heavy regions.

Combined with the issue's v0.11.0 baseline, parse time falls from 2,637 ms to about 669 ms. At the measured ~1,537 ms lex time, the Rust total is about 2,206 ms versus the issue's 2,656 ms C++ total.

Closes #113.

Validation

  • cargo test --locked --all-targets --all-features
  • cargo clippy --locked --all-targets --all-features -- -D warnings
  • cargo run --release --quiet --bin antlr4-runtime-testsuite (357 passed, 0 failed, 0 skipped)
  • tests/kotlin-parity/run.sh ... (9/9 snippets match)
  • tests/javascript-parity/run.sh ... (6/6 snippets match)
  • tests/typescript-parity/run.sh ... (5/5 snippets match)
  • cargo fmt --check -- src/parser.rs
  • git diff --check
  • same-host MySQL A/B benchmarks with zero syntax failures

Summary by CodeRabbit

  • Performance

    • Improved parser memoization heuristics to make recognition more efficient across different parsing scenarios.
    • Enhanced cache handling when switching between grammars, helping maintain accurate parsing behavior.
  • Reliability

    • Improved handling of cyclic and empty parsing paths.
    • Expanded automated coverage for memoization modes, cache resets, and grammar transitions.

@coderabbitai

coderabbitai Bot commented Jul 18, 2026

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

The parser replaces single-outcome memoization with adaptive clean-memo probe modes, integrates those modes into fast-recognizer lookup and insertion, and makes empty-cycle caching ATN-sensitive. Tests cover mode transitions, sparse memo behavior, and cache reuse across resets and ATN changes.

Changes

Parser memoization and cache behavior

Layer / File(s) Summary
Clean-memo state and probe transitions
src/parser.rs
Introduces CleanMemoMode, probe thresholds, adaptive mode transitions, and updated parser initialization and per-parse reset state.
Fast-recognizer memo lookup and insertion
src/parser.rs
Gates clean-pass memo reuse by key and disables memo insertion in Sparse mode; tests cover sparse behavior and mode transitions.
ATN-sensitive empty-cycle cache
src/parser.rs
Tracks the ATN associated with empty-cycle results, invalidates the cache when ATNs change, and tests reset and recomputation behavior.

Estimated code review effort: 3 (Moderate) | ~20 minutes

Possibly related PRs

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Linked Issues check ✅ Passed The parser memoization and cache changes directly address #113's parsing-prediction overhead goals.
Out of Scope Changes check ✅ Passed The changes stay focused on parser memoization, caching, and tests with no evident unrelated scope.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately captures the main goal of reducing parser memoization overhead.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch codex/reduce-parser-memo-overhead

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@tinovyatkin
tinovyatkin marked this pull request as ready for review July 18, 2026 23:10
@github-actions

Copy link
Copy Markdown

Copy/Paste Detection

Found 16 duplication(s) across 1 changed Rust file(s) (threshold: 100 tokens).

Show duplications

Found a 27 line (145 tokens) duplication in the following files:

  • Starting at line 14508 of src/parser.rs
  • Starting at line 14641 of src/parser.rs
    fn generated_match_token_recovers_missing_token_from_context_follow() {
        let atn = generated_match_recovery_atn();
        let data = RecognizerData::new(
            "Mini.g4",
            Vocabulary::new(
                [None, Some("'X'"), Some("'Y'")],
                [None, Some("X"), Some("Y")],
                [None::<&str>, None, None],
            ),
        );
        let mut parser = BaseParser::new(
            CommonTokenStream::new(Source {
                tokens: vec![TestToken::eof("parser-test", 3, 1, 3)],
                index: 0,
            }),
            data,
        );
        parser.rule_context_stack = vec![
            RuleContextFrame {
                rule_index: 0,
                invoking_state: 0,
            },
            RuleContextFrame {
                rule_index: 1,
                invoking_state: 1,
            },
        ];

Found a 27 line (127 tokens) duplication in the following files:

  • Starting at line 13466 of src/parser.rs
  • Starting at line 13538 of src/parser.rs
        let mut atn = ParserAtnBuilder::new(2);
        assert_eq!(
            atn.add_state(AtnStateKind::RuleStart, Some(0))
                .expect("state")
                .index(),
            0
        );
        assert_eq!(
            atn.add_state(AtnStateKind::BlockStart, Some(0))
                .expect("state")
                .index(),
            1
        );
        assert_eq!(
            atn.add_state(AtnStateKind::Basic, Some(0))
                .expect("state")
                .index(),
            2
        );
        assert_eq!(
            atn.add_state(AtnStateKind::Basic, Some(0))
                .expect("state")
                .index(),
            3
        );
        assert_eq!(
            atn.add_state(AtnStateKind::BlockEnd, Some(0))

Found a 33 line (126 tokens) duplication in the following files:

  • Starting at line 8264 of src/parser.rs
  • Starting at line 8302 of src/parser.rs
                    if self.fast_parser_predicate_matches(predicate_context, transition, index) {
                        let boundary = left_recursive_boundary(atn, state, target);
                        outcomes.extend(
                            self.recognize_state_fast(
                                atn,
                                FastRecognizeRequest {
                                    state_number: target,
                                    stop_state,
                                    index,
                                    rule_start_index,
                                    decision_start_index: next_decision_start_index,
                                    precedence,
                                    depth: depth + 1,
                                    recovery_symbols: Rc::clone(&epsilon_recovery_symbols),
                                    recovery_state: epsilon_recovery_state,
                                },
                                FastRecognizeScratch {
                                    predicate_context,
                                    visiting,
                                    memo,
                                    expected,
                                },
                            )
                            .into_iter()
                            .map(|mut outcome| {
                                if let Some(rule_index) = boundary {
                                    let boundary = self.arena_boundary_node(rule_index);
                                    self.defer_fast_outcome_node(&mut outcome, boundary);
                                }
                                outcome
                            }),
                        );
                    } else {

Found a 32 line (125 tokens) duplication in the following files:

  • Starting at line 8229 of src/parser.rs
  • Starting at line 8265 of src/parser.rs
  • Starting at line 8303 of src/parser.rs
                    let boundary = left_recursive_boundary(atn, state, target);
                    outcomes.extend(
                        self.recognize_state_fast(
                            atn,
                            FastRecognizeRequest {
                                state_number: target,
                                stop_state,
                                index,
                                rule_start_index,
                                decision_start_index: next_decision_start_index,
                                precedence,
                                depth: depth + 1,
                                recovery_symbols: Rc::clone(&epsilon_recovery_symbols),
                                recovery_state: epsilon_recovery_state,
                            },
                            FastRecognizeScratch {
                                predicate_context,
                                visiting,
                                memo,
                                expected,
                            },
                        )
                        .into_iter()
                        .map(|mut outcome| {
                            if let Some(rule_index) = boundary {
                                let boundary = self.arena_boundary_node(rule_index);
                                self.defer_fast_outcome_node(&mut outcome, boundary);
                            }
                            outcome
                        }),
                    );
                }

Found a 22 line (125 tokens) duplication in the following files:

  • Starting at line 13492 of src/parser.rs
  • Starting at line 13564 of src/parser.rs
            atn.add_state(AtnStateKind::BlockEnd, Some(0))
                .expect("state")
                .index(),
            4
        );
        assert_eq!(
            atn.add_state(AtnStateKind::RuleStop, Some(0))
                .expect("state")
                .index(),
            5
        );
        atn.set_rule_to_start_state(vec![0])
            .expect("rule start states");
        atn.set_rule_to_stop_state(vec![5])
            .expect("rule stop states");
        atn.add_decision_state(1).expect("decision state");
        atn.add_transition(0, ParserTransitionSpec::Epsilon { target: 1 })
            .expect("transition");
        atn.add_transition(
            1,
            ParserTransitionSpec::Atom {
                target: 2,

Found a 33 line (115 tokens) duplication in the following files:

  • Starting at line 9183 of src/parser.rs
  • Starting at line 9256 of src/parser.rs
                        outcomes.extend(
                            self.recognize_state(
                                atn,
                                RecognizeRequest {
                                    state_number: *target,
                                    stop_state,
                                    index,
                                    rule_start_index,
                                    decision_start_index: next_decision_start_index,
                                    init_action_rules,
                                    predicates,
                                    semantics,
                                    rule_args,
                                    member_actions,
                                    return_actions,
                                    local_int_arg,
                                    member_values: member_values.clone(),
                                    return_values: return_values.clone(),
                                    rule_alt_number: next_alt_number,
                                    track_alt_numbers,
                                    consumed_eof,
                                    precedence,
                                    depth: depth + 1,
                                    recovery_symbols: epsilon_recovery_symbols.clone(),
                                    recovery_state: epsilon_recovery_state,
                                },
                                visiting,
                                memo,
                                expected,
                            )
                            .into_iter()
                            .map(|mut outcome| {
                                prepend_decision(&mut outcome, decision);

Found a 15 line (113 tokens) duplication in the following files:

  • Starting at line 14566 of src/parser.rs
  • Starting at line 14850 of src/parser.rs
    fn generated_match_token_counts_single_token_deletion_recovery() {
        let atn = generated_match_recovery_atn();
        let data = RecognizerData::new(
            "Mini.g4",
            Vocabulary::new(
                [None, Some("'X'"), Some("'Y'"), Some("'Z'")],
                [None, Some("X"), Some("Y"), Some("Z")],
                [None::<&str>, None, None, None],
            ),
        );
        let mut parser = BaseParser::new(
            CommonTokenStream::new(Source {
                tokens: vec![
                    TestToken::new(3).with_text("z"),
                    TestToken::new(2).with_text("y"),

Found a 12 line (112 tokens) duplication in the following files:

  • Starting at line 12447 of src/parser.rs
  • Starting at line 12530 of src/parser.rs
        let mut atn = ParserAtnBuilder::new(1);
        for (state, kind, rule) in [
            (0, AtnStateKind::RuleStart, 0),
            (1, AtnStateKind::StarLoopEntry, 0),
            (2, AtnStateKind::Basic, 0), // ops hub
            (3, AtnStateKind::Basic, 0), // shift prec
            (4, AtnStateKind::Basic, 0), // shift first >
            (5, AtnStateKind::Basic, 0), // shift second >
            (6, AtnStateKind::Basic, 0), // rel prec
            (7, AtnStateKind::Basic, 0), // rel >
            (8, AtnStateKind::LoopEnd, 0),
            (9, AtnStateKind::RuleStop, 0),

Found a 22 line (112 tokens) duplication in the following files:

  • Starting at line 13903 of src/parser.rs
  • Starting at line 14104 of src/parser.rs
    fn predicate_after_token_atn() -> Atn {
        let mut atn = ParserAtnBuilder::new(2);
        assert_eq!(
            atn.add_state(AtnStateKind::RuleStart, Some(0))
                .expect("state")
                .index(),
            0
        );
        assert_eq!(
            atn.add_state(AtnStateKind::Basic, Some(0))
                .expect("state")
                .index(),
            1
        );
        assert_eq!(
            atn.add_state(AtnStateKind::Basic, Some(0))
                .expect("state")
                .index(),
            2
        );
        assert_eq!(
            atn.add_state(AtnStateKind::Basic, Some(0))

Found a 22 line (111 tokens) duplication in the following files:

  • Starting at line 12312 of src/parser.rs
  • Starting at line 13903 of src/parser.rs
  • Starting at line 14104 of src/parser.rs
    fn left_recursive_loop_with_caller_follow_atn(caller_symbol: i32) -> Atn {
        let mut atn = ParserAtnBuilder::new(2);
        assert_eq!(
            atn.add_state(AtnStateKind::RuleStart, Some(0))
                .expect("state")
                .index(),
            0
        );
        assert_eq!(
            atn.add_state(AtnStateKind::Basic, Some(0))
                .expect("state")
                .index(),
            1
        );
        assert_eq!(
            atn.add_state(AtnStateKind::Basic, Some(0))
                .expect("state")
                .index(),
            2
        );
        assert_eq!(
            atn.add_state(AtnStateKind::RuleStart, Some(1))

Found a 14 line (110 tokens) duplication in the following files:

  • Starting at line 14237 of src/parser.rs
  • Starting at line 16691 of src/parser.rs
    fn parser_matches_token_and_reports_mismatch() {
        let source = Source {
            tokens: vec![
                TestToken::new(1).with_text("x"),
                TestToken::eof("parser-test", 1, 1, 1),
            ],
            index: 0,
        };
        let data = RecognizerData::new(
            "Mini.g4",
            Vocabulary::new([None, Some("'x'")], [None, Some("X")], [None::<&str>, None]),
        );
        let mut parser = BaseParser::new(CommonTokenStream::new(source), data);
        let matched = parser.match_token(1).expect("token 1 should match");

Found a 13 line (109 tokens) duplication in the following files:

  • Starting at line 14237 of src/parser.rs
  • Starting at line 16716 of src/parser.rs
    fn parser_matches_token_and_reports_mismatch() {
        let source = Source {
            tokens: vec![
                TestToken::new(1).with_text("x"),
                TestToken::eof("parser-test", 1, 1, 1),
            ],
            index: 0,
        };
        let data = RecognizerData::new(
            "Mini.g4",
            Vocabulary::new([None, Some("'x'")], [None, Some("X")], [None::<&str>, None]),
        );
        let mut parser = BaseParser::new(CommonTokenStream::new(source), data);

Found a 22 line (108 tokens) duplication in the following files:

  • Starting at line 7131 of src/parser.rs
  • Starting at line 7519 of src/parser.rs
    ) -> Option<RecognizeOutcome> {
        let (error_index, message) = self.expected_error_message(rule_index, start_index, expected);
        let diagnostic = diagnostic_for_token(self.token_at(error_index), message);
        let mut next_index = error_index;
        loop {
            let symbol = self.token_type_at(next_index);
            if sync_symbols.contains(&symbol) {
                if next_index == error_index {
                    return None;
                }
                break;
            }
            if symbol == TOKEN_EOF {
                break;
            }
            let after = self.consume_index(next_index, symbol);
            if after == next_index {
                break;
            }
            next_index = after;
        }
        let mut nodes = NodeSeqId::EMPTY;

Found a 15 line (108 tokens) duplication in the following files:

  • Starting at line 16970 of src/parser.rs
  • Starting at line 16994 of src/parser.rs
    fn outcome_ties_keep_later_non_recursive_alternative() {
        let arena = RecognitionArena::default();
        let first = RecognizeOutcome {
            index: 1,
            consumed_eof: false,
            alt_number: 0,
            member_values: BTreeMap::new(),
            return_values: BTreeMap::new(),
            diagnostics: DiagnosticSeqId::EMPTY,
            decisions: Vec::new(),
            actions: vec![ParserAction::new(1, 0, 0, None)],
            nodes: NodeSeqId::EMPTY,
        };
        let second = RecognizeOutcome {
            actions: vec![ParserAction::new(2, 0, 0, None)],

Found a 16 line (105 tokens) duplication in the following files:

  • Starting at line 6328 of src/parser.rs
  • Starting at line 6898 of src/parser.rs
        let start_state = atn.rule_to_start_state().get(rule_index).ok_or_else(|| {
            AntlrError::Unsupported(format!("rule {rule_index} has no start state"))
        })?;
        let stop_state = atn
            .rule_to_stop_state()
            .get(rule_index)
            .filter(|state| *state != usize::MAX)
            .ok_or_else(|| {
                AntlrError::Unsupported(format!("rule {rule_index} has no stop state"))
            })?;

        let start_index = self.current_visible_index();
        self.clear_prediction_diagnostics();
        self.reset_per_parse_caches();
        self.reset_recognition_arena();
        let caller_follow_state = self.pending_invoking_follow_state(atn);

Found a 13 line (100 tokens) duplication in the following files:

  • Starting at line 6063 of src/parser.rs
  • Starting at line 6087 of src/parser.rs
        let mut expected = BTreeSet::new();
        for index in (1..self.rule_context_stack.len()).rev() {
            let invoking_state = self.rule_context_stack[index].invoking_state;
            let Ok(state_number) = usize::try_from(invoking_state) else {
                continue;
            };
            let Some(Transition::Rule { follow_state, .. }) = atn
                .state(state_number)
                .and_then(|state| state.transitions().first())
                .map(ParserTransition::data)
            else {
                continue;
            };

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request refactors the clean-pass memoization strategy in the parser to support periodic reprobing in sparse mode, renaming single-outcome memo fields and constants to a more general clean memoization terminology. Additionally, it optimizes the empty_cycle_cache to survive parser resets by associating it with a companion ATN key, invalidating it lazily only when the ATN changes. Feedback suggests using atn.state_count() instead of atn.states().len() to avoid constructing a temporary iterator when resizing the cycle cache.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment thread src/parser.rs Outdated
@tinovyatkin

Copy link
Copy Markdown
Contributor Author

@coderabbitai review

@tinovyatkin

Copy link
Copy Markdown
Contributor Author

@codex review

@coderabbitai

coderabbitai Bot commented Jul 18, 2026

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@tinovyatkin

Copy link
Copy Markdown
Contributor Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: a01eca70ce

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/parser.rs Outdated
@tinovyatkin

Copy link
Copy Markdown
Contributor Author

@codex review

@tinovyatkin tinovyatkin changed the title [codex] Reduce clean parser memo overhead Reduce clean parser memo overhead Jul 18, 2026
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Hooray!

Reviewed commit: 9d198d5fe4

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@tinovyatkin
tinovyatkin merged commit 01e0eb7 into main Jul 18, 2026
11 checks passed
@tinovyatkin
tinovyatkin deleted the codex/reduce-parser-memo-overhead branch July 18, 2026 23:33
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

perf(parser): ~8x slower than antlr4-cpp at parsing (same-machine) — broad prediction gap, not just sakila

1 participant