Skip to content

#473 — refactor raw_scan masking to the sqlparser tokenizer (evaluation → adopt) - #475

Merged
cmbays merged 4 commits into
mainfrom
adapters-473-sqlparser-tokenizer
Jun 23, 2026
Merged

cmbays merged 4 commits into
mainfrom
adapters-473-sqlparser-tokenizer

Conversation

@cmbays

@cmbays cmbays commented Jun 23, 2026 •

Copy link
Copy Markdown
Contributor

Evaluation outcome for #473: ADOPT. Reimplements raw_scan's SQL-region lexical masking on sqlparser's tokenizer (the same engine the compiled side already parses with), deleting the hand-rolled SQL string/comment/dollar-quote lexer. Behavior-preserving: all goldens byte-identical.

What changed

  • mask_regions is now two passes: pass 1 mask_jinja (unchanged — sqlparser doesn't understand Jinja; reuses render::find_close/find_expr_close so the raw-span path and the zone path agree on Jinja boundaries on the same model), pass 2 mask_sql_tokens — Tokenizer::new(&GenericDialect{}, …).tokenize_with_location() then blank the byte span of every string / comment / dollar-quote / quoted-identifier token via the now-shared cte_engine::ByteIndex::byte_of.
  • Deleted the hand-rolled SQL machinery: scan_sql_quoted (CC 14, the doubled-quote + prefixed-backslash escape logic), scan_dollar_quote (CC 12), the line/block comment scanners, quote_has_string_prefix, and the SQL arms of classify_opener (which now only dispatches Jinja). mask_regions CC 16 → 10.
  • The single home of SQL string-escape handling is now the tokenizer itself — no hand-rolled escape code to drift, and raw & compiled cannot diverge on SQL lexing (both GenericDialect).

Soundness (the load-bearing property)

The masker may only ever OVER-mask (a name omitted → honest "no anchor") and must never UNDER-mask (a string/comment interior name leaks as a live span → a false anchor). Preserved on every axis:

  • TokenizerError (unterminated string, malformed U&'…\…') → blank the entire text = maximal over-mask (tokenizer_error_fails_closed_blanks_everything).
  • Malformed Jinja → whole model emits nothing (unchanged fail-closed).
  • Quoted identifiers (Word{quote_style:Some}, e.g. ")") are blanked, not left live — two honest-direction reasons: (a) paren-balance soundness (a ) inside select ")" as x is not a structural paren; leaving it live truncated the CTE body span — a false span, now pinned by paren_inside_quoted_identifier_does_not_truncate_cte_span); (b) behavior-parity (such names are already honestly omitted at the fill layer). An unquoted Word stays live (the only Word carrying a matchable name).
  • Line-comment trailing \n preserved — the tokenizer's SingleLineComment span includes its terminating newline; blanking it shifted downstream line/col (byte offsets stayed correct). Caught mid-build via 4 churned goldens, diagnosed not blindly regenerated, fixed to match the old scanner exactly → goldens back to byte-identical.

Verification

  • Independent adversarial soundness pass: no false anchor found across an exhaustive exotic-form sweep — raw/byte/triple-quote/national/hex strings, prefixed E'/U&'/N'/X', Postgres $func$-body nesting, mismatched dollar tags, CRLF/CR/U+2028 line endings, comment-glued names, placeholders, backtick identifiers. Verdict: sound; fail-closed verified.
  • 62 tests (was 57; +5: fail-closed-at-tokenizer + fail-closed-at-fill-layer + paren-inside-quoted-ident + unquoted-word-stays-live + quoted-ident-honestly-omitted).
  • Goldens byte-identical (git diff --exit-code -- examples/ clean — independently re-verified).
  • Gates green: fmt --check, clippy --all-targets --locked -D warnings, full nextest, crap4rs, cargo deny, cargo doc -D warnings.

Notes for review

  • One honest deviation from the issue's delete-list: balanced_close was kept — it's the structural paren-matcher for CTE-body extent (operates over the masked text, not a region scanner); deleting it would break cte_raw_span. Still sound because quoted-identifier interiors are now masked, so only live parens survive.
  • Second commit is a doc-only fix correcting stale "left live" quoted-identifier prose in the mask_regions module doc that contradicted is_maskable_token's (sound) behavior.
  • Follow-up Mechanically enforce is_maskable_token exhaustiveness vs sqlparser Token variants (dep-bump under-mask guard) #474 (priority:later, non-blocking): is_maskable_token is exhaustive over sqlparser 0.62 but uses matches! with an implicit wildcard, so a future dep bump adding a string variant could silently under-mask — tracked to add a compile-time exhaustiveness guard.

Settles the masker before S2 (#470) inherits it.

Closes #473.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • Refactor

    • Improved internal code organization and reusability of utility components.
  • Tests

    • Extended test coverage for internal masking and tokenization logic.

cmbays and others added 2 commits June 23, 2026 00:02
…e hand-rolled SQL lexer)

Reimplement raw_scan's SQL-region masking via sqlparser::tokenizer::Tokenizer
(GenericDialect — the same dialect cte_engine parses with, so raw & compiled
never diverge on SQL lexing). Two passes: (1) Jinja masking via the existing
offset-preserving render::find_close/find_expr_close helpers (sqlparser does not
understand Jinja); (2) tokenize the Jinja-masked text and blank the byte span of
every string/comment/dollar-quote token, with the tokenizer's own escape
handling. Fail-closed on a TokenizerError (blank the whole text — nothing leaks).

Deletes the hand-rolled SQL string/comment/dollar-quote machinery — scan_sql_quoted
(CC 14), scan_dollar_quote (CC 12), the SQL arms of classify_opener (CC 16→5),
scan_line_comment, scan_block_comment, quote_has_string_prefix. The subtle SQL
string-escape surface (doubled-quote, prefixed-backslash E'…'/U&'…') now lives in
the tokenizer, not in this module.

Soundness preserved (over-mask-only, never under-mask): a quoted identifier
(Token::Word{quote_style:Some}) is blanked too — both for paren-balance soundness
(an interior `)` is not a structural paren) and behavior-parity with the old
masker (quoted-ident names are already honestly omitted at the fill layer). A
line comment's terminating `\n` is preserved so emitted line/col stays faithful
to the true raw source (byte offsets are the anchor; the `\n` carries no name).

Goldens byte-identical (examples/ unchanged). 62 raw_scan tests (was 57: +5
pinning tokenizer-error fail-closed, paren-in-quoted-ident, bare-Word-live,
dialect-parity). Full nextest + bdd + fmt + clippy --locked + deny + crap4rs
green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0199AmBCec5kyVEEF1Qd7TSF
…cs to match masking behavior

The mask_regions module doc + lexical-region table still described quoted
identifiers as 'left live', contradicting is_maskable_token (which blanks
them, w.quote_style.is_some()) and its own detailed soundness rationale.
A doc asserting the inverse of the code on a soundness-load-bearing
decision is a correctness hazard; align the docs with the (sound,
over-masking) behavior. Doc-only — goldens + gates unaffected.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0199AmBCec5kyVEEF1Qd7TSF
@qodo-code-review

Copy link
Copy Markdown

Qodo reviews are paused for this user.

Troubleshooting steps vary by plan Learn more →

On a Teams plan?
Reviews resume once this user has a paid seat and their Git account is linked in Qodo.
Link Git account →

Using GitHub Enterprise Server, GitLab Self-Managed, or Bitbucket Data Center?
These require an Enterprise plan - Contact us
Contact us →

@coderabbitai

coderabbitai Bot commented Jun 23, 2026 •

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@cmbays, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 18 minutes and 7 seconds. Learn how PR review limits work.

Your organization has used up its prepaid credits, and credit purchases are no longer available. Enable the review add-on in the billing tab to keep reviews running — you're only billed for reviews past your plan's rate limits ($0.25/file).

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based credits.

🚦 How do rate limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan refill rate.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, the refill rate gradually slows as usage increases. The highest same-day bursts are limited more strictly.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: b94eb93c-b734-4ad0-b3f0-504dd644bb82

📥 Commits

Reviewing files that changed from the base of the PR and between a0e2f42 and 8763a0c.

📒 Files selected for processing (1)
  • src/adapters/raw_scan.rs
📝 Walkthrough

Walkthrough

ByteIndex in cte_engine.rs is widened to pub(crate). raw_scan.rs masking is refactored from a single forward pass into an explicit two-pass model: a Jinja pass followed by a sqlparser::Tokenizer-based SQL pass that replaces all hand-rolled SQL string/comment/dollar-quote scanning. New helpers mask_sql_tokens and is_maskable_token implement the SQL pass with fail-closed behavior.

Changes

Two-pass masking refactor

Layer / File(s) Summary
ByteIndex widened to pub(crate)
src/adapters/cte_engine.rs
ByteIndex struct, new, and byte_of are changed from private to pub(crate) so raw_scan can import and reuse the byte-offset helpers.
Imports and classify_opener simplification
src/adapters/raw_scan.rs
Imports add ByteIndex, GenericDialect, and tokenizer types; classify_opener is narrowed to {-led Jinja only, removing SQL branches that are now handled by the tokenizer pass.
Two-pass mask_regions, mask_sql_tokens, and is_maskable_token
src/adapters/raw_scan.rs
mask_regions gains explicit two-pass docs and control flow; mask_sql_tokens tokenizes via sqlparser::Tokenizer under GenericDialect, fails closed on error, uses ByteIndex for span math, and preserves trailing newlines on single-line comments; is_maskable_token enumerates all maskable Token variants.
Test updates and new tokenizer-pass contracts
src/adapters/raw_scan.rs
A test comment clarifies newline preservation; five new tests cover tokenizer-error blanking, fill omission on failure, quoted-identifier parens, unquoted word anchoring, and GenericDialect dollar-quote alignment.

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly related issues

Possibly related PRs

  • breezy-bays-labs/cute-dbt#46: Directly introduced ByteIndex for span-based raw_sql slicing; this PR widens its visibility to pub(crate) for reuse in raw_scan.
  • breezy-bays-labs/cute-dbt#452: Extended ByteIndex and span-slicing utilities in cte_engine.rs with SourceSpan; this PR builds on the same helpers by exposing them crate-wide.
  • breezy-bays-labs/cute-dbt#472: Introduced the original raw_scan masking/anchoring machinery that this PR refactors into the two-pass model.

Poem

🐇 Hop, hop, two passes now!
First Jinja's masked with a careful paw,
Then sqlparser takes its bow —
No hand-rolled quotes to gnaw.
ByteIndex shared, the bytes align,
Every span precisely fine! ✨

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly references the main refactoring effort (moving from hand-rolled SQL lexing to sqlparser tokenizer) and indicates the decision outcome (evaluation → adopt), directly corresponding to the core changes in raw_scan.rs.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch adapters-473-sqlparser-tokenizer

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

github-actions Bot commented Jun 23, 2026 •

Copy link
Copy Markdown
Contributor

📄 Rendered report preview

All golden examples regenerated cleanly.

🟡 Golden examples

Committed to examples/ and byte-identity gated — the canonical reports contributors and consumers browse. Stable across PRs.

Report View Download
prdiff-minidag-report.html ▶ Open ↗ ⬇ Download
jaffle-shop-report.html ▶ Open ↗ ⬇ Download
diff-showcase-report.html ▶ Open ↗ ⬇ Download
seed-showcase-report.html ▶ Open ↗ ⬇ Download
playground-report.html ▶ Open ↗ ⬇ Download
macro-heavy-report.html ▶ Open ↗ ⬇ Download

🐶 Live dogfood preview

This PR doesn't touch dbt-project/, so there's no live dogfood preview.

🧭 Explore preview

The two-page cute-dbt explore explorer — dag.html (model lineage) + tests.html (unit-test viewer). Same golden/live split as the report.

🟡 Golden explore

The committed examples/explore/ playground golden (the full synthetic playground manifest). Byte-identity gated in Example report check. Stable across PRs.

Page View Download
explore/dag.html ▶ Open ↗ ⬇ Download
explore/tests.html ▶ Open ↗ ⬇ Download

🐶 Live explore

This PR doesn't touch dbt-project/, so there's no live explore preview.

▶ Open ↗ opens the report or explorer in your browser in one
click — published to this repo's GitHub Pages under
/pr-475/.
⬇ Download fetches the same self-contained HTML as a workflow
artifact (auth-gated; works fully offline). Either way the report
makes zero external resource requests.

The Pages preview may take ~1 min to update after this comment
posts. On PRs from forks the Open link is unavailable (read-only
token) — use Download.

Alternative: GitHub CLI
# gh CLI >= 2.63 extracts into ./report-preview-playground/.
gh run download 28003299556 -R breezy-bays-labs/cute-dbt -n report-preview-playground
open report-preview-playground/playground-report.html

Posted by report-preview.yml for 8763a0c26a4c54ff462a7d90a0fe85c15096f49b. Affordance only — never blocks merge.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request refactors the SQL masking logic in src/adapters/raw_scan.rs into a robust two-pass process: a Jinja-masking pass followed by a SQL-tokenization pass using sqlparser's tokenizer. This replaces several hand-rolled scanners, improving consistency and correctness. The review feedback recommends adding a defensive guard to ensure start < end before calling blank to prevent potential panics.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment thread src/adapters/raw_scan.rs
cmbays and others added 2 commits June 23, 2026 00:55
Per gemini review on PR #475: guard blank() against an inverted (from >= to)
span. The upper bound was already clamped (to.min(len)); the inverted case
was unreachable today (byte_of clamps to len + token spans are ordered + the
line-comment end-=1 is guarded by end > start) but would panic if any future
caller passed an inverted span. The masker must fail closed, never crash the
render on a manifest-derived span, so make blank() total. Pinned by
blank_is_total_on_malformed_spans. Behavior on valid spans unchanged (goldens
byte-stable).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0199AmBCec5kyVEEF1Qd7TSF
…ingle_char_names)

The blank totality test used 5 single-char bindings (a..e), tripping
clippy::many_single_char_names under -D warnings. Rename to valid/clamped/
empty/inverted/past_end. Test-only; no behavior change.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0199AmBCec5kyVEEF1Qd7TSF
@cmbays
cmbays merged commit 593c10f into main Jun 23, 2026
45 checks passed
@cmbays
cmbays deleted the adapters-473-sqlparser-tokenizer branch June 23, 2026 05:05
github-actions Bot added a commit that referenced this pull request Jun 23, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Refactor raw_scan masking to the sqlparser tokenizer (evaluation — before S2)

1 participant