Skip to content

adapters: askama renderer reproducing Claude Design HTML (#11) - #38

Merged
cmbays merged 8 commits into
mainfrom
adapters-11-renderer
May 23, 2026
Merged

cmbays merged 8 commits into
mainfrom
adapters-11-renderer

Conversation

@cmbays

@cmbays cmbays commented May 23, 2026 •

Copy link
Copy Markdown
Contributor

Closes #11.

PR 8b — implement the v0.1 interactive report.html via askama 0.16,
reproducing the Claude Design hand-back.

What ships

  • src/adapters/render.rs — composition layer between domain
    types, the CTE engine, and the askama template. Owns per-model
    payload assembly, node-role classification (final / import /
    transform), clean-import-CTE binding via ref('NAME') parsing,
    and the run-loop integration.
  • templates/report.html — askama 0.16 template emitting the
    design's full DOM contract: diff-scope banner, model + test cascade
    selectors, CTE DAG host, two-panel region (Node detail / All inputs
    · Expected), JSON payload carrier
    (<script type="application/json" id="cute-dbt-data">). The
    vendored asset bundle (Sakura · jQuery · DataTables · Mermaid UMD)
    is inlined via include_str! constants from asset_embed.rs.
  • interaction.js adapted from the design hand-back verbatim,
    with the dormant dark-mode infrastructure stripped (YAGNI) and the
    Mermaid <g> selector runtime-constructed
    ([id^="${svgId}-flowchart-"], fallback [id*='-flowchart-'])
    per ADR-4 amendment 2026-05-22.
  • tests/render_integration.rs — fixture-driven integration
    coverage: asset bundle assembly, design DOM contract, payload
    shape, JSON carrier integrity, structured resource-ref egress
    lint, and a chrome-only insta snapshot of the rendered
    jaffle-shop report.

What was removed

  • MERMAID_INIT constant + its pin test from asset_embed.rs — the
    real init lives in the inlined interaction script with
    startOnLoad: false (ADR-4 amendment).
  • smoke_report_html + its tests + the escape_label /
    mermaid_block helpers — the real renderer replaces them.
  • tests/asset_embed.rs — replaced by tests/render_integration.rs.
  • ARCHITECTURE.md §6 (synthetic-only fixture invariant) — relocated
    to CONTRIBUTING.md (human contributors) + AGENTS.md (AI agents).
    The mechanism (tests/fixtures/MANIFEST.toml + listed-file gate)
    is unchanged; the rule is contributor/build-hygiene, not
    architectural.

Run-loop wiring

src/cli/mod.rs::render widens to thread current / in_scope /
models_in_scope / baseline_label into
adapters::render::render_report. parse_ctes() stays as a named
no-op call site; per-model CTE parsing happens inside the renderer
during payload assembly, so the four-stage diagram
(scope → preflight_compiled → parse_ctes → render) stays greppable.

Discovery #1 — folded in

features/report_generation.feature dropped the stale "And the
banner states the v0.1 fidelity limit 'model body changes'" line.
The body-checksum scoping note moved to README.md. The other four
.feature files were scanned and need no corrections.

Acceptance criteria

  • askama 0.16 template emitting the design's DOM contract
  • All vendored assets inlined via include_str!
  • interaction.js inlined verbatim, dark-mode stripped, escape
    hardening preserved (askama's | json filter escapes <>&')
  • DataTables init: paging: false, info: false,
    searching: false, scrollX: false, ordering: true
  • Node-role classification at the render layer
  • Clean-import-CTE binding via ref('NAME') parsing
  • Model-first JSON blob shape
  • "0 unit tests wired" empty state surfaced per-model
  • Edge color rendering keyed by EdgeType (incl. union_all +
    union_distinct as dashed orange)
  • Drop the v0.1 fidelity banner from the rendered HTML
  • Delete ARCHITECTURE.md §6; relocate to CONTRIBUTING.md +
    AGENTS.md
  • Drop MERMAID_INIT + smoke_report_html + their tests
  • Wire renderer into run-loop render step
  • Insta snapshot of rendered HTML for the jaffle-shop fixture
  • Strict CRAP threshold (15) on every strict-keyed file touched
  • cargo-mutants on the strict scope: 0 surviving mutants

Validation gates (all green locally)

Gate Result
cargo test 254 passed across 12 suites
cargo clippy --all-targets --locked -- -D warnings clean
cargo fmt --check clean
cargo deny check advisories ok, bans ok, licenses ok, sources ok
cargo doc -D warnings clean
cargo llvm-cov nextest --fail-under-lines 85 98.94% lines / 99.08% functions
lefthook run pre-push --all-files 9/9 green
crap4rs strict (15) on src/adapters/render.rs max 8.19 — PASS
cargo mutants --file src/adapters/render.rs 161 mutants, 132 caught, 0 missed, 29 unviable

🤖 Generated with Claude Code

Summary by CodeRabbit

Release Notes

  • New Features

    • Introduced interactive HTML report generation with click-to-inspect CTE dependency graphs, SQL syntax highlighting, and copy-to-clipboard functionality.
    • Added edge coloring in dependency DAG visualizations based on relationship types.
    • Included example jaffle-shop report demonstrating full end-to-end report rendering.
  • Documentation

    • Expanded README with clearer report functionality descriptions.
    • Added comprehensive fixture guidelines in contributing documentation.
  • Tests

    • Added integration tests validating HTML report generation, asset inlining, and payload structure.
  • Chores

    • Updated CI workflow to validate report consistency.

Review Change Stack

cmbays and others added 2 commits May 23, 2026 11:45
PR 8b — implement the v0.1 interactive `report.html` via askama 0.16.

Closes #11.

## What ships

- `src/adapters/render.rs` — composition layer between domain types, the
  CTE engine, and the askama template. Owns per-model payload assembly,
  node-role classification (`final` / `import` / `transform`), clean-import-CTE
  binding via `ref('NAME')` parsing, and the run-loop integration. CRAP
  threshold strict (15); the highest score on this file is 8.19.
- `templates/report.html` — askama 0.16 template emitting the design's
  full DOM contract: diff-scope banner, model + test cascade selectors,
  CTE DAG host, two-panel region, JSON payload carrier
  (`<script type="application/json" id="cute-dbt-data">`). The vendored
  asset bundle (Sakura, jQuery, DataTables, Mermaid UMD) is inlined via
  `include_str!` constants from `asset_embed.rs`. `interaction.js`
  adapted verbatim from the design hand-back with the dormant dark-mode
  infrastructure stripped (YAGNI) and the Mermaid `<g>` selector
  runtime-constructed (`[id^="${svgId}-flowchart-"]`) per ADR-4
  amendment 2026-05-22.
- `tests/render_integration.rs` — fixture-driven integration coverage:
  asset bundle assembly, design DOM contract, payload shape, JSON
  carrier integrity, structured resource-ref egress lint, and a chrome-
  only insta snapshot of the rendered jaffle-shop report.

## What was removed

- `MERMAID_INIT` constant + its pin test from `asset_embed.rs` — the
  real init lives in the inlined interaction script with `startOnLoad:
  false` (ADR-4 amendment); a Rust-side constant would invite drift.
- `smoke_report_html` + its tests + the `escape_label` / `mermaid_block`
  helpers — the real renderer replaces them.
- `tests/asset_embed.rs` — replaced by `tests/render_integration.rs`
  which exercises the real renderer end-to-end.
- ARCHITECTURE.md §6 (synthetic-only fixture invariant) — the rule is
  contributor/build-hygiene rather than architectural; relocated to
  CONTRIBUTING.md (human contributors) and AGENTS.md (AI agents).
  Section numbers tightened.

## Run-loop wiring

`src/cli/mod.rs::render` widens to thread `current` / `in_scope` /
`models_in_scope` / `baseline_label` into `adapters::render::render_report`.
`parse_ctes()` stays as a named no-op call site; per-model CTE parsing
happens inside the renderer during payload assembly, so the four-stage
diagram (`scope → preflight_compiled → parse_ctes → render`) stays
greppable.

## Discovery #1

Folded into this PR: `features/report_generation.feature` dropped the
stale "And the banner states the v0.1 fidelity limit 'model body
changes'" line — the diff-scope banner no longer carries that wording
(per AC). The body-checksum scoping note moved to README.md.

## Validation gates

- `cargo test`: 254 passed across 12 suites
- `cargo clippy --all-targets --locked -- -D warnings`: clean
- `cargo fmt --check`: clean
- `cargo deny check`: advisories ok, bans ok, licenses ok, sources ok
- `RUSTDOCFLAGS="-D warnings" cargo doc --no-deps --document-private-items
  --locked`: clean
- `lefthook run pre-push --all-files`: 9/9 gates green (fmt, clippy,
  test, deny, docs, baseline-required-grep, feature-count,
  fixture-manifest-gate, non-mirror-guard)
- crap4rs strict scoring on `src/adapters/render.rs`: max 8.19,
  threshold 15 — PASS
- `cargo mutants --file src/adapters/render.rs`: 161 mutants tested,
  132 caught, 0 missed, 29 unviable — 100% kill on the strict scope

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
The previous XSS test embedded `</script><script>alert(1)</script>` into
a `WITH "<name>" AS (...)` clause. sqlparser refuses that as an
identifier so `parse_cte_graph` returns Err, the renderer falls back to
an empty graph, and the hostile compiled_code lands in `compiled_sql`
under the model's bare name instead of taking the CTE-name-payload
path the docstring claimed.

Rewrite the test so the hostile value rides on a unit-test `tags` entry
— manifest-side YAML metadata that does not pass through any parser
before landing in the payload. Same structural property exercised
(askama `| json` escapes `<>&'`), now under the path the docstring
describes. Also tighten the assertion: the payload carrier `<script
type="application/json" id="cute-dbt-data">` opens exactly once, so
the hostile string can't smuggle a second carrier-open tag in.

Tests: 254 passed (still 100% green).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
@gemini-code-assist

Copy link
Copy Markdown

Warning

Gemini encountered an error creating the review. You can try again by commenting /gemini review.

@coderabbitai

coderabbitai Bot commented May 23, 2026 •

Copy link
Copy Markdown

Warning

Review limit reached

@cmbays, we couldn't start this review because you've used your available PR reviews for now.

Your plan currently allows 1 review/hour. Refill in 38 minutes and 14 seconds.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more review capacity refills, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans have higher rate limits than trial, open-source, and free plans. In all cases, review capacity refills continuously over time.

Please see our FAQ for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: e6fdbd4b-41f2-47d8-8061-28d6f24e0610

📥 Commits

Reviewing files that changed from the base of the PR and between 8a7093a and 31ddda6.

⛔ Files ignored due to path filters (1)
  • tests/snapshots/render_integration__rendered_chrome_jaffle_shop.snap is excluded by !**/*.snap
📒 Files selected for processing (4)
  • examples/jaffle-shop-report.html
  • src/adapters/render.rs
  • templates/report.html
  • tests/render_integration.rs
📝 Walkthrough

Walkthrough

This PR delivers the v0.1 interactive HTML report by implementing an Askama 0.16 template and Rust renderer module that compiles manifest/CTE data into a per-model, per-test JSON payload and final HTML with DAG visualization, clickable node inspection, and SQL syntax highlighting. The renderer integrates into the CLI pipeline, removes obsolete stubs, and includes comprehensive integration tests with artifact integrity guardrails.

Changes

Interactive HTML Report Renderer (Askama 0.16)

Layer / File(s) Summary
Renderer data contract and core types
src/adapters/render.rs (lines 1–340)
ReportPayload, ModelPayload, NodePayload, EdgePayload, DagPayload, test/fixture payloads, and NodeRole enum (Final/Import/Transform) define the shape consumed by the template. edge_type_wire_key() maps EdgeType variants to snake_case JS keys. parse_ref_name(), classify_node_role(), and bind_import_to_given() implement node role classification and import-CTE binding logic.
Renderer payload assembly
src/adapters/render.rs (lines 351–669)
build_payload() walks in-scope models and assembles the complete ReportPayload; build_model_payload() constructs DAG nodes/edges and test payloads; build_test_payload() parses ref('NAME') from unit-test inputs and binds them to import CTEs via name match (pass 1) or leaf table reference match (pass 2). Graceful fallback to full compiled SQL when CTE graph is empty or parse fails.
Renderer unit tests
src/adapters/render.rs (lines 670–1387)
Comprehensive coverage: ref-name parsing, import binding with body-table fallback, node-role classification (terminal detection, import-shape heuristic), edge wire-key/serde consistency, banner text singular/plural, payload compilation under parse failure, XSS resistance of JSON script carrier, and asset egress validation.
Askama template and interactive UI
templates/report.html (lines 1–861)
HTML document with inline Sakura CSS styling, DOM contract matching Claude Design. JavaScript layer (wrapped in {% raw %}) parses #cute-dbt-data JSON, renders model/test dropdowns and diff-scope banner, builds and renders Mermaid DAG with securityLevel: strict, binds Mermaid runtime node IDs to original DAG nodes, renders left-panel SQL inspection with copy-to-clipboard affordance (Clipboard API + textarea fallback), SQL syntax highlighting, and responsive panel reflow. DataTables init with paging/searching disabled.
Adapter module export and dependency
src/adapters/mod.rs, Cargo.toml
Adds pub mod render; export and askama 0.16 with serde_json feature to production dependencies.
Renderer integration tests
tests/render_integration.rs (215 lines), .cargo/mutants.toml, crap4rs.toml
Five test scenarios validate asset inlining (Sakura, jQuery, DataTables, Mermaid), DOM contract (sections, classes, selectors, payload carrier), JSON payload structure (baseline, models, tests, dag), golden-snapshot stability, and zero external resource egress (no src=, @import, http URLs; data: favicon only). Mutation testing and CRAP threshold 15 configured for src/adapters/render.rs.
CLI run-loop integration
src/cli/mod.rs (lines 6–178)
Wires render_report() into the render stage; derives baseline_label from baseline manifest path and delegates report generation to the Askama renderer. Updated execute() to pass full rendering context (current/baseline manifests, in-scope sets, output path).
Asset embedding cleanup
src/adapters/asset_embed.rs (lines 2–88)
Removes MERMAID_INIT constant and smoke_report_html() function (now handled by inlined template JS). Updates documentation and narrows tests to asset-embedding version assertions and empty favicon URI check.
CI artifact integrity and vocabulary guards
.github/workflows/ci.yml (lines 179–489)
edge-vocab-completeness job validates EdgeType coverage across render pipeline: extracts variants from src/domain/cte.rs, checks presence in src/adapters/render.rs via edge_type_wire_key, converts to snake_case, and verifies keys exist in templates/report.html's JOIN_COLORS. New example-report-up-to-date job regenerates examples/jaffle-shop-report.html via CLI and enforces byte-identical output.
Git configuration for generated artifacts
.gitattributes
Marks examples/*.html with -diff linguist-generated to suppress unreadable HTML diffs.
Documentation, specifications, and release notes
ARCHITECTURE.md, CONTRIBUTING.md, AGENTS.md, README.md, CHANGELOG.md, examples/README.md, features/report_generation.feature
Relocates synthetic-only fixture invariant from Architecture to CONTRIBUTING/AGENTS; updates README with v0.1 fidelity notes and edge-type details; documents examples and CI artifact guardrail; updates CHANGELOG with feature/maintenance entries; updates feature spec banner assertion.

Sequence Diagram(s)

sequenceDiagram
  participant CLI as CLI execute()
  participant Scope as scope stage
  participant Preflight as preflight_compiled
  participant ParseCTE as parse_ctes (no-op)
  participant Render as render stage
  participant RenderReport as render_report()
  participant Askama as Askama::render()
  participant Output as report.html
  
  CLI->>Scope: current, baseline, models_in_scope
  Scope->>Preflight: in_scope, models_in_scope
  Preflight->>ParseCTE: in_scope, models_in_scope (no mutation)
  ParseCTE->>Render: ready for render
  Render->>Render: derive baseline_label from path
  Render->>RenderReport: current, baseline_label, in_scope, models_in_scope, out_path
  RenderReport->>RenderReport: build_payload (walk models, bind tests)
  RenderReport->>Askama: ReportTemplate + ReportPayload
  Askama->>Askama: inline assets, embed JSON, render template
  Askama->>Output: write HTML
  Output-->>CLI: io::Result
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~60 minutes

Possibly related issues

  • breezy-bays-labs/cute-dbt#11: The PR directly implements the acceptance criteria for the Askama renderer with all required data shapes, node-role classification, import-CTE binding via ref('NAME') parsing, edge-color rendering by EdgeType, empty-state handling, asset inlining, integration tests, and CI artifact guards.

Possibly related PRs

  • breezy-bays-labs/cute-dbt#26: This PR removes the MERMAID_INIT constant and smoke_report_html() function that were introduced earlier; the new renderer replaces both with Askama-based inlined initialization and template-driven HTML generation.

Poem

🐰 A template springs forth with Askama's grace,
Where JSON blooms and DAGs find their place,
Click nodes to inspect the SQL within,
While edge-types paint their colors thin—
From CLI's run loop to browser's view,
The cute-dbt report is finally true! ✨

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The PR title clearly identifies the main change: implementing an askama renderer for the interactive HTML report using the Claude Design.
Linked Issues check ✅ Passed All major acceptance criteria from issue #11 are satisfied: askama template, inlined assets, interaction.js integration, DataTables config, node-role classification, import-CTE binding, model-first JSON payload, empty states, edge coloring, fixture hygiene relocation, constant/test removal, run-loop wiring, snapshot testing, and strict CRAP/mutant gates.
Out of Scope Changes check ✅ Passed All changes directly support the renderer implementation and documentation updates. CI workflow updates ensure artifact stability; config/manifest/doc changes document or enforce the new feature appropriately.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch adapters-11-renderer

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

PR 8b dropped `edge_label` (and `smoke_report_html`) from
`src/adapters/asset_embed.rs`; the `edge-vocab-completeness` CI gate's
grep against that file always failed.

- Add `edge_type_wire_key` to `src/adapters/render.rs` — an exhaustive
  Rust match producing the snake_case wire key per `EdgeType` variant.
  Compile-time exhaustiveness is now the primary mechanism (a new
  variant fails to build); CI greps this function as belt-and-braces.
- Add a unit test pinning every variant's wire key against
  `serde_json`'s snake_case serialization. If a future
  `#[serde(rename_all)]` edit drifts the wire shape, the test fails
  before the JS palette ships a broken color.
- Update the CI gate to enforce BOTH coverage paths:
  1. Every variant appears in `render.rs::edge_type_wire_key`.
  2. Every snake_case wire key appears in `templates/report.html`'s
     `JOIN_COLORS` map (covers the JS palette + legend).

The two-side gate keeps the renderer's edge palette in lockstep with
the classifier vocabulary without re-introducing the Rust→Mermaid
label string the original `edge_label` function carried.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
@cmbays

cmbays commented May 23, 2026

Copy link
Copy Markdown
Contributor Author

/gemini review

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request replaces the placeholder smoke renderer with a full-featured askama 0.16 implementation, introducing src/adapters/render.rs and templates/report.html. The new renderer manages per-model payload assembly, CTE graph node-role classification, and the binding of unit test fixtures to import-CTEs. The CLI run loop has been updated to integrate this renderer, and documentation regarding synthetic-only fixtures was relocated to CONTRIBUTING.md. Review feedback highlighted logic issues in the SQL parsing heuristics, specifically noting that keyword detection in is_simple_from_select is fragile due to strict spacing requirements and that parse_ref_name lacks support for valid dbt whitespace variations.

Comment thread src/adapters/render.rs
Comment on lines +171 to +181
fn is_simple_from_select(sql: &str) -> bool {
let lower = sql.to_ascii_lowercase();
if !lower.contains("select") {
return false;
}
if lower.contains(" join ") || lower.contains("\njoin ") {
return false;
}
let from_count = lower.matches(" from ").count() + lower.matches("\nfrom ").count();
from_count == 1
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The is_simple_from_select heuristic contains a logic bug that will cause it to fail for most standard SQL strings. The use of matches(" from ") and contains(" join ") requires specific surrounding spaces that are often absent (e.g., at the end of a string or before a semicolon). For example, "select * from table" will result in a from_count of 0 because there is no space after table. Using split_whitespace() to iterate over words is a much more robust approach for keyword detection.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 31ddda6. Replaced the substring-with-required-spaces heuristic with a whitespace-tokenizing pass: sql.split_whitespace() → lowercase + strip trailing ,;) → classify as select/from/join. Trailing-punctuation and end-of-string edge cases (e.g. select * from x;, select 1 from x\n) now classify correctly. Verification: is_simple_from_select_* unit tests at src/adapters/render.rs:1130-1165.

Comment thread src/adapters/render.rs
Comment on lines +108 to +119
pub fn parse_ref_name(input: &str) -> Option<&str> {
let trimmed = input.trim();
let rest = trimmed.strip_prefix("ref(").or_else(|| {
trimmed
.strip_prefix("REF(")
.or_else(|| trimmed.strip_prefix("Ref("))
})?;
let rest = rest.strip_suffix(')')?;
let inner = rest.trim();
let name = inner.strip_prefix('\'')?.strip_suffix('\'')?;
if name.is_empty() { None } else { Some(name) }
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The current implementation of parse_ref_name is fragile. It only supports three specific case variants of the ref keyword and does not allow for whitespace between the keyword and the opening parenthesis (e.g., ref ('model')), which is valid in dbt. Additionally, using to_ascii_lowercase() on the prefix check would be more robust than hardcoding variants.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 31ddda6. parse_ref_name now (1) uses eq_ignore_ascii_case("ref") for case-insensitive keyword matching across any byte casing, and (2) calls trim_start() between the keyword and the opening paren so ref ('x'), REF\t('y'), etc. parse correctly. Verification: new tests parse_ref_name_tolerates_whitespace_between_ref_and_paren and the updated parse_ref_name_accepts_case_variant_keyword.

cmbays and others added 3 commits May 23, 2026 12:13
Two parallel changes in this commit:

## examples/

End-to-end demonstration of cute-dbt's output, generated against the
already-committed jaffle-shop fixture pair. A reader evaluating the
tool opens `examples/jaffle-shop-report.html` directly and sees the
real renderer behavior — Mermaid DAG, DataTables panels, JSON payload,
the in-scope `stg_customers` model with its 1 unit test carrying
populated `tags` / `meta` / `defined_in`.

- `examples/jaffle-shop-report.html` (3.6 MB) — generated by
  `cargo run --bin cute-dbt -- --manifest tests/fixtures/...current.json
  --baseline-manifest tests/fixtures/...baseline.json --out ...`.
- `examples/README.md` — what each file demonstrates + how to
  regenerate + why we commit the artifact instead of generating in CI.
- `.gitattributes` — marks `examples/*.html` as `-diff
  linguist-generated` so `git diff` and the GitHub UI suppress the
  unreadable wall of inlined-asset bytes; the new CI guard is the
  authoritative change detector.
- CI job `example-report-up-to-date` — regenerates the report against
  the same committed fixtures and asserts byte-identical to
  `examples/jaffle-shop-report.html`. Drift fails the gate with a
  one-line remediation command. (The two previous skip-stub jobs for
  the headless zero-egress proof stay skipped; their tracking issue
  is unchanged.)

The renderer is deterministic over fixture input (BTreeMap +
BTreeSet-backed iteration in payload assembly), so byte-equality is
the correct gate.

## Comment sweep

Removed pipeline-private references that are not self-contained for a
reader of the public repo:

- "PR 8b" / "PR 8a" / "PR N" labels in module docs and integration
  test docs → either dropped or replaced with the GitHub issue number
  where the public anchor is meaningful.
- "ADR-N" references → replaced with the matching `ARCHITECTURE.md §N`
  section (which IS public). Architectural intent is preserved.
- "ADR-4 amendment 2026-05-22" date-stamped breadcrumbs → dropped; the
  in-file comments describe the actual behavior they were pointing at
  (Mermaid `<g>` id selector form, `securityLevel: 'strict'` constraints,
  the JS `JOIN_COLORS` keying contract).

The insta chrome snapshot was regenerated to reflect the JS comment
updates (no behavior change).

## Tests

255 passed (was 254 — one new `edge_type_wire_key_*` test added
earlier this PR). Build / clippy / fmt clean.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
…elog

- examples/README.md → link the richer-examples follow-up issue (#39).
- CHANGELOG.md → drop "PR 8b" / "ADR-4 amendment" labels that don't
  read for someone arriving at the file from outside the build
  pipeline; add an Added bullet for the examples/ directory.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
…count, union split

Five small fixes from a live walk-through of the example report:

1. **Compiled-SQL block: top padding.** The floating `.sql-copy` button
   was overlapping the first line of SQL. Added `padding-top: 2.5rem`
   to `.sql-block` — the button (~30 px from the wrapper's top edge)
   now clears the first line cleanly.

2. **Import-CTE binding: pass-2 body match.** The existing jaffle-shop
   fixture's import CTE is named `source` (dbt's idiomatic unwrapper
   name), and the unit test's `given.input` is `ref('raw_customers')`.
   Pass-1 name match couldn't resolve `raw_customers` against `source`,
   so the rendered report showed the empty "no fixture provided" state.

   Extended `find_import_node_id` with a pass-2 body-leaf match: when
   pass-1 misses, scan each import-CTE's `raw_sql` for table refs (via
   the new `extract_table_leaf_refs` helper) and pick the CTE whose
   body references the ref'd model. The dbt-idiomatic compiled-SQL
   shape `with source as (select * from "db"."schema"."MODEL")` now
   binds cleanly without forcing fixtures to use the design's
   sample-data CTE-name=model-name convention.

3. **Drop redundant test-count UI.** The model selector option labels
   already carry each model's test count (`stg_customers (1 test)`).
   The right-side `.test-select-count` element next to the unit-test
   selector was duplicating the same signal; removed the span, the
   JS update, and the (now-orphaned) CSS rule.

4. **UNION variants get distinct colors.** Previously `union_all` and
   `union_distinct` both rendered dashed orange. Split: `union_all`
   stays orange (duplicates preserved), `union_distinct` is now blue
   (deduplicated). Both stay dashed so the row-concatenation cue is
   preserved.

5. **Regenerated `examples/jaffle-shop-report.html`** — the report now
   surfaces the `ref('raw_customers')` Given table when you click the
   `source` import node, instead of the empty-state copy.

## Tests

261 passed (was 255 — 6 new):
- `build_payload_given_binds_import_cte_via_body_table_reference`
- 5 unit tests for `extract_table_leaf_refs`

Pre-push battery: 9/9 green. CI guard locally byte-identical against
the regenerated example.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 4

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In @.github/workflows/ci.yml:
- Around line 465-489: The CI jobs resource-ref-lint and headless-zero-egress
are currently disabled by the literal gate `if: false`; remove or change that
gate so the jobs run (for example remove the `if: false` lines or replace with a
proper conditional like `if: github.event_name == 'pull_request' ||
github.event_name == 'push'`), ensuring the Resource-ref lint job
(`resource-ref-lint`) executes to reject real-loading constructs and the
headless browser proof job (`headless-zero-egress`) runs with network blocked to
assert zero outbound requests; keep their steps (checkout + run) intact and
adjust any timeouts or runner settings if needed.
- Around line 445-447: The workflow uses mutable tags for third‑party
actions—replace each tag-based reference (the uses lines "actions/checkout@v4",
"dtolnay/rust-toolchain@stable", and "Swatinem/rust-cache@v2") with the
corresponding commit SHA pinned references; locate the exact commit SHA for each
action repository (e.g. the commit that matches the tag you intend to lock) and
update the uses entries in the new job in .github/workflows/ci.yml to use the
full SHA (repo@sha) so the CI job is pinned to immutable revisions.

In `@src/adapters/render.rs`:
- Around line 1233-1258: The test function
render_report_does_not_emit_external_resource_constructs is using fragile
substring greps to detect external references; replace those raw
contains/replace checks with the project's structured resource-ref lint utility
(the same linter used elsewhere to reject <script src>, <link href>, <img src>,
CSS `@import` and url(), and protocol-relative //), invoking it against the
generated chrome content from render_report (and still stripping the known
inlined assets like SAKURA_CSS, DATATABLES_CSS, JQUERY_JS, DATATABLES_JS,
MERMAID_JS beforehand); assert the linter reports no external references rather
than relying on plain .contains("http://")/.contains(" src=\"") greps.

In `@templates/report.html`:
- Line 239: The inline JSON script with id "cute-dbt-data" currently uses {{
payload|json|safe }} which allows a payload containing "</script>" to break out
and cause XSS; change the rendering to a safe JSON-encoding mechanism (e.g., use
the template engine's safe JSON serializer such as Jinja2's tojson filter or
otherwise escape closing script tags) so the payload is serialized as valid JSON
and any "</script>" sequences are escaped (for example by replacing "</script>"
with "<\/script>") before insertion; update the template expression that renders
payload (the {{ payload... }} expression) accordingly and remove any raw |safe
usage so the JSON blob cannot terminate the script tag.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 7d41ee93-b9cb-4380-ac82-bd27e96543a9

📥 Commits

Reviewing files that changed from the base of the PR and between 116a721 and 8a7093a.

⛔ Files ignored due to path filters (2)
  • Cargo.lock is excluded by !**/*.lock
  • tests/snapshots/render_integration__rendered_chrome_jaffle_shop.snap is excluded by !**/*.snap
📒 Files selected for processing (20)
  • .cargo/mutants.toml
  • .gitattributes
  • .github/workflows/ci.yml
  • AGENTS.md
  • ARCHITECTURE.md
  • CHANGELOG.md
  • CONTRIBUTING.md
  • Cargo.toml
  • README.md
  • crap4rs.toml
  • examples/README.md
  • examples/jaffle-shop-report.html
  • features/report_generation.feature
  • src/adapters/asset_embed.rs
  • src/adapters/mod.rs
  • src/adapters/render.rs
  • src/cli/mod.rs
  • templates/report.html
  • tests/asset_embed.rs
  • tests/render_integration.rs
💤 Files with no reviewable changes (2)
  • tests/asset_embed.rs
  • features/report_generation.feature

Comment thread .github/workflows/ci.yml
Comment thread .github/workflows/ci.yml
Comment thread src/adapters/render.rs
Comment thread templates/report.html Outdated
cmbays and others added 2 commits May 23, 2026 12:39
cargo-mutants surfaced one surviving mutant after the body-match
extension: relaxing the pass-1 predicate from `name match && role
match` to `name match || role match` would spuriously bind a
wrong-named import CTE to a unit test whose ref doesn't actually match
it. No existing test reached that code path because every fixture
either had ALL import-named-correctly or NO import-named-correctly.

Add a discriminating fixture: an import CTE with the WRONG name + a
transform CTE with the RIGHT name + a body that references an
unrelated table. The correct return is `None` (neither pass-1 nor
pass-2 matches); under the mutant, pass-1's `||` would accept the
first import-role node regardless of name and return its identifier.

Mutation testing on `src/adapters/render.rs`: 175 mutants, 146 caught,
0 missed, 29 unviable.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Resolves the actionable findings from CodeRabbit + Gemini on PR #38.

## Rust-side `</` + `<!` escaping for the JSON payload

CodeRabbit flagged that `{{ payload|json|safe }}` relied on askama's
filter to escape script-terminating sequences. While the askama 0.16
book documents that `|json` escapes `<>&'`, encoding the safety
property in askama's filter is an indirection that breaks the moment
the template engine swaps or the docs drift.

Move the escape to Rust:
- New `payload_json_for_html_script(&ReportPayload) -> Result<String, _>`
  serializes via serde_json, then re-escapes any `<` followed by `/`
  or `!` as the JSON unicode escape `<`. These are the only two
  sequences HTML5's script-data state machine treats specially
  (`</alpha` terminates, `<!--` enters script-data-double-escape).
- Template now takes `payload_json: &str` (pre-escaped) and renders it
  with `|safe`. The safety property is the Rust-side escape, no longer
  the filter.
- Template no longer reads `payload.baseline`; the renderer passes
  `baseline_label: &str` directly. Cleaner separation: the template
  doesn't know the payload's structure.

Tests cover the escape (`</script>` and `<!--` both get unicode-
escaped, bare `<` stays bare, output round-trips through `JSON.parse`).

## Matcher robustness (Gemini)

- `parse_ref_name` is now byte-case-insensitive on the `ref` keyword
  (`rEf` accepted, not just the hardcoded `ref`/`REF`/`Ref` triple)
  and tolerates whitespace between the keyword and the opening
  paren (`ref ('x')`, `REF\t('y')`) — matches dbt/Jinja semantics.
- `is_simple_from_select` switches from substring-with-spaces
  (`" from "` / `"\nfrom "`) to a whitespace-tokenizing pass that
  classifies cleanly regardless of trailing punctuation or line
  endings. The token-stripping handles `from;` / `from,` / `from)`
  edge cases the original substring scan would miss.

## Widened egress test patterns (CodeRabbit)

`render_report_does_not_emit_external_resource_constructs` and its
integration-test sibling now match the structured-lint pattern set:
`<script src` / `<link href` / `<img` / ` src="` / `@import` / `url(`
/ `http://` / `https://` / protocol-relative `//`. Comment references
the canonical structured lint at `#12` so a
future reader knows where the real gate lives.

## Validation

- 267 tests pass (was 263 — 4 new for `payload_json_for_html_script`).
- `cargo mutants --file src/adapters/render.rs`: 182 mutants,
  153 caught, 0 missed, 29 unviable.
- Example report regenerated; CI guard locally byte-identical.
- clippy + fmt + docs + cargo-deny clean.

## Dispositions for non-fixed findings (posted as PR replies)

- CodeRabbit: pin GitHub Actions to commit SHAs — won't fix; the rest
  of `.github/workflows/ci.yml` uses tag refs, so pinning only the
  new job is inconsistent. A workflow-wide SHA-pinning sweep belongs
  in its own PR.
- CodeRabbit: enable `resource-ref-lint` / `headless-zero-egress` —
  deferred to `#12` (pre-existing stubs with
  `tracked:` comments; the real harness lands there).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
@cmbays

cmbays commented May 23, 2026

Copy link
Copy Markdown
Contributor Author

@coderabbitai review

/gemini review

@cmbays
cmbays merged commit 05adb18 into main May 23, 2026
20 checks passed
@cmbays
cmbays deleted the adapters-11-renderer branch May 23, 2026 17:04
@coderabbitai

coderabbitai Bot commented May 23, 2026

Copy link
Copy Markdown

Tip

For best results, initiate chat on the files or code changes.

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
cat src/adapters/render.rs

Length of output: 62919


🏁 Script executed:

#!/bin/bash
cat templates/report.html

Length of output: 31974


🏁 Script executed:

#!/bin/bash
cat src/cli/mod.rs
cat src/adapters/asset_embed.rs
cat tests/render_integration.rs
cat .github/workflows/ci.yml

Length of output: 42410


@cmbays — this is a substantial, well-engineered PR. The renderer architecture is clean, the two-pass import-CTE binding is a nice design, and the test coverage (unit + integration + insta snapshot + egress guard) is thorough. A few issues below, ranging from dead payload weight to a visual rendering bug.


🔴 Issues worth fixing before merge

1. highlightSql — keyword regex double-wraps keywords inside string literals

In templates/report.html, the highlightSql function applies the string-literal regex before the keyword regex:

s = s.replace(/('[^'\n]*')/g, '<span class="sql-string">$1</span>');
// ...
s = s.replace(SQL_KEYWORDS_RE, '<span class="sql-keyword">$1</span>');

A literal like 'from me' becomes <span class="sql-string">'from me'</span> after the first pass, then SQL_KEYWORDS_RE matches from inside that span and wraps it again, producing:

<span class="sql-string">'<span class="sql-keyword">from</span> me'</span>

This is malformed HTML (a <span> inside an already-rendered <span>) and will cause incorrect colour display whenever a SQL keyword appears in a string value. The fix is to skip the keyword regex for content already inside a span — easiest by applying keywords before strings, or by using a negative look-behind/ahead, or by walking the string in segments.


2. NodePayload::ref_name is serialized but never consumed by the JS

NodePayload carries a ref_name: Option<String> (serialized as "ref") for import nodes:

let ref_name = if role == NodeRole::Import {
    Some(node.name().to_owned())
} else {
    None
};

Grepping templates/report.html's {% raw %} block, the JS never reads node.ref — neither in buildMermaidSource, renderNodeDetail, nor anywhere else. The bound_to_node field on GivenPayload is the actual binding mechanism. This field inflates every model's payload for no effect. Either wire it into the node-detail panel (e.g., label the import node with its upstream ref) or remove it before the JSON contract stabilises.


🟡 Issues worth addressing soon

3. buildMermaidSource(m, t) — t is a declared-but-unused parameter

function buildMermaidSource(m, t) {   // `t` never referenced in the body
    var lines = ["graph LR"];
    // ...
}

The callsite passes only m:

var src = buildMermaidSource(m);

Either remove the parameter or use it (perhaps to highlight the currently-selected-test's bound nodes). As-is it suggests abandoned intent and will trip any future JSHint/ESLint pass.


4. mermaid.initialize is called on every renderDag() invocation

function renderDag() {
    // ...
    window.mermaid.initialize({
        startOnLoad: false,
        securityLevel: "strict",
        theme: "base",
        fontFamily: '...'
    });
    var id = "mermaid-" + Math.random().toString(36).slice(2, 9);
    window.mermaid.render(id, src).then(...);
}

mermaid.initialize is called every time a model or test is selected. Mermaid handles re-initialization without error, but it re-parses its config on each call. Move the initialize call into the $(function() { ... }) startup block — call once, render many.


5. unwrapRows silently drops non-array formats

function unwrapRows(rows) {
    if (!rows) return [];
    if (Array.isArray(rows)) return rows;
    return [];   // format: sql, format: csv, or any future shape → silent empty
}

dbt unit tests support format: sql (rows is a SQL SELECT string) and potentially other formats. The GivenPayload::format field is already carried in the JSON, but the JS never branches on it. A user whose fixture uses format: sql will see an empty table with no error. Either render a "format: {{ format }} is not yet visualised" notice when format is non-null and Array.isArray(rows) is false, or document this as a known v0.1 limit in the template.


6. _model_name in find_import_node_id is unused

fn find_import_node_id(graph: &CteGraph, ref_name: &str, _model_name: &str) -> Option<String> {

The _ prefix acknowledges the suppression. The parameter has no use in either pass. If pass-3 (matching against the model's own bare name) is planned, a // TODO comment would clarify intent; otherwise removing the parameter keeps the call sites cleaner.


🟢 Nits / low-priority

7. is_simple_from_select allocates a String per token

let token = raw
    .to_ascii_lowercase()
    .trim_end_matches([',', ';', ')'])
    .to_owned();          // ← String per token
match token.as_str() { ... }

to_ascii_lowercase() already returns a String. trim_end_matches on it returns a &str slice, then .to_owned() re-allocates. The match arm could operate on the &str slice directly (bind to a variable before the match), avoiding one allocation per token.

8. DAG nodes are not keyboard-accessible

g.style.cursor = "pointer";
g.addEventListener("click", function () { ... });

No role="button", tabindex="0", or keydown handler — DAG nodes are click-only. Reasonable for v0.1 but worth tracking as a follow-up.


✅ Things that look solid

  • The payload_json_for_html_script \u003c escaping is correct and well-tested for both </ and <! sequences.
  • The two-pass import-CTE binding (name-match → body-match) is well-designed and the build_payload_pass_1_requires_both_name_match_and_import_role test is a great regression pin.
  • The edge-vocab-completeness CI guard (exhaustive match + awk snake_case derivation + JS palette check) is a clever structural enforcement.
  • The example-report-up-to-date byte-equality gate is exactly the right approach for a committed example artifact.
  • Removing MERMAID_INIT and moving init into the template's inlined script eliminates a class of drift bugs.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request implements a production-ready HTML report renderer using the askama template engine, replacing the previous placeholder smoke renderer. It introduces per-model payload assembly, CTE node role classification, and logic to bind unit test inputs to specific CTE nodes. The UI is enhanced with an interactive Mermaid DAG, searchable DataTables for test data, and a colorblind-safe legend for edge types. Feedback focuses on improving the robustness of the renderer's internal logic, specifically recommending the use of the sqlparser dependency instead of fragile whitespace-based heuristics for SQL classification and table reference extraction. Additionally, the reviewer suggests using expect() for node index lookups to document invariants and fail fast in case of logic errors, rather than providing silent fallbacks.

Comment thread src/adapters/render.rs
Comment on lines +157 to +159
let Some(node) = graph.nodes().get(node_index) else {
return NodeRole::Transform;
};

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

This function silently falls back to NodeRole::Transform if the node index is out of bounds. Since this function is called with indices derived from the graph itself, an out-of-bounds index indicates a logic error or invariant violation. Using expect() would be more appropriate to document this invariant and fail fast if it is violated.

    let node = graph.nodes().get(node_index).expect("node index is valid");
References
  1. Prioritize explicit defensive guards that document code invariants over more concise idiomatic alternatives if the guard clarifies intent and ensures robustness against future changes in logic.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Won't fix. Defensive fallback documents "this is unreachable but we don't panic" intent. classify_node_role is called from payload assembly where an out-of-bounds index is logically unreachable, but expect() would take down the entire report for a single-node anomaly. The current behavior degrades gracefully (the node simply gets classified as transform and renders with the neutral palette) rather than crashing the process. Full disposition: see my summary comment on the thread.

Comment thread src/adapters/render.rs
Comment on lines +184 to +200
fn is_simple_from_select(sql: &str) -> bool {
let mut has_select = false;
let mut from_count = 0usize;
for raw in sql.split_whitespace() {
let token = raw
.to_ascii_lowercase()
.trim_end_matches([',', ';', ')'])
.to_owned();
match token.as_str() {
"select" => has_select = true,
"join" => return false,
"from" => from_count += 1,
_ => {}
}
}
has_select && from_count == 1
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The whitespace-tokenizing heuristic for identifying simple SELECT statements is fragile and will incorrectly classify SQL containing keywords like FROM or JOIN inside string literals or comments (e.g., SELECT 'from' AS col FROM table). Since sqlparser is already a project dependency and used in the cte_engine, it should be used here to ensure robust classification. This is important for correctly identifying 'Import' nodes and displaying their associated unit test inputs in the UI.

fn is_simple_from_select(sql: &str) -> bool {
    use sqlparser::dialect::GenericDialect;
    use sqlparser::parser::Parser;
    use sqlparser::ast::{Statement, SetExpr, TableFactor};

    let Ok(ast) = Parser::parse_sql(&GenericDialect {}, sql) else { return false; };
    if let [Statement::Query(q)] = ast.as_slice() {
        if let SetExpr::Select(s) = &*q.body {
            return s.from.len() == 1
                && s.from[0].joins.is_empty()
                && matches!(s.from[0].relation, TableFactor::Table { .. });
        }
    }
    false
}

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tracked to #40. Agreed — the whitespace heuristic does false-positive on select 'from' as col from x and on SELECT '/* from */' FROM x. Folded into the combined renderer-hygiene tracking issue alongside G-4 (same root cause: use sqlparser AST). The post-AST implementation will subsume CodeRabbit's CR-7 micro-perf nit (per-token String allocation) — no more tokenization at all. Full disposition: see my summary comment on the thread.

Comment thread src/adapters/render.rs
Comment on lines +539 to +541
let Some(node) = graph.nodes().get(index) else {
return String::new();
};

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Returning an empty string when a node index is not found can lead to broken IDs in the JSON payload, which may cause the Mermaid diagram or interactive JS logic to fail silently. If this state is considered unreachable due to invariants in CteGraph, using expect() is preferred to document the invariant and ensure that any violation is caught during development or testing.

    let node = graph.nodes().get(index).expect("edge references valid node index");
References
  1. Prioritize explicit defensive guards that document code invariants over more concise idiomatic alternatives if the guard clarifies intent and ensures robustness against future changes in logic.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Won't fix. Same philosophical disposition as G-1: an empty-string node id would cause one Mermaid edge to render with a stub endpoint (graceful), while expect() would crash the whole render (catastrophic). The CTE engine's acyclicity invariant (from < to) plus the bounds-check at construction time makes this branch logically unreachable; the silent fallback is intentional defense-in-depth. Full disposition: see my summary comment on the thread.

Comment thread src/adapters/render.rs
Comment on lines +670 to +697
fn extract_table_leaf_refs(sql: &str) -> Vec<String> {
let tokens: Vec<&str> = sql.split_whitespace().collect();
let mut out = Vec::new();
for (i, tok) in tokens.iter().enumerate() {
let lower = tok.to_ascii_lowercase();
if lower != "from" && lower != "join" {
continue;
}
let Some(next) = tokens.get(i + 1) else {
continue;
};
// Strip surrounding `(` and trailing `,`/`)`/`;`/`(`; take the
// last `.`-delimited segment; strip `"` quotes.
let cleaned = next
.trim_start_matches('(')
.trim_end_matches([',', ')', ';']);
let leaf = cleaned.rsplit('.').next().unwrap_or("");
let leaf = leaf.trim_matches('"');
if leaf.is_empty() {
continue;
}
if !leaf.chars().all(|c| c.is_alphanumeric() || c == '_') {
continue;
}
out.push(leaf.to_ascii_lowercase());
}
out
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Similar to is_simple_from_select, this function uses a whitespace-based heuristic that is susceptible to false positives and negatives when SQL comments, string literals, or quoted identifiers with spaces are present. This can lead to failures in binding unit test inputs to the correct CTE nodes, causing the report to show 'no fixture provided' even when fixtures are defined. Leveraging the existing sqlparser dependency to traverse the AST would provide a much more reliable implementation for extracting table references.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tracked to #40. Agreed — same robustness class as G-2 (whitespace-tokenizing heuristic misses table references inside string literals + quoted identifiers with spaces). Folded into the combined renderer-hygiene tracking issue. Likely needs the CTE engine to expose the per-CTE body's Query AST rather than the renderer re-parsing the string — captured in #40's acceptance criteria. Full disposition: see my summary comment on the thread.

@cmbays

cmbays commented May 23, 2026

Copy link
Copy Markdown
Contributor Author

@coderabbitai @gemini-code-assist — dispositions on the post-merge findings (per memory feedback_bot-review-disposition-comments; one comment here since CodeRabbit's findings are in a single walkthrough comment, Gemini's are threaded inline — replying to those individually below):

Tracked to #40 (combined renderer-hygiene issue, fold into next renderer PR):

  • CR-1 highlightSql double-wraps SQL keywords inside string literals — real visual bug, agreed.
  • CR-2 NodePayload::ref_name serialized but unused by the JS — real dead-payload weight, agreed.
  • CR-3 buildMermaidSource(m, t) — t parameter unused — agreed, trivial.
  • CR-4 mermaid.initialize called on every renderDag() — agreed, move to startup block.
  • CR-5 unwrapRows silently drops format: sql / format: csv rows — real silent-failure bug, agreed.
  • CR-6 _model_name parameter unused in find_import_node_id — agreed, no third pass planned.
  • G-2 is_simple_from_select should use sqlparser AST instead of whitespace heuristic — agreed, classification fragile against select 'from' as col from x etc. Folded under renderer hygiene: fold post-merge bot findings on PR #38 into the next render-layer touch #40 alongside G-4.
  • G-4 extract_table_leaf_refs should use AST traversal instead of whitespace heuristic — agreed, same robustness class.

Tracked separately to #41 (different scope class — accessibility, not renderer hygiene):

  • CR-8 DAG nodes not keyboard-accessible (no tabindex/role/keydown) — agreed, priority:later.

Won't fix:

  • CR-7 is_simple_from_select allocates String per token — micro-perf nit and moot after the G-2 AST refactor lands under renderer hygiene: fold post-merge bot findings on PR #38 into the next render-layer touch #40. The post-AST implementation won't have this allocation pattern at all.
  • G-1 classify_node_role out-of-bounds fallback to Transform should be expect() — keeping the silent fallback. The function is called from payload assembly where an out-of-bounds index is logically unreachable, but choosing not to panic in a renderer is a defensible posture: the side panels degrade gracefully (the node simply gets classified as transform and renders with the neutral palette) rather than the entire process crashing. expect() would document the invariant but at the cost of taking down the whole report for a single-node anomaly.
  • G-3 endpoint_id empty-string fallback should be expect() — same philosophical disposition as G-1. An empty-string node id would cause one Mermaid edge to render with a stub endpoint; expect() would crash the whole render. The current behavior is preferred.

Things you flagged that look solid — appreciated the explicit callouts on payload_json_for_html_script's < escaping (both </ and <! paths), the two-pass import-CTE binding regression pin, the edge-vocab-completeness CI structural guard, and the example-report-up-to-date byte-equality gate. Those design choices were the higher-effort parts of this PR and the bot-review pass is the strongest external validation of them.

@gemini-code-assist

Copy link
Copy Markdown

Thank you for the detailed review and the thoughtful breakdown of the findings. I have noted the dispositions and the tracking issues (#40, #41) for the follow-up work. I appreciate the validation of the design choices regarding the payload escaping and the CI structural guards.

@coderabbitai

coderabbitai Bot commented May 23, 2026

Copy link
Copy Markdown

Tip

For best results, initiate chat on the files or code changes.

@cmbays — clean triaging, thanks for the thorough write-up. Dispositions all make sense:

  • #40 bundle (CR-1–6, G-2, G-4): tracked and agreed.
  • #41 (CR-8): good to have it isolated by scope class.
  • Won't fix — G-1 / G-3: the graceful-degradation rationale is sound. A renderer that produces a subtly wrong palette for one anomalous node is far preferable to one that crashes the whole report. The expect() pattern documents invariants well in library/domain code, but opt-out-silent in adapter/presentation code is the right call here.
  • Won't fix — CR-7: moot once the AST refactor lands; no point optimising an allocation pattern that the G-2/G-4 work will delete entirely.

Glad the escaping, two-pass binding, and CI guard callouts were useful — those were the parts most worth an external validation pass.

🐇 ✨


✏️ Learnings added
Learnt from: cmbays
URL: https://github.com/breezy-bays-labs/cute-dbt/pull/38

Timestamp: 2026-05-23T17:17:29.544Z
Learning: In the cute-dbt project (breezy-bays-labs/cute-dbt), the deliberate rendering philosophy is to prefer silent graceful fallback over panicking in adapter/presentation code (e.g., `classify_node_role` out-of-bounds → `Transform`, `endpoint_id` empty-string fallback). Crashing the whole renderer for a single anomalous node is considered worse than a subtly degraded panel or edge. Use `expect()` for domain/library invariants but not for renderer payload assembly.

Learnt from: cmbays
URL: https://github.com/breezy-bays-labs/cute-dbt/pull/38

Timestamp: 2026-05-23T17:17:29.544Z
Learning: In cute-dbt, CR-7 (micro-alloc in `is_simple_from_select` per-token `String`) is intentionally deferred as won't-fix because it will be eliminated by the G-2/G-4 AST-based refactor of `is_simple_from_select` and `extract_table_leaf_refs` tracked in issue `#40`.

Note: Learnings are effective only in the context of similar code segments. To apply general coding guidelines, please configure review instructions. You can manage existing learnings in the app.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

adapters: askama renderer reproducing Claude Design HTML

1 participant