Skip to content

#259 — coverage truthfulness: severity/where/disabled/singular attribution honesty - #283

Merged
cmbays merged 5 commits into
mainfrom
domain-259-coverage-truthfulness
Jun 12, 2026
Merged

cmbays merged 5 commits into
mainfrom
domain-259-coverage-truthfulness

Conversation

@cmbays

@cmbays cmbays commented Jun 12, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Hardens verdict honesty in the coverage-intelligence checks over the #258 test-config wire (per-AC delivery below). Domain-only check logic in src/domain/checks.rs, a pure-rendering findings-panel surface, and a leak-checked fixture splice so every new cue is visible in the committed goldens.

Closes #259

Per-AC delivery

AC1 — severity/where/limit-degraded tests attribute as degraded backing. New DegradedBacking POD rides Finding.degraded (serde-skipped when empty — payload byte-stability): a covering uniqueness test that is severity: warn, where-filtered, or limit-capped still attributes in verdict.by, with every cause enumerated per test in domain-composed copy. Verdict vocabulary stays exactly three-valued (Covered/Uncovered/Unknown — never a fourth verdict, never a percentage); the cue is in-row beside the attribution per the #262 copy principles. An unrecognized severity surfaces its raw value, never guessed at. The findings panel shows a degraded backing summary chip only when every attributing test is weakened (a partially degraded attribution keeps the summary quiet — no false alarm) under the #146/#188 tooltip contract, and lists the per-test causes either way.

AC2 — disabled tests surface as "exists but disabled", distinct from absent. A disabled uniqueness test on the declared grain never counts as coverage, but surfaces as an exists but disabled evidence row naming the test id + columns. Both disabled surfaces are scanned: nodes-map tests carrying config.enabled: false (synthetic manifests) and the manifest disabled map (where both real engines put them), via a shared uniqueness_columns recognizer over TestMetadata + the column_name fallback — the same linkage DisabledEntry keeps on disabled generic tests.

AC3 — singular tests participate via depends_on linkage. With no enabled generic uniqueness backing, an enabled singular (SQL-file) test referencing the model through depends_on.nodes (the only wire linkage singular tests carry — #258, live-probed on both engines) degrades the verdict to honest UNKNOWN: evidence states what WAS checked (generic backing) and enumerates the singular tests, and no recommendation fires. Interpretation note (the issue's Discovery question): a singular test's SQL is not statically classifiable, so claiming it as Covered backing would overclaim on a TOTAL-tier check — the honest-UNKNOWN reading is implemented: singular tests participate by preventing the false Uncovered nag and by being named in evidence, never by entering by. Likewise the founder-ping question "warn = partial or annotated-full": implemented as annotated-full (attribution stays; the weakening is enumerated in-row) — a warn test still runs and still asserts, it just doesn't gate.

AC4 — existing check pins extended, none weakened. The #258 ingestion pin grows the fourth disabled entry's linkage assertions; all prior grain/union/join/incremental pins and both insta payload snapshots pass unchanged (the new degraded key is serde-skipped on full-strength findings). Spec conditions/exclusions prose mirrors the new predicate; heuristics/registry.toml + the book check page are regenerated from SPECS (byte-gated).

Dogfood-alongside (fixture splice + goldens)

playground-current.json (current side only; node-level audit confirms every other byte identical): unique_dim_payers_payer_key → severity warn; unique_fct_provider_metrics_provider_metrics_key → where + limit 100 (each with the authored unrendered_config twin); int_patients__never_admitted → unique_key: patient_id + a hand-authored synthetic singular test (assert_never_admitted_patients_distinct, cloned field-for-field from the engine-emitted singular specimen, depends_on-linked only); disabled map → a disabled unique test on mart_dq_summary.entity_type cloned from the engine-emitted disabled-generic specimen. MANIFEST.toml sha256 + provenance updated.

Golden-diff audit (payload-level, structural): intended-only deltas —

  • playground-report.html: dim_payers grain → covered + degraded(severity warn); int_patients__never_admitted → NEW grain finding unknown + generic-backing/singular evidence; mart_dq_summary grain → uncovered + exists-but-disabled evidence
  • diff-showcase-report.html: fct_provider_metrics grain → covered + degraded(where, limit); mart_dq_summary as above
  • explore/tests.html: same payload facts + check_specs prose; explore/dag.html byte-identical
  • jaffle-shop-report.html: CSS/JS bundle only (no payload change — no degraded/disabled/singular shapes there)

Tests

  • 13 new domain unit tests (causes enumeration incl. unrecognized severity, degraded ⊆ by, serde key-omission, both disabled surfaces + irrelevant-entry negatives, singular unknown/never-attributes/disabled-singular/sorted)
  • 4 new check_engine real-fixture pins + extended adapters: ingest test-config semantics + disabled map + singular-test linkage #258 ingestion pin
  • 4 new BDD scenarios over the subprocess wire (incl. the disabled-MAP injection arm in the builders)
  • 1 new headless test (chip presence + tooltip contract + causes + quiet-partial + exists-but-disabled text + unknown verdict)
  • 1 new render payload test (flattened degraded wire shape)

Gates (run directly, not via lefthook)

cargo fmt clean · clippy --all-targets --locked -D warnings clean · nextest 1423 passed · BDD 172 scenarios / 1104 steps passed · headless pair (headless_zero_egress 10, headless_toggle 77) passed · cargo doc --no-deps --locked -D warnings (incl. --document-private-items) clean · cargo deny check ok · heuristics ledger + example byte-gates green via regen.

🤖 Generated with Claude Code


Open in Stage

github-actions Bot and others added 3 commits June 12, 2026 01:39
…honest-UNKNOWN in the grain check

The grain.unique-key-unbacked verdict stops overclaiming on the #258
test-config wire (the cute-dbt#259 truthfulness pack):

- DegradedBacking POD + Finding.degraded (serde-skipped when empty —
  payload byte-stability): a covering test that is warn-severity,
  where-filtered, or limit-capped still attributes, with every cause
  enumerated per test in domain-composed copy. Never a fourth verdict,
  never a percentage — the three-valued covered/uncovered/unknown
  vocabulary stays the trust contract; the cue rides in-row beside the
  attribution (#262 copy principles). An unrecognized severity surfaces
  its raw value, never guessed at.
- exists-but-disabled evidence: a disabled uniqueness test on the
  declared grain (config.enabled: false in nodes, or a generic-test
  entry in the Manifest.disabled map — the shared uniqueness_columns
  recognizer now serves both linkage shapes) never counts as coverage
  but surfaces as a distinct fact from absent.
- singular-test linkage: with no enabled generic uniqueness backing, an
  enabled singular (SQL-file) test referencing the model via depends_on
  (the only wire linkage singular tests carry, #258) degrades the
  verdict to honest UNKNOWN — evidence states what WAS checked and
  enumerates the singular tests — never a false Uncovered nag on
  singular-test shops. Disabled-map singular entries carry no linkage
  (both engines empty depends_on on disabled nodes) and are declared
  out in the spec exclusions.

Spec conditions/exclusions mirror the new predicate; registry.toml +
the book check page are regenerated from SPECS (byte-gated).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… the findings panel

FindingPayload flattens the domain Finding, so the new degraded array
rides the wire unchanged (pinned by a build_payload test). The findings
panel renders it pure-presentationally: a summary 'degraded backing'
chip ONLY when every attributing test is weakened (no full-strength
backing at all — a partially degraded attribution keeps the summary
quiet, no false alarm), under the #146/#188 tooltip contract (focusable
trigger, aria-label, CSS-positioned bubble on hover AND focus, never a
native title); plus the per-test enumerated causes beside the Covered-by
attribution either way (in-row honesty). The exists-but-disabled and
singular-test cues render through the existing generic evidence list —
no new affordance needed.

Chip styling pairs always-AA --text with the non-text amber edge token,
so no per-theme contrast stand-in is needed; the latte .f-label deepening
extends to the new label. Chrome snapshot re-accepted (CSS/JS bundle
only).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…pins + golden regen

playground-current.json (current side only, leak-checked node-level
splice — every other byte verified identical): unique_dim_payers_payer_key gains severity warn,
unique_fct_provider_metrics_provider_metrics_key gains where
('year_actual >= 2024') + limit 100 (each with the matching authored
unrendered_config twin), int_patients__never_admitted gains
unique_key = 'patient_id' plus a hand-authored synthetic SINGULAR test
(assert_never_admitted_patients_distinct, cloned field-for-field from
the engine-emitted assert_patient_dates_valid shape, depends_on-linked
only), and the disabled map gains a fourth entry — a disabled unique
test on mart_dq_summary.entity_type cloned from the engine-emitted
disabled-generic specimen. MANIFEST.toml sha256 + provenance updated.

Pins: four new check_engine real-fixture tests (warn-degraded covered,
where+limit causes, exists-but-disabled on the UNCOVERED row,
singular-only honest UNKNOWN); the #258 ingestion pin extended to the
fourth disabled entry (none weakened); four new BDD scenarios over the
subprocess wire (degraded marks ⊆ by, disabled-MAP placement, singular
unknown) with the builders growing a disabled-map injection arm; one
new headless test driving the chip/tooltip/causes/quiet-partial DOM.

Goldens regenerated per the ci.yml recipes; payload-level diff audit
confirms intended-only deltas: dim_payers covered+degraded(warn) and
int_patients__never_admitted unknown(singular) and mart_dq_summary
uncovered+exists-but-disabled in playground-report,
fct_provider_metrics covered+degraded(where,limit) + mart_dq_summary in
diff-showcase, check_specs prose everywhere, explore tests.html payload
only; jaffle diff is the CSS/JS bundle alone.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Jun 12, 2026 •

Copy link
Copy Markdown

Warning

Review limit reached

@cmbays, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 46 minutes and 15 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more credits in the billing tab to continue.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 3cf6ecf6-b5c9-4317-9c55-a5e84b1f537c

📥 Commits

Reviewing files that changed from the base of the PR and between 97e8d38 and 523e5f5.

⛔ Files ignored due to path filters (1)
  • tests/snapshots/render_integration__rendered_chrome_jaffle_shop.snap is excluded by !**/*.snap
📒 Files selected for processing (20)
  • book/src/checks/grain.unique-key-unbacked.md
  • examples/diff-showcase-report.html
  • examples/explore/tests.html
  • examples/jaffle-shop-report.html
  • examples/playground-report.html
  • features/coverage_checks.feature
  • heuristics/registry.toml
  • src/adapters/render.rs
  • src/domain/checks.rs
  • src/domain/mod.rs
  • templates/interaction.js
  • templates/report.css
  • tests/check_engine.rs
  • tests/fixtures/MANIFEST.toml
  • tests/fixtures/playground-current.json
  • tests/headless_toggle.rs
  • tests/manifest_ingestion.rs
  • tests/steps/builders.rs
  • tests/steps/coverage_checks.rs
  • tests/steps/world.rs
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch domain-259-coverage-truthfulness

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions

github-actions Bot commented Jun 12, 2026 •

Copy link
Copy Markdown
Contributor

📄 Rendered report preview

All golden examples regenerated cleanly.

🟡 Golden examples

Committed to examples/ and byte-identity gated — the canonical reports contributors and consumers browse. Stable across PRs.

Report View Download
playground-report.html ▶ Open ↗ ⬇ Download
jaffle-shop-report.html ▶ Open ↗ ⬇ Download
diff-showcase-report.html ▶ Open ↗ ⬇ Download

🐶 Live dogfood preview

This PR doesn't touch dbt-project/, so there's no live dogfood preview.

▶ Open ↗ opens the report in your browser in one click —
published to this repo's GitHub Pages under /pr-283/.
⬇ Download fetches the same self-contained HTML as a workflow
artifact (auth-gated; works fully offline). Either way the report
makes zero external resource requests.

The Pages preview may take ~1 min to update after this comment
posts. On PRs from forks the Open link is unavailable (read-only
token) — use Download.

Alternative: GitHub CLI
# gh CLI >= 2.63 extracts into ./report-preview-playground/.
gh run download 27397421641 -R breezy-bays-labs/cute-dbt -n report-preview-playground
open report-preview-playground/playground-report.html

Posted by report-preview.yml for 523e5f52ec42cb09bd95e6e37d20efa96a02d6dc. Affordance only — never blocks merge.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request implements coverage truthfulness features for the unique-key unbacked check in cute-dbt (issue #259). It introduces support for identifying and rendering degraded backing (due to warn severity, where filters, or limit caps), surfacing disabled uniqueness tests as distinct evidence, and degrading verdicts to unknown when singular tests are present. The changes span the domain logic, rendering adapters, UI templates, and extensive test suites. The review feedback focuses on performance optimizations in src/domain/checks.rs, specifically recommending the use of borrowed &str slices instead of owned String allocations during filtering, sorting, and helper function returns to reduce unnecessary memory allocations.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment thread src/domain/checks.rs Outdated
Comment thread src/domain/checks.rs Outdated
Review follow-up (Gemini, PR #283): the grain detector's covering scan
collects (&str, &Node) and sorts on the borrowed key — ids become owned
Strings only at the POD ownership boundary (verdict.by /
DegradedBacking.by), dropping the intermediate owned collection;
singular_tests_on returns Vec<&str> borrowed from the manifest — its
sole consumer (grain_fallback_verdict) only formats the ids into
evidence copy, so owning them allocated just to discard. Pure refactor:
no payload change, goldens byte-identical.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@cmbays
cmbays merged commit b8cffa9 into main Jun 12, 2026
33 checks passed
@cmbays
cmbays deleted the domain-259-coverage-truthfulness branch June 12, 2026 05:56
github-actions Bot added a commit that referenced this pull request Jun 12, 2026
@cmbays

cmbays commented Jun 12, 2026

Copy link
Copy Markdown
Contributor Author

Overnight orchestration wrap-up

Merged as b8cffa9 (squash) under Christopher's overnight authority — CI fully green, 2 Gemini threads applied with honest dispositions (the covering-scan borrow refactor applied as idiomatic cleanup with the same-allocation-count caveat stated plainly; the singular_tests_on borrowed return was a genuine allocate-to-discard win). Issue #259 auto-closed ✔.

What shipped: the coverage truthfulness pack — DegradedBacking with per-test cause enumeration (warn severity / where filter / limit cap) under the honest-UNKNOWN posture (summary chip only on full degradation; in-row causes beside "Covered by"), exists-but-disabled evidence scanned on both surfaces (nodes map + the disabled map), and the singular-test honest-UNKNOWN flip (singular tests never enter by — TOTAL-tier soundness; evidence states what WAS checked). Dogfood visible across playground + diff-showcase goldens (degraded-warn, degraded-where/limit, exists-but-disabled, singular-UNKNOWN all rendered in committed examples).

🤖 Generated with Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

domain: coverage truthfulness pack — severity/where/disabled/singular-test attribution honesty

1 participant