fix(core): recalibrate the architecture thresholds against measured Svelte code - #311
Conversation
Correct the changeset's release note, which claimed a lowered Architecture score from more info findings; the scorer averages per-file scores across a project, so the score barely moves unless over half a project's components get flagged. Document the real breaking case instead: --fail-on info / failOn: 'info' (including the Vite plugin's build mode) now fails on components that passed before, since failOn defaults to 'critical'. Also: reframe the design doc's Problem section around findings visibility rather than a score effect, add sample-size and Tailwind-confounder caveats to the design doc's measurement sections, make the prop-count doc pages' percentile claim exact (scoped to countable components), and fix the plan's git add command to include the ja/ rules directory.
The final review found the changeset overstated the score effect; the same wrong sentence appeared twice more in the plan document (its Global Constraints and its copy of the changeset body). Corrected both to match: component-scoped rules score per file and the per-file scores are averaged, so the Architecture score barely moves — the real breaking case is `--fail-on info` / `failOn: 'info'`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
Warning Review limit reached
Next review available in: 28 minutes Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available. How can I continue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews. How do review limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (5)
📝 WalkthroughWalkthroughThe PR recalibrates ChangesArchitecture threshold recalibration
Estimated code review effort: 2 (Simple) | ~10 minutes Possibly related PRs
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
Comment |
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In
`@docs/superpowers/specs/2026-07-25-architecture-threshold-recalibration-design.md`:
- Around line 194-212: Add the text language identifier to the fenced raw-output
block containing the percentile and threshold results, changing its opening
fence to a text fence while preserving all output content unchanged.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro Plus
Run ID: c85b074b-aa5c-4bba-8046-72b1da9a8561
📒 Files selected for processing (10)
.changeset/architecture-threshold-recalibration.mddocs/src/content/docs/ja/rules/architecture/component-size.mddocs/src/content/docs/ja/rules/architecture/prop-count.mddocs/src/content/docs/rules/architecture/component-size.mddocs/src/content/docs/rules/architecture/prop-count.mddocs/superpowers/plans/2026-07-25-architecture-threshold-recalibration.mddocs/superpowers/specs/2026-07-25-architecture-threshold-recalibration-design.mdpackages/core/src/rules/architecture/component-size.tspackages/core/src/rules/architecture/prop-count.tspackages/core/test/architecture-rules.test.ts
The published rule pages quoted "2,239 components in 7 codebases", which invites the reader to judge the sample rather than the result. Two changes: - Widened the survey from 7 repositories to 13 (10 with enough runes components to yield a percentile), adding appwrite/console, SvelteKit's own repo, and xyflow. The per-repository p90 median for prop count stayed at exactly 6, and the windmill-excluded pool stayed at 6 too; component-size moved only 124 -> 132 (p90) and 179 -> 183 (p95), leaving 200 untouched. Neither constant changes. - The rule pages now state the basis qualitatively and note that widening the survey did not move the number — which is the stronger claim anyway. The corpus, per-repository tables, and raw output stay in the design doc, which is internal and not published to the docs site. Also refreshed: melt-ui and open-webui turned out to still be on Svelte 4 `export let`, so they contribute nothing — recorded, since it shows the rule simply does not apply before a runes migration. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
There was a problem hiding this comment.
Actionable comments posted: 3
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In @.changeset/architecture-threshold-recalibration.md:
- Line 10: Update the changeset sentence describing threshold derivation: state
that prop-count uses the median repository p90, while component-size uses a
conservative value above the measured p95. Remove the claim that both thresholds
are based on the median p90, while preserving the benchmark-survey context.
In `@docs/src/content/docs/rules/architecture/prop-count.md`:
- Line 12: Update the threshold explanation in prop-count.md to state that 6 is
the median of the per-repository 90th-percentile prop counts, rather than the
pooled 90th percentile across all components. Remove or revise the claim that 7
or more props exceeds roughly nine in ten surveyed components, and preserve the
note that widening the survey did not change the threshold.
In
`@docs/superpowers/specs/2026-07-25-architecture-threshold-recalibration-design.md`:
- Around line 154-157: Update the architecture threshold recalibration document
comment describing the pooled repository distribution: replace the stale 56%
outlier figure with approximately 49%, reflecting windmill-labs/windmill’s 1,265
of 2,591 countable components. Preserve the surrounding component count and
distribution context.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro Plus
Run ID: 876d2262-dd70-4ac8-8490-52a313b7ad15
📒 Files selected for processing (8)
.changeset/architecture-threshold-recalibration.mddocs/src/content/docs/ja/rules/architecture/component-size.mddocs/src/content/docs/ja/rules/architecture/prop-count.mddocs/src/content/docs/rules/architecture/component-size.mddocs/src/content/docs/rules/architecture/prop-count.mddocs/superpowers/specs/2026-07-25-architecture-threshold-recalibration-design.mdpackages/core/src/rules/architecture/component-size.tspackages/core/src/rules/architecture/prop-count.ts
🚧 Files skipped from review as they are similar to previous changes (5)
- packages/core/src/rules/architecture/prop-count.ts
- docs/src/content/docs/ja/rules/architecture/component-size.md
- docs/src/content/docs/rules/architecture/component-size.md
- docs/src/content/docs/ja/rules/architecture/prop-count.md
- packages/core/src/rules/architecture/component-size.ts
Both review findings were correct: - The changeset described both numbers as "the median of each repository's 90th percentile", but only prop-count's 6 is that. component-size's 200 is deliberately above the measured p90 and p95. Split into two clauses. - The prop-count page said 6 is "the 90th percentile … so a component with 7+ props is wider than roughly nine in ten of the components". That holds per repository, not across the pooled survey — pooled, p90 is 9 and > 6 covers 16.3%. Kept the nine-in-ten intuition (it is the useful part) but scoped it to "in a typical project" and named the statistic as the median per-repository 90th percentile. Same correction in the ja page. Not taken: LanguageTool's "almost never is wordy" nit — it reads naturally and the alternatives are worse. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The review flagged a stale "56% of the sample" in the design doc's example of the MAX_PROPS comment. It was stale, and checking the rest turned up the wider problem behind it: both illustrated comments had drifted from what the source files actually carry, since only the source was updated when the corpus was widened. The design doc's example exists to show the code, so both blocks are now byte-identical to prop-count.ts and component-size.ts (verified programmatically, not by eye). The plan document quotes the pre-widening figures too. That one is left as-is — it records what was executed at the time — with a note at the top pointing at the design doc and source for the current numbers. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Summary
Both Architecture rules carried thresholds picked without measurement, and both sat so far out in the tail that on a typical Svelte project neither rule fired at all:
architecture/prop-countarchitecture/component-sizeHow the numbers were derived
svelte-vitals' own parser was run over 7 real Svelte 5 codebases (4 component libraries, 3 applications; 6,460
.sveltefiles, 2,239 components with a countable prop count), mirroring the benchmark-based threshold selection used by ReactSniffer (Ferreira & Valente, IST 2023) — take a percentile of the measured distribution as the threshold.Pooling everything gives p90 = 9, but one project (windmill) contributes 56% of the sample and is itself the corpus's highest outlier. The median of the per-repository p90 is 6, and the windmill-excluded pool agrees at every percentile — that agreement across the whole curve is what makes the aggregation choice principled rather than result-shopped.
For line count the per-repo p90 median is 124 and p95 median is 179; 200 sits deliberately above both, because length is a weaker and more context-dependent smell than a wide prop surface (tables, forms, and generated markup are legitimately long).
Note this lands well below React's empirical 13 — consistent with Svelte passing content through snippets, state through
bind:, and shared state through context, all of which are props in React.The corpus, raw distributions, and the measurement script are recorded in the design doc, and each constant carries a doc comment naming the corpus, statistic, and date — so the next person who feels 6 is "too strict" has the evidence rather than a fresh argument.
Design doc:
docs/superpowers/specs/2026-07-25-architecture-threshold-recalibration-design.mdPlan:
docs/superpowers/plans/2026-07-25-architecture-threshold-recalibration.mdImpact on existing projects
More
infofindings. The Architecture score barely moves — component-scoped rules score per file and the per-file scores are averaged, so over half a project's components must be flagged to lose even one point. Nothing fails by default (failOndefaults tocritical), but anyone running--fail-on info/failOn: 'info'— including the Vite plugin's build mode — will newly fail on components that passed before. The changeset says all of this.Test plan
propCount6 passes / 7 flags,loc200 passes / 201 flags. The pre-existing cases (15/3, 500/50) gave the same verdict under both old and new values, so nothing pinned the thresholds before this PRpnpm lint && pnpm typecheck && pnpm build && pnpm test && pnpm check:publish— all green (1,643 tests)Reviewer note
The whole-branch review caught a real error in my own release note: it claimed each
infofinding would drop the Architecture score by a point. Verified againstpackages/core/src/scoring/score.ts— the scorer averages per-file scores, so that is arithmetically impossible at any realistic flag rate. Corrected, along with the omission of the one genuine breaking case (--fail-on info).Two things are deliberately out of scope and recorded in the design doc as follow-ups: the 37% of
$props()components thatcountPropscannot count (...rest/ non-destructured), and per-rule configurable thresholds — the measured per-repo p90 ranged from 3 to 10, so a component library and an application genuinely want different numbers.🤖 Generated with Claude Code
Summary by CodeRabbit
New Features
Documentation
Bug Fixes
Tests