Skip to content

fix(core): recalibrate the architecture thresholds against measured Svelte code - #311

Merged
oekazuma merged 11 commits into
mainfrom
worktree-architecture-thresholds
Jul 26, 2026
Merged

oekazuma merged 11 commits into
mainfrom
worktree-architecture-thresholds

Conversation

@oekazuma

@oekazuma oekazuma commented Jul 26, 2026 •

Copy link
Copy Markdown
Owner

Summary

Both Architecture rules carried thresholds picked without measurement, and both sat so far out in the tail that on a typical Svelte project neither rule fired at all:

Rule Was Now
architecture/prop-count > 10 props > 6
architecture/component-size > 400 lines > 200

How the numbers were derived

svelte-vitals' own parser was run over 7 real Svelte 5 codebases (4 component libraries, 3 applications; 6,460 .svelte files, 2,239 components with a countable prop count), mirroring the benchmark-based threshold selection used by ReactSniffer (Ferreira & Valente, IST 2023) — take a percentile of the measured distribution as the threshold.

Pooling everything gives p90 = 9, but one project (windmill) contributes 56% of the sample and is itself the corpus's highest outlier. The median of the per-repository p90 is 6, and the windmill-excluded pool agrees at every percentile — that agreement across the whole curve is what makes the aggregation choice principled rather than result-shopped.

For line count the per-repo p90 median is 124 and p95 median is 179; 200 sits deliberately above both, because length is a weaker and more context-dependent smell than a wide prop surface (tables, forms, and generated markup are legitimately long).

Note this lands well below React's empirical 13 — consistent with Svelte passing content through snippets, state through bind:, and shared state through context, all of which are props in React.

The corpus, raw distributions, and the measurement script are recorded in the design doc, and each constant carries a doc comment naming the corpus, statistic, and date — so the next person who feels 6 is "too strict" has the evidence rather than a fresh argument.

Design doc: docs/superpowers/specs/2026-07-25-architecture-threshold-recalibration-design.md
Plan: docs/superpowers/plans/2026-07-25-architecture-threshold-recalibration.md

Impact on existing projects

More info findings. The Architecture score barely moves — component-scoped rules score per file and the per-file scores are averaged, so over half a project's components must be flagged to lose even one point. Nothing fails by default (failOn defaults to critical), but anyone running --fail-on info / failOn: 'info' — including the Vite plugin's build mode — will newly fail on components that passed before. The changeset says all of this.

Test plan

  • Boundary tests added that actually pin both constants: propCount 6 passes / 7 flags, loc 200 passes / 201 flags. The pre-existing cases (15/3, 500/50) gave the same verdict under both old and new values, so nothing pinned the thresholds before this PR
  • pnpm lint && pnpm typecheck && pnpm build && pnpm test && pnpm check:publish — all green (1,643 tests)
  • Rule doc pages updated en + ja for both rules, each explaining where its number comes from
  • Changeset added (minor × 4)
  • Built via subagent-driven-development: 3 tasks, each independently reviewed, plus a whole-branch review

Reviewer note

The whole-branch review caught a real error in my own release note: it claimed each info finding would drop the Architecture score by a point. Verified against packages/core/src/scoring/score.ts — the scorer averages per-file scores, so that is arithmetically impossible at any realistic flag rate. Corrected, along with the omission of the one genuine breaking case (--fail-on info).

Two things are deliberately out of scope and recorded in the design doc as follow-ups: the 37% of $props() components that countProps cannot count (...rest / non-destructured), and per-rule configurable thresholds — the measured per-repo p90 ranged from 3 to 10, so a component library and an application genuinely want different numbers.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Updated Architecture checks to flag components with more than 6 props or over 200 lines, based on recalibrated (empirically derived) thresholds.
  • Documentation

    • Refreshed English and Japanese rule pages with the new limits and their percentile-based rationale.
    • Added/updated architecture threshold recalibration guidance and plans.
  • Bug Fixes

    • Documented expected behavior changes for existing projects, including how additional findings may affect runs when configured to treat “info” as failures.
  • Tests

    • Added boundary tests to ensure exact threshold passes and one over threshold fails for both rules.

oekazuma and others added 7 commits July 26, 2026 00:00
Correct the changeset's release note, which claimed a lowered Architecture
score from more info findings; the scorer averages per-file scores across a
project, so the score barely moves unless over half a project's components
get flagged. Document the real breaking case instead: --fail-on info /
failOn: 'info' (including the Vite plugin's build mode) now fails on
components that passed before, since failOn defaults to 'critical'.

Also: reframe the design doc's Problem section around findings visibility
rather than a score effect, add sample-size and Tailwind-confounder caveats
to the design doc's measurement sections, make the prop-count doc pages'
percentile claim exact (scoped to countable components), and fix the plan's
git add command to include the ja/ rules directory.
The final review found the changeset overstated the score effect; the same
wrong sentence appeared twice more in the plan document (its Global
Constraints and its copy of the changeset body). Corrected both to match:
component-scoped rules score per file and the per-file scores are averaged,
so the Architecture score barely moves — the real breaking case is
`--fail-on info` / `failOn: 'info'`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Jul 26, 2026 •

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@oekazuma, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 28 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: dcb5e27b-6246-4215-be2c-6ea0976e14a8

📥 Commits

Reviewing files that changed from the base of the PR and between 1ec8b37 and d6727b9.

📒 Files selected for processing (5)
  • .changeset/architecture-threshold-recalibration.md
  • docs/src/content/docs/ja/rules/architecture/prop-count.md
  • docs/src/content/docs/rules/architecture/prop-count.md
  • docs/superpowers/plans/2026-07-25-architecture-threshold-recalibration.md
  • docs/superpowers/specs/2026-07-25-architecture-threshold-recalibration-design.md
📝 Walkthrough

Walkthrough

The PR recalibrates architecture/prop-count from 10 to 6 props and architecture/component-size from 400 to 200 lines. It updates rule comments, boundary tests, empirical rationale, English/Japanese documentation, and package release metadata.

Changes

Architecture threshold recalibration

Layer / File(s) Summary
Empirical threshold rationale
docs/superpowers/specs/..., docs/superpowers/plans/...
Documents benchmark-based percentile derivation, selected thresholds, rule invariants, test requirements, and reproducibility details.
Rule thresholds and boundary validation
packages/core/src/rules/architecture/*, packages/core/test/architecture-rules.test.ts
Lowers the constants to 6 props and 200 lines, and verifies exact-threshold passes plus one-over-threshold findings.
Documentation and release updates
docs/src/content/docs/**/rules/architecture/*, .changeset/architecture-threshold-recalibration.md
Synchronizes English and Japanese rule descriptions and records minor releases for the affected packages.

Estimated code review effort: 2 (Simple) | ~10 minutes

Possibly related PRs

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly summarizes the main change: recalibrating core architecture thresholds using measured Svelte code.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In
`@docs/superpowers/specs/2026-07-25-architecture-threshold-recalibration-design.md`:
- Around line 194-212: Add the text language identifier to the fenced raw-output
block containing the percentile and threshold results, changing its opening
fence to a text fence while preserving all output content unchanged.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: c85b074b-aa5c-4bba-8046-72b1da9a8561

📥 Commits

Reviewing files that changed from the base of the PR and between b1c6f80 and 7ae4595.

📒 Files selected for processing (10)
  • .changeset/architecture-threshold-recalibration.md
  • docs/src/content/docs/ja/rules/architecture/component-size.md
  • docs/src/content/docs/ja/rules/architecture/prop-count.md
  • docs/src/content/docs/rules/architecture/component-size.md
  • docs/src/content/docs/rules/architecture/prop-count.md
  • docs/superpowers/plans/2026-07-25-architecture-threshold-recalibration.md
  • docs/superpowers/specs/2026-07-25-architecture-threshold-recalibration-design.md
  • packages/core/src/rules/architecture/component-size.ts
  • packages/core/src/rules/architecture/prop-count.ts
  • packages/core/test/architecture-rules.test.ts

Comment thread docs/superpowers/specs/2026-07-25-architecture-threshold-recalibration-design.md Outdated
oekazuma and others added 2 commits July 26, 2026 12:32
The published rule pages quoted "2,239 components in 7 codebases", which
invites the reader to judge the sample rather than the result. Two changes:

- Widened the survey from 7 repositories to 13 (10 with enough runes
  components to yield a percentile), adding appwrite/console, SvelteKit's own
  repo, and xyflow. The per-repository p90 median for prop count stayed at
  exactly 6, and the windmill-excluded pool stayed at 6 too; component-size
  moved only 124 -> 132 (p90) and 179 -> 183 (p95), leaving 200 untouched.
  Neither constant changes.
- The rule pages now state the basis qualitatively and note that widening the
  survey did not move the number — which is the stronger claim anyway. The
  corpus, per-repository tables, and raw output stay in the design doc, which
  is internal and not published to the docs site.

Also refreshed: melt-ui and open-webui turned out to still be on Svelte 4
`export let`, so they contribute nothing — recorded, since it shows the rule
simply does not apply before a runes migration.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In @.changeset/architecture-threshold-recalibration.md:
- Line 10: Update the changeset sentence describing threshold derivation: state
that prop-count uses the median repository p90, while component-size uses a
conservative value above the measured p95. Remove the claim that both thresholds
are based on the median p90, while preserving the benchmark-survey context.

In `@docs/src/content/docs/rules/architecture/prop-count.md`:
- Line 12: Update the threshold explanation in prop-count.md to state that 6 is
the median of the per-repository 90th-percentile prop counts, rather than the
pooled 90th percentile across all components. Remove or revise the claim that 7
or more props exceeds roughly nine in ten surveyed components, and preserve the
note that widening the survey did not change the threshold.

In
`@docs/superpowers/specs/2026-07-25-architecture-threshold-recalibration-design.md`:
- Around line 154-157: Update the architecture threshold recalibration document
comment describing the pooled repository distribution: replace the stale 56%
outlier figure with approximately 49%, reflecting windmill-labs/windmill’s 1,265
of 2,591 countable components. Preserve the surrounding component count and
distribution context.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 876d2262-dd70-4ac8-8490-52a313b7ad15

📥 Commits

Reviewing files that changed from the base of the PR and between cf1ce24 and 1ec8b37.

📒 Files selected for processing (8)
  • .changeset/architecture-threshold-recalibration.md
  • docs/src/content/docs/ja/rules/architecture/component-size.md
  • docs/src/content/docs/ja/rules/architecture/prop-count.md
  • docs/src/content/docs/rules/architecture/component-size.md
  • docs/src/content/docs/rules/architecture/prop-count.md
  • docs/superpowers/specs/2026-07-25-architecture-threshold-recalibration-design.md
  • packages/core/src/rules/architecture/component-size.ts
  • packages/core/src/rules/architecture/prop-count.ts
🚧 Files skipped from review as they are similar to previous changes (5)
  • packages/core/src/rules/architecture/prop-count.ts
  • docs/src/content/docs/ja/rules/architecture/component-size.md
  • docs/src/content/docs/rules/architecture/component-size.md
  • docs/src/content/docs/ja/rules/architecture/prop-count.md
  • packages/core/src/rules/architecture/component-size.ts

Comment thread .changeset/architecture-threshold-recalibration.md Outdated
Comment thread docs/src/content/docs/rules/architecture/prop-count.md Outdated
Comment thread docs/superpowers/specs/2026-07-25-architecture-threshold-recalibration-design.md Outdated
oekazuma and others added 2 commits July 26, 2026 12:52
Both review findings were correct:

- The changeset described both numbers as "the median of each repository's
  90th percentile", but only prop-count's 6 is that. component-size's 200 is
  deliberately above the measured p90 and p95. Split into two clauses.
- The prop-count page said 6 is "the 90th percentile … so a component with 7+
  props is wider than roughly nine in ten of the components". That holds per
  repository, not across the pooled survey — pooled, p90 is 9 and > 6 covers
  16.3%. Kept the nine-in-ten intuition (it is the useful part) but scoped it
  to "in a typical project" and named the statistic as the median
  per-repository 90th percentile. Same correction in the ja page.

Not taken: LanguageTool's "almost never is wordy" nit — it reads naturally and
the alternatives are worse.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The review flagged a stale "56% of the sample" in the design doc's example of
the MAX_PROPS comment. It was stale, and checking the rest turned up the wider
problem behind it: both illustrated comments had drifted from what the source
files actually carry, since only the source was updated when the corpus was
widened. The design doc's example exists to show the code, so both blocks are
now byte-identical to prop-count.ts and component-size.ts (verified
programmatically, not by eye).

The plan document quotes the pre-widening figures too. That one is left as-is —
it records what was executed at the time — with a note at the top pointing at
the design doc and source for the current numbers.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@oekazuma
oekazuma merged commit 0b2fd98 into main Jul 26, 2026
7 checks passed
@oekazuma
oekazuma deleted the worktree-architecture-thresholds branch July 26, 2026 04:01
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant