Skip to content

feat(core): report each route's score per category - #366

Merged
oekazuma merged 12 commits into
mainfrom
feat/route-category-scores
Aug 5, 2026
Merged

oekazuma merged 12 commits into
mainfrom
feat/route-category-scores

Conversation

@oekazuma

@oekazuma oekazuma commented Aug 4, 2026 •

Copy link
Copy Markdown
Owner

Why

The report says what each route scored and what each category scored. It does not say what a route scored in a category, so a category that looks wrong cannot be traced to the routes that produced it.

That is not hypothetical. A field report observed a category displaying 100 while carrying 276 findings, and could neither confirm nor reject its own hypothesis from the output, because per-route scores were aggregated before they were exposed. It arrived as a question rather than a bug report for that reason.

Note precisely what was and was not missing. Every issue already carries its category, so identifying the routes with findings in a category is mechanical from the report as it stands. What is unrecoverable is magnitude: a key's deficit is (100 × failedWeight) / inventoryWeight, and the inventory denominators — 110 for seo::route, 28 for performance::route, 5 for seo::component — appear nowhere in the output. Two keys each carrying exactly one warning can be 95 points apart, and nothing in the report lets a reader see why.

What

{
  "route": "/blog",
  "score": 96,
  "categories": { "seo": 95, "performance": 100 },
  "issues": [ … ]
}

scoresByCategory gains an optional third ScoreOptions parameter forwarded to computeScore; the reporter calls it per route with { applyCriticalCap: false }. Every existing caller passes nothing and is unaffected — which matters, because computeHealth depends on a category capped at 79 by a critical continuing to pull Health down.

Three decisions, each of which could reasonably have gone the other way:

  • Scores only, no scoreModel. A route's results contain no project-scoped findings, so sitePenalty is structurally 0, and routeAverage restates score because there is one key.
  • Only the categories that produced a result on that route. A route with no architecture result must not appear with architecture: 100 — that would claim a measurement that never happened. "Fill in the missing categories as 100" is the obvious-looking improvement and it is wrong.
  • The critical cap stays off, matching routes[].score. A cap holding a whole category at 79 is a site-level signal; applying it per route would make one route's critical look like every route's problem.

routes[].score is not guaranteed to be the mean of routes[].categories

A key agrees exactly when every category on it scores the same ratio. Two very different-looking keys are both that case: a single-category key, where the bucket is the whole result set, and a clean multi-category key, where every ratio is zero. Clean keys are most keys on a healthy project.

When the ratios differ the mean usually disagrees, systematically once the gap exceeds a point:

value denominator result
routes[].score the union of observed pairs, 110 + 28 96
categories.seo 110 95
categories.performance 28 100

The mean of 95 and 100 is 97.5, not 96. score is one ratio against everything the route was measured against; each category score is a ratio against that category's own inventory.

Three exceptions to the disagreement are real, and two are ordinary — a sub-point gap collapses under flooring; equal observed inventories force agreement whatever the ratios, and for two categories those are provably the only exact coincidences besides equal ratios; and sporadic coincidences exist even under the default registry. The user docs therefore say not guaranteed, in both directions.

How the shape of the corpus fooled three drafts

This section of the design was rejected four times, and the shape of the error moved each time. Recorded here because the pattern is more useful than the conclusion:

  1. It stated the non-coincidence flatly and generalised it to "a mean of ratios with different denominators is not the ratio of the sums" — not a theorem; equal ratios coincide for any denominators.
  2. It then called the disagreement a rare exception, on the strength of the repo's own fixtures: 51 keys, 50 agreeing. But 46 of those 51 carry a single category, which is a property of minimal test fixtures rather than of the model. A 200-page project from the repo's own bench generator gives 413 keys, 400 of them multi-category, and 52% agreement with uniform six-point gaps.
  3. It then keyed the rule to category count — which contradicted its own table, because a clean multi-category key agrees. That rule predicts 3% agreement for a corpus measured at 52%.
  4. It carried a coincidence count and a worked example inherited from a reviewer's first search; the search had silently truncated, and the example was attached to a key shape that cannot occur.

The count is now omitted rather than corrected: a figure this design cannot reproduce by a stated method is the kind of figure it had already got wrong twice, and the existence claim is what the section needs.

Verification

core 1218, cli 805, vite 206 tests pass; tsc --noEmit clean in all three after rebuilding core; pnpm smoke 8/8; lint and format clean. The fixture's report grows 67,656 → 69,003 bytes, 61 per route over 22 routes — matching the design's measurement exactly.

Both mechanisms were confirmed by mutation rather than assumed: reverting the categories line fails all four new tests, and changing { applyCriticalCap: false } to {} in that one call fails exactly one test of 1,218, with expected 79 to be 86. The cap test asserts routes[].categories.seo specifically — asserting routes[].score would pass on a capped implementation, since that path is already cap-free.

Making categories required rather than optional forced eight literal route constructions in rendering fixtures to gain the field, four in core and four in vite, none of which a text search for the type name finds. That is the point: optional would have let a future literal carry undefined where {} is the honest value.

Out of scope, recorded in the design

Rendering the field in the HTML report or the dev dashboard — both receive it through app-shell.ts's route spread and ignore it, and whether it belongs in either UI is a question about those surfaces. Also: making the aggregates re-derivable, and a per-route scoreModel.

Design: docs/superpowers/specs/2026-08-04-route-category-scores-design.md. Plan: docs/superpowers/plans/2026-08-04-route-category-scores.md.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • JSON reports now include category-specific scores for each route.
    • Categories appear only when results are available and are calculated from category-specific measurements.
    • Category scores may differ from the overall route score and are not subject to the overall critical-score cap.
  • Documentation

    • Updated reporting guides with the new per-route categories data and scoring details.
  • Tests

    • Added coverage for category scoring, missing categories, critical-score handling, and report preservation.

oekazuma added 11 commits August 4, 2026 18:57
Five findings, all reproduced. The prescribed mechanism contradicted the cap decision:
scoresByCategory takes no options and defaults applyCriticalCap to true, so a route
with a critical came back 79 where routes[].score gives 86 — and the claim that
criticalCap is always null was false. The byte cost was a guess 35% low. And the
non-derivability claim, stated flatly, is false on 21 of the fixture's 22 routes.
The second draft called it a rare exception on the strength of the repo's fixtures,
where 46 of 51 keys carry one category. A 200-page project from the repo's own bench
generator gives 413 keys, 400 two-category, 52% agreement and systematic 6-point gaps.
Single-category agrees by construction, multi-category disagrees routinely; the fixture
is the atypical corpus. Also replace an invented 15-point spread with the measured 95.
The third draft's rule contradicted the table above it: 200 of the synthetic corpus's
400 multi-category keys agree, all of them clean, so a rule keyed on category count
predicts 3% agreement for a corpus measured at 52%. A key agrees when every category on
it scores the same ratio — single-category and clean multi-category are both that case,
and clean keys dominate a healthy project. Three real exceptions to the disagreement are
stated, including the k=2 theorem that equal inventories or equal ratios are the only
ones. Also drop an unmeasured field-scale prediction.
…example

The sporadic-coincidence exception carried three errors inherited from a reviewer's
first search: a count that search had silently truncated, an example attached to the
wrong key shape (it needs performance observed at 37, not its component-only 9), and
"bit-identical" for doubles that sit one ulp apart. The count is now omitted rather
than replaced, since it cannot be reproduced here. Also generalise the equal-inventory
exception to any k, scope its uniqueness proof to k=2, and add the flooring qualifier
the docs clause needed — exception 1 falsified it as written.
The caution presented a key observing performance at 37 as a case that can occur. It
cannot: a key is either a route id or a file path, so no key observes both of a
category's scopes, and performance contributes 28 or 9. 37 is precisely what a reader
gets by summing the registry — which is the trap the caution is about, so saying so
makes the example do its job instead of implying a state that does not exist.
@coderabbitai

coderabbitai Bot commented Aug 4, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 010aae67-29da-4cec-b09b-b72f604a85d3

📥 Commits

Reviewing files that changed from the base of the PR and between ff43663 and d673093.

📒 Files selected for processing (2)
  • docs/src/content/docs/ja/guides/(reporting)/reporters.md
  • packages/core/src/reporter/json.ts
🚧 Files skipped from review as they are similar to previous changes (2)
  • docs/src/content/docs/ja/guides/(reporting)/reporters.md
  • packages/core/src/reporter/json.ts

📝 Walkthrough

Walkthrough

The JSON reporter now adds per-route category scores. Category scoring accepts optional score options, including disabling the critical cap. Existing overall route scores remain unchanged. Tests, HTML snapshots, fixtures, documentation, and release metadata cover the new route field.

Changes

Route category score reporting

Layer / File(s) Summary
Category scoring and JSON serialization
packages/core/src/scoring/score.ts, packages/core/src/reporter/json.ts, docs/superpowers/specs/..., docs/superpowers/plans/...
scoresByCategory accepts optional ScoreOptions. JSON route entries now include category score maps computed from route results.
Report propagation and validation
packages/core/test/*report.test.ts, packages/core/test/score.test.ts, packages/vite/test/*
Tests cover category calculation, omission, uncapped critical scores, route-score independence, HTML snapshot propagation, and updated fixtures.
Documentation and release metadata
docs/src/content/docs/guides/(reporting)/*, docs/src/content/docs/ja/guides/(reporting)/*, .changeset/route-category-scores.md
The reporter documentation describes the categories field and its scoring rules. The changeset records minor releases for three packages.

Estimated code review effort: 3 (Moderate) | ~20 minutes

Sequence Diagram(s)

sequenceDiagram
  participant RouteResults
  participant scoresByCategory
  participant buildJsonReport
  participant formatHtmlReport
  RouteResults->>scoresByCategory: calculate category scores with options
  scoresByCategory->>buildJsonReport: return route categories
  buildJsonReport->>buildJsonReport: preserve the overall route score
  buildJsonReport->>formatHtmlReport: provide the serialized report
  formatHtmlReport->>formatHtmlReport: embed route categories in the snapshot
Loading

Possibly related PRs

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely describes the main change: reporting each route's score by category.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@packages/core/src/reporter/json.ts`:
- Around line 87-88: Update the explanatory comment in
packages/core/src/reporter/json.ts at lines 87-88 to state that the category
ratios are not guaranteed to average to score. Synchronize
docs/src/content/docs/ja/guides/(reporting)/reporters.md at line 85 by removing
のみ and documenting that equal ratios guarantee equality while single-category
routes, equal observed inventories, and other coincidence cases may also produce
equality.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 18e88150-3689-4737-b885-5729230e72c8

📥 Commits

Reviewing files that changed from the base of the PR and between 872cf85 and ff43663.

📒 Files selected for processing (12)
  • .changeset/route-category-scores.md
  • docs/src/content/docs/guides/(reporting)/reporters.md
  • docs/src/content/docs/ja/guides/(reporting)/reporters.md
  • docs/superpowers/plans/2026-08-04-route-category-scores.md
  • docs/superpowers/specs/2026-08-04-route-category-scores-design.md
  • packages/core/src/reporter/json.ts
  • packages/core/src/scoring/score.ts
  • packages/core/test/html-report.test.ts
  • packages/core/test/json-report.test.ts
  • packages/core/test/score.test.ts
  • packages/vite/test/app-shell-static.test.ts
  • packages/vite/test/ui-dashboard.test.ts

Comment thread packages/core/src/reporter/json.ts Outdated
…core

The design spent four review passes arriving at "not guaranteed", and both the code
comment and the Japanese guide then restated it as a certainty — the comment flatly,
and the Japanese with のみ, which makes equal ratios necessary as well as sufficient.
The English said "whenever ... and can differ otherwise" and was right, so the two
languages had also drifted apart.
@oekazuma
oekazuma merged commit 28d51e9 into main Aug 5, 2026
8 checks passed
@oekazuma
oekazuma deleted the feat/route-category-scores branch August 5, 2026 00:49
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant