feat(core): report each route's score per category - #366
Conversation
Five findings, all reproduced. The prescribed mechanism contradicted the cap decision: scoresByCategory takes no options and defaults applyCriticalCap to true, so a route with a critical came back 79 where routes[].score gives 86 — and the claim that criticalCap is always null was false. The byte cost was a guess 35% low. And the non-derivability claim, stated flatly, is false on 21 of the fixture's 22 routes.
The second draft called it a rare exception on the strength of the repo's fixtures, where 46 of 51 keys carry one category. A 200-page project from the repo's own bench generator gives 413 keys, 400 two-category, 52% agreement and systematic 6-point gaps. Single-category agrees by construction, multi-category disagrees routinely; the fixture is the atypical corpus. Also replace an invented 15-point spread with the measured 95.
The third draft's rule contradicted the table above it: 200 of the synthetic corpus's 400 multi-category keys agree, all of them clean, so a rule keyed on category count predicts 3% agreement for a corpus measured at 52%. A key agrees when every category on it scores the same ratio — single-category and clean multi-category are both that case, and clean keys dominate a healthy project. Three real exceptions to the disagreement are stated, including the k=2 theorem that equal inventories or equal ratios are the only ones. Also drop an unmeasured field-scale prediction.
…example The sporadic-coincidence exception carried three errors inherited from a reviewer's first search: a count that search had silently truncated, an example attached to the wrong key shape (it needs performance observed at 37, not its component-only 9), and "bit-identical" for doubles that sit one ulp apart. The count is now omitted rather than replaced, since it cannot be reproduced here. Also generalise the equal-inventory exception to any k, scope its uniqueness proof to k=2, and add the flooring qualifier the docs clause needed — exception 1 falsified it as written.
The caution presented a key observing performance at 37 as a case that can occur. It cannot: a key is either a route id or a file path, so no key observes both of a category's scopes, and performance contributes 28 or 9. 37 is precisely what a reader gets by summing the registry — which is the trap the caution is about, so saying so makes the example do its job instead of implying a state that does not exist.
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (2)
🚧 Files skipped from review as they are similar to previous changes (2)
📝 WalkthroughWalkthroughThe JSON reporter now adds per-route category scores. Category scoring accepts optional score options, including disabling the critical cap. Existing overall route scores remain unchanged. Tests, HTML snapshots, fixtures, documentation, and release metadata cover the new route field. ChangesRoute category score reporting
Estimated code review effort: 3 (Moderate) | ~20 minutes Sequence Diagram(s)sequenceDiagram
participant RouteResults
participant scoresByCategory
participant buildJsonReport
participant formatHtmlReport
RouteResults->>scoresByCategory: calculate category scores with options
scoresByCategory->>buildJsonReport: return route categories
buildJsonReport->>buildJsonReport: preserve the overall route score
buildJsonReport->>formatHtmlReport: provide the serialized report
formatHtmlReport->>formatHtmlReport: embed route categories in the snapshot
Possibly related PRs
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches📝 Generate docstrings
Comment |
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@packages/core/src/reporter/json.ts`:
- Around line 87-88: Update the explanatory comment in
packages/core/src/reporter/json.ts at lines 87-88 to state that the category
ratios are not guaranteed to average to score. Synchronize
docs/src/content/docs/ja/guides/(reporting)/reporters.md at line 85 by removing
のみ and documenting that equal ratios guarantee equality while single-category
routes, equal observed inventories, and other coincidence cases may also produce
equality.
🪄 Autofix
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro Plus
Run ID: 18e88150-3689-4737-b885-5729230e72c8
📒 Files selected for processing (12)
.changeset/route-category-scores.mddocs/src/content/docs/guides/(reporting)/reporters.mddocs/src/content/docs/ja/guides/(reporting)/reporters.mddocs/superpowers/plans/2026-08-04-route-category-scores.mddocs/superpowers/specs/2026-08-04-route-category-scores-design.mdpackages/core/src/reporter/json.tspackages/core/src/scoring/score.tspackages/core/test/html-report.test.tspackages/core/test/json-report.test.tspackages/core/test/score.test.tspackages/vite/test/app-shell-static.test.tspackages/vite/test/ui-dashboard.test.ts
…core The design spent four review passes arriving at "not guaranteed", and both the code comment and the Japanese guide then restated it as a certainty — the comment flatly, and the Japanese with のみ, which makes equal ratios necessary as well as sufficient. The English said "whenever ... and can differ otherwise" and was right, so the two languages had also drifted apart.
Why
The report says what each route scored and what each category scored. It does not say what a route scored in a category, so a category that looks wrong cannot be traced to the routes that produced it.
That is not hypothetical. A field report observed a category displaying 100 while carrying 276 findings, and could neither confirm nor reject its own hypothesis from the output, because per-route scores were aggregated before they were exposed. It arrived as a question rather than a bug report for that reason.
Note precisely what was and was not missing. Every issue already carries its
category, so identifying the routes with findings in a category is mechanical from the report as it stands. What is unrecoverable is magnitude: a key's deficit is(100 × failedWeight) / inventoryWeight, and the inventory denominators — 110 forseo::route, 28 forperformance::route, 5 forseo::component— appear nowhere in the output. Two keys each carrying exactly onewarningcan be 95 points apart, and nothing in the report lets a reader see why.What
{ "route": "/blog", "score": 96, "categories": { "seo": 95, "performance": 100 }, "issues": [ … ] }scoresByCategorygains an optional thirdScoreOptionsparameter forwarded tocomputeScore; the reporter calls it per route with{ applyCriticalCap: false }. Every existing caller passes nothing and is unaffected — which matters, becausecomputeHealthdepends on a category capped at 79 by acriticalcontinuing to pull Health down.Three decisions, each of which could reasonably have gone the other way:
scoreModel. A route's results contain no project-scoped findings, sositePenaltyis structurally 0, androuteAveragerestatesscorebecause there is one key.architectureresult must not appear witharchitecture: 100— that would claim a measurement that never happened. "Fill in the missing categories as 100" is the obvious-looking improvement and it is wrong.routes[].score. A cap holding a whole category at 79 is a site-level signal; applying it per route would make one route'scriticallook like every route's problem.routes[].scoreis not guaranteed to be the mean ofroutes[].categoriesA key agrees exactly when every category on it scores the same ratio. Two very different-looking keys are both that case: a single-category key, where the bucket is the whole result set, and a clean multi-category key, where every ratio is zero. Clean keys are most keys on a healthy project.
When the ratios differ the mean usually disagrees, systematically once the gap exceeds a point:
routes[].scorecategories.seocategories.performanceThe mean of 95 and 100 is 97.5, not 96.
scoreis one ratio against everything the route was measured against; each category score is a ratio against that category's own inventory.Three exceptions to the disagreement are real, and two are ordinary — a sub-point gap collapses under flooring; equal observed inventories force agreement whatever the ratios, and for two categories those are provably the only exact coincidences besides equal ratios; and sporadic coincidences exist even under the default registry. The user docs therefore say not guaranteed, in both directions.
How the shape of the corpus fooled three drafts
This section of the design was rejected four times, and the shape of the error moved each time. Recorded here because the pattern is more useful than the conclusion:
The count is now omitted rather than corrected: a figure this design cannot reproduce by a stated method is the kind of figure it had already got wrong twice, and the existence claim is what the section needs.
Verification
core1218,cli805,vite206 tests pass;tsc --noEmitclean in all three after rebuilding core;pnpm smoke8/8; lint and format clean. The fixture's report grows 67,656 → 69,003 bytes, 61 per route over 22 routes — matching the design's measurement exactly.Both mechanisms were confirmed by mutation rather than assumed: reverting the
categoriesline fails all four new tests, and changing{ applyCriticalCap: false }to{}in that one call fails exactly one test of 1,218, withexpected 79 to be 86. The cap test assertsroutes[].categories.seospecifically — assertingroutes[].scorewould pass on a capped implementation, since that path is already cap-free.Making
categoriesrequired rather than optional forced eight literal route constructions in rendering fixtures to gain the field, four incoreand four invite, none of which a text search for the type name finds. That is the point: optional would have let a future literal carryundefinedwhere{}is the honest value.Out of scope, recorded in the design
Rendering the field in the HTML report or the dev dashboard — both receive it through
app-shell.ts's route spread and ignore it, and whether it belongs in either UI is a question about those surfaces. Also: making the aggregates re-derivable, and a per-routescoreModel.Design:
docs/superpowers/specs/2026-08-04-route-category-scores-design.md. Plan:docs/superpowers/plans/2026-08-04-route-category-scores.md.🤖 Generated with Claude Code
Summary by CodeRabbit
New Features
Documentation
categoriesdata and scoring details.Tests