From eaba0ba3b8f9bd9b067eef6070e908c1d2280e1d Mon Sep 17 00:00:00 2001 From: oekazuma Date: Tue, 4 Aug 2026 18:57:31 +0900 Subject: [PATCH 01/12] docs: design per-route category scores --- ...2026-08-04-route-category-scores-design.md | 124 ++++++++++++++++++ 1 file changed, 124 insertions(+) create mode 100644 docs/superpowers/specs/2026-08-04-route-category-scores-design.md diff --git a/docs/superpowers/specs/2026-08-04-route-category-scores-design.md b/docs/superpowers/specs/2026-08-04-route-category-scores-design.md new file mode 100644 index 000000000..205513ca6 --- /dev/null +++ b/docs/superpowers/specs/2026-08-04-route-category-scores-design.md @@ -0,0 +1,124 @@ +# Per-route category scores in the JSON report — design + +**Date:** 2026-08-04 +**Status:** approved +**Origin:** the third follow-up recorded by `2026-07-31-score-honesty-design.md`, whose successor +`2026-08-04-score-proportionality-design.md` sharpened it: "a reader who wants to know why a category moved +now has to reconstruct a ratio per key rather than a subtraction." + +## The problem + +The report says what each route scored and what each category scored. It does not say **what a route scored in +a category**, so the one number a reader wants when a category looks wrong — which routes dragged it there — +has to be guessed from the issue list. + +That is not hypothetical. The field report that started this whole line of work observed a category displaying +100 while carrying 276 findings, and **could neither confirm nor reject its own hypothesis from the output**, +because per-route scores were aggregated before they were exposed. It arrived as a question rather than a bug +report for that reason. + +Two changes since have made the gap worse rather than better: + +- **`computeHealth` is deliberately not re-derivable** from the displayed category scores — it averages + unrounded values and floors once, so it can sit up to a point above their mean. +- **A key's score is now a ratio, not a subtraction.** Under the old model a reader could add up `DEDUCTION` + values and check the arithmetic by hand. Now they would have to know the severity-weighted inventory of + every `(category, scope)` pair the route touched. + +So the report's numbers are less checkable than they were, and this is the cheapest thing that makes them +checkable again. + +## Design + +Each entry in `routes[]` gains a `categories` map: + +```jsonc +{ + "route": "/blog", + "score": 96, + "categories": { "seo": 95, "performance": 100 }, + "issues": [ … ] +} +``` + +`JsonReport['routes']` becomes +`Array<{ route: string; score: number; categories: Record; issues: JsonIssue[] }>`, filled by +calling the already-exported `scoresByCategory` on that route's own results. + +Three decisions, each of which could reasonably have gone the other way: + +**Scores only, no `scoreModel`.** The top-level `categories` carries `{ score, scoreModel }`, and mirroring it +here would be symmetric and useless: a route's results contain no project-scoped findings, so `sitePenalty` is +always 0, and the critical cap is disabled on this path, so `criticalCap` is always null. `routeAverage` would +restate `score`. Three fields that cannot vary are noise, not symmetry. + +**Only the categories that produced a result on that route.** A route with no `architecture` result must not +appear with `architecture: 100`. That would claim a measurement that never happened — the same dishonesty +`2026-07-31-score-honesty-design.md` exists to remove, at a smaller scale. `scoresByCategory` already buckets +only what is present, so this is its behaviour rather than an added filter, and it is worth stating precisely +because "add the missing categories as 100" is the obvious-looking improvement someone will propose. + +**The critical cap stays off**, matching `routes[].score`. A cap that holds a whole category at 79 is a +site-level signal; applying it per route would make one route's `critical` look like every route's problem. + +`routes[].score` itself does not change. + +## The relationship that will surprise a reader, stated so it is not discovered as a bug + +**`routes[].score` is not the mean of `routes[].categories`.** Measured on one route carrying a failing `seo` +`warning` beside a passing `performance` route rule: + +| value | denominator | result | +| ------------------------ | ------------------------------------- | ------- | +| `routes[].score` | the union of observed pairs, 110 + 28 | **96** | +| `categories.seo` | 110 | **95** | +| `categories.performance` | 28 | **100** | + +The mean of 95 and 100 is 97.5, which is not 96. `routes[].score` is one ratio computed against everything the +route was measured against; each category score is a ratio computed against that category's own inventory. +Both are correct and they answer different questions. + +This is the third place in the scoring model where an aggregate is not re-derivable from the parts below it — +after `computeHealth` over category scores, and a category score over its key scores. The pattern is the same +each time and has the same cause: flooring and division happen at each level, and a mean of ratios with +different denominators is not the ratio of the sums. Recorded here rather than left for a reader to file. + +## Cost + +About 45 bytes per route — measured against the repo's own fixture, 22 routes, roughly 1 KB on a 67,656-byte +report. At the field project's 351 routes, roughly 16 KB. + +That cost is only safe to take because `872cf859` fixed the truncation that discarded everything past the +first 65,536 bytes of a piped report. Growing the report before that fix would have moved more of it past the +cliff. + +`packages/core/src/reporter/app-shell.ts` spreads each route (`{ ...route, issues: … }`), so the field reaches +the HTML report's embedded snapshot and the dev dashboard without further wiring. Nothing renders it yet; that +is a separate decision about those surfaces, not a gap this design leaves. + +## Testing + +1. **A route's category score matches scoring that route's results in that category alone.** Assert an exact + value derived from the formula, not from what the reporter printed. +2. **A category absent from a route's results is absent from its map** — not present as 100. This is the + dishonesty guard; a test asserting only "seo is present" would pass on an implementation that filled every + category in. +3. **`routes[].score` is unchanged** by this addition, on a fixture where it and the category scores disagree. + That disagreement is the point of the section above, so the test pins both numbers on the same input. +4. **The critical cap is off per route.** A route carrying a `critical` scores what the ratio gives, not 79. +5. **The field survives the HTML report's snapshot**, since `app-shell.ts` spreads the route object — one + assertion that the embedded snapshot carries `categories` for a route, so the spread is not silently + replaced by an explicit field list later. +6. **The documented shape matches.** `docs/src/content/docs/guides/(reporting)/reporters.md` and its Japanese + counterpart show the `routes[]` shape; the sample must gain the field, and the prose must say what it means + and that it does not average to `score`. + +## Deliberately not solved + +- **Rendering it.** The HTML report and the dev dashboard receive the field and ignore it. Whether a per-route + category breakdown belongs in either UI is a design question about those surfaces. +- **Making the aggregates re-derivable.** The three-level non-derivability described above is a consequence of + flooring at each level, and every alternative trades it for something worse — the predecessor spec worked + through this for `computeHealth` and chose deficit space precisely to keep the top of the scale honest. + Exposing the parts, which is what this change does, is the answer available. +- **A per-route `scoreModel`.** See above; every field would be constant. From e9b161e6b78e4c90f7449cfca5ad60f50545f8d9 Mon Sep 17 00:00:00 2001 From: oekazuma Date: Tue, 4 Aug 2026 19:17:49 +0900 Subject: [PATCH 02/12] docs: fix the route-category-scores design after adversarial review MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Five findings, all reproduced. The prescribed mechanism contradicted the cap decision: scoresByCategory takes no options and defaults applyCriticalCap to true, so a route with a critical came back 79 where routes[].score gives 86 — and the claim that criticalCap is always null was false. The byte cost was a guess 35% low. And the non-derivability claim, stated flatly, is false on 21 of the fixture's 22 routes. --- ...2026-08-04-route-category-scores-design.md | 95 ++++++++++++++----- 1 file changed, 70 insertions(+), 25 deletions(-) diff --git a/docs/superpowers/specs/2026-08-04-route-category-scores-design.md b/docs/superpowers/specs/2026-08-04-route-category-scores-design.md index 205513ca6..691170498 100644 --- a/docs/superpowers/specs/2026-08-04-route-category-scores-design.md +++ b/docs/superpowers/specs/2026-08-04-route-category-scores-design.md @@ -9,8 +9,14 @@ now has to reconstruct a ratio per key rather than a subtraction." ## The problem The report says what each route scored and what each category scored. It does not say **what a route scored in -a category**, so the one number a reader wants when a category looks wrong — which routes dragged it there — -has to be guessed from the issue list. +a category**, so the reader who wants to know why a category moved cannot get at the size of each route's +contribution. + +Note precisely what is and is not missing. Every issue already carries its `category`, so _identifying_ the +routes with findings in a category is mechanical from the report as it stands. What is unrecoverable is +**magnitude**: a route's deficit is `(100 × failedWeight) / inventoryWeight`, and the inventory denominators — +110 for `seo::route`, 28 for `performance::route`, and so on — appear nowhere in the output. Two routes with +one `warning` each can differ by 15 points and the report gives no way to see it. That is not hypothetical. The field report that started this whole line of work observed a category displaying 100 while carrying 276 findings, and **could neither confirm nor reject its own hypothesis from the output**, @@ -42,15 +48,30 @@ Each entry in `routes[]` gains a `categories` map: ``` `JsonReport['routes']` becomes -`Array<{ route: string; score: number; categories: Record; issues: JsonIssue[] }>`, filled by -calling the already-exported `scoresByCategory` on that route's own results. +`Array<{ route: string; score: number; categories: Record; issues: JsonIssue[] }>`. + +**The wiring needs one change, and a first draft of this section got it wrong.** `scoresByCategory` is exported +and buckets by category exactly as needed, but it takes no `ScoreOptions` and calls +`computeScore(rs, config)` — so `applyCriticalCap` defaults to **true**. Used as exported it would contradict +the cap decision below: measured, a route carrying one failing `seo` `critical` comes back as `seo: 79` with +`scoreModel.criticalCap: 79`, where the same route's `routes[].score` is **86**. + +So `scoresByCategory` gains an optional third parameter, `options: ScoreOptions = {}`, forwarded to +`computeScore`. Every existing caller — `computeHealth` among them — passes nothing and is unaffected, which +matters because the predecessor spec requires a capped category to keep pulling Health down. The reporter +passes `{ applyCriticalCap: false }`. + +The alternative, bucketing inside `buildJsonReport` and calling `computeScore` per bucket, was rejected: it +would duplicate `scoresByCategory`'s `r.category ?? 'seo'` fallback in a second place, where the two could +drift. Three decisions, each of which could reasonably have gone the other way: **Scores only, no `scoreModel`.** The top-level `categories` carries `{ score, scoreModel }`, and mirroring it -here would be symmetric and useless: a route's results contain no project-scoped findings, so `sitePenalty` is -always 0, and the critical cap is disabled on this path, so `criticalCap` is always null. `routeAverage` would -restate `score`. Three fields that cannot vary are noise, not symmetry. +here would add nothing: a route's results contain no project-scoped findings, so `sitePenalty` is structurally +0, and `routeAverage` restates `score` because there is one key. `criticalCap` would be null too — but only +because the cap is switched off above, which is a decision rather than a structural fact; the first draft of +this paragraph asserted it as one and was wrong. **Only the categories that produced a result on that route.** A route with no `architecture` result must not appear with `architecture: 100`. That would claim a measurement that never happened — the same dishonesty @@ -63,10 +84,13 @@ site-level signal; applying it per route would make one route's `critical` look `routes[].score` itself does not change. -## The relationship that will surprise a reader, stated so it is not discovered as a bug +## `routes[].score` is not guaranteed to be the mean of `routes[].categories` -**`routes[].score` is not the mean of `routes[].categories`.** Measured on one route carrying a failing `seo` -`warning` beside a passing `performance` route rule: +It usually is, which is exactly why the exception needs writing down. Measured with the field implemented, on +the repo's own fixture: **21 of 22 routes have `floor(mean(categories)) === score`**, and 17 of the 22 carry +only one category, where the two are equal by construction. One route disagrees. + +That one route is enough, because the disagreement is not a rounding artefact: | value | denominator | result | | ------------------------ | ------------------------------------- | ------- | @@ -74,27 +98,40 @@ site-level signal; applying it per route would make one route's `critical` look | `categories.seo` | 110 | **95** | | `categories.performance` | 28 | **100** | -The mean of 95 and 100 is 97.5, which is not 96. `routes[].score` is one ratio computed against everything the -route was measured against; each category score is a ratio computed against that category's own inventory. -Both are correct and they answer different questions. +The mean of 95 and 100 is 97.5, not 96. `routes[].score` is one ratio against everything the route was +measured against; each category score is a ratio against that category's own inventory. Both are correct and +they answer different questions. + +A first draft of this section stated the non-coincidence flatly and generalised it to "a mean of ratios with +different denominators is not the ratio of the sums", which is not a theorem — equal ratios coincide for any +denominators. The honest form is the one in this heading: not guaranteed, commonly true anyway, and the docs +must say it that way too, because a user comparing their own report against a flat "it does not average" would +find the report contradicting the documentation on most routes. -This is the third place in the scoring model where an aggregate is not re-derivable from the parts below it — -after `computeHealth` over category scores, and a category score over its key scores. The pattern is the same -each time and has the same cause: flooring and division happen at each level, and a mean of ratios with -different denominators is not the ratio of the sums. Recorded here rather than left for a reader to file. +This is the third level of the scoring model where an aggregate is not guaranteed to be re-derivable from the +parts below it — after `computeHealth` over category scores, and a category score over its key scores. Same +cause each time: division and flooring happen at each level. ## Cost -About 45 bytes per route — measured against the repo's own fixture, 22 routes, roughly 1 KB on a 67,656-byte -report. At the field project's 351 routes, roughly 16 KB. +**61 bytes per route**, measured with the field implemented rather than estimated: the fixture's report goes +from 67,656 to 69,003 bytes over 22 routes — 1,347 bytes. At the field project's 351 routes that is roughly +21 KB. The first draft guessed 45 bytes and was 35% low across all three figures. That cost is only safe to take because `872cf859` fixed the truncation that discarded everything past the first 65,536 bytes of a piped report. Growing the report before that fix would have moved more of it past the cliff. `packages/core/src/reporter/app-shell.ts` spreads each route (`{ ...route, issues: … }`), so the field reaches -the HTML report's embedded snapshot and the dev dashboard without further wiring. Nothing renders it yet; that -is a separate decision about those surfaces, not a gap this design leaves. +the HTML report's embedded snapshot and the dev dashboard without further wiring at runtime. Nothing renders +it yet; that is a separate decision about those surfaces, not a gap this design leaves. + +**At compile time it is not free.** A required `categories` on `JsonReport['routes']` breaks eight literal +route constructions that a text search for the type name cannot find — four in +`packages/core/test/html-report.test.ts`, and four across `packages/vite/test/app-shell-static.test.ts` and +`packages/vite/test/ui-dashboard.test.ts`. The CLI and the markdown reporter compile clean. So the +implementation must `tsc --noEmit` **each package separately** after rebuilding core, which is how the +equivalent break was found late on an earlier branch. ## Testing @@ -105,13 +142,21 @@ is a separate decision about those surfaces, not a gap this design leaves. category in. 3. **`routes[].score` is unchanged** by this addition, on a fixture where it and the category scores disagree. That disagreement is the point of the section above, so the test pins both numbers on the same input. -4. **The critical cap is off per route.** A route carrying a `critical` scores what the ratio gives, not 79. -5. **The field survives the HTML report's snapshot**, since `app-shell.ts` spreads the route object — one +4. **The critical cap is off per route, asserted on `routes[].categories[cat]` specifically.** A route carrying + one failing `seo` `critical` must show `categories.seo` at the ratio's value — 86 on the current registry — + not 79. Asserting `routes[].score` instead would pass on a broken implementation, because that path is + already cap-free; this test only holds anything if it names the category map. +5. **`scoresByCategory`'s existing callers are unchanged.** Called without options it must still cap, because + `computeHealth` depends on a capped category pulling Health down. Assert the capped value through + `scoresByCategory(rs, config)` and the uncapped one through + `scoresByCategory(rs, config, { applyCriticalCap: false })` on the same input. +6. **The field survives the HTML report's snapshot**, since `app-shell.ts` spreads the route object — one assertion that the embedded snapshot carries `categories` for a route, so the spread is not silently replaced by an explicit field list later. -6. **The documented shape matches.** `docs/src/content/docs/guides/(reporting)/reporters.md` and its Japanese +7. **The documented shape matches.** `docs/src/content/docs/guides/(reporting)/reporters.md` and its Japanese counterpart show the `routes[]` shape; the sample must gain the field, and the prose must say what it means - and that it does not average to `score`. + and that it is **not guaranteed** to average to `score` — not that it does not, which the reader's own + report would contradict on most routes. ## Deliberately not solved From 76f4401cb2b6f004dca25ba1f276217092f24aca Mon Sep 17 00:00:00 2001 From: oekazuma Date: Tue, 4 Aug 2026 19:25:11 +0900 Subject: [PATCH 03/12] docs: measure both corpora before calling the mean disagreement rare The second draft called it a rare exception on the strength of the repo's fixtures, where 46 of 51 keys carry one category. A 200-page project from the repo's own bench generator gives 413 keys, 400 two-category, 52% agreement and systematic 6-point gaps. Single-category agrees by construction, multi-category disagrees routinely; the fixture is the atypical corpus. Also replace an invented 15-point spread with the measured 95. --- ...2026-08-04-route-category-scores-design.md | 39 +++++++++++++------ 1 file changed, 28 insertions(+), 11 deletions(-) diff --git a/docs/superpowers/specs/2026-08-04-route-category-scores-design.md b/docs/superpowers/specs/2026-08-04-route-category-scores-design.md index 691170498..034edde4b 100644 --- a/docs/superpowers/specs/2026-08-04-route-category-scores-design.md +++ b/docs/superpowers/specs/2026-08-04-route-category-scores-design.md @@ -15,8 +15,10 @@ contribution. Note precisely what is and is not missing. Every issue already carries its `category`, so _identifying_ the routes with findings in a category is mechanical from the report as it stands. What is unrecoverable is **magnitude**: a route's deficit is `(100 × failedWeight) / inventoryWeight`, and the inventory denominators — -110 for `seo::route`, 28 for `performance::route`, and so on — appear nowhere in the output. Two routes with -one `warning` each can differ by 15 points and the report gives no way to see it. +110 for `seo::route`, 28 for `performance::route`, 5 for `seo::component` — appear nowhere in the output. So two +keys each carrying exactly one `warning` can be 95 points apart: against `seo::component`'s inventory of 5 that +one warning scores the key **0**, and against `seo::route`'s 110 it scores **95**. Nothing in the report lets a +reader see why. That is not hypothetical. The field report that started this whole line of work observed a category displaying 100 while carrying 276 findings, and **could neither confirm nor reject its own hypothesis from the output**, @@ -86,11 +88,24 @@ site-level signal; applying it per route would make one route's `critical` look ## `routes[].score` is not guaranteed to be the mean of `routes[].categories` -It usually is, which is exactly why the exception needs writing down. Measured with the field implemented, on -the repo's own fixture: **21 of 22 routes have `floor(mean(categories)) === score`**, and 17 of the 22 carry -only one category, where the two are equal by construction. One route disagrees. +**Whether they agree depends on the shape of the project, and the two available measurements point opposite +ways** — so neither can be presented as the typical case. -That one route is enough, because the disagreement is not a rounding artefact: +| corpus | keys | agreement | +| ------------------------------------------------------ | ----------------------------- | ------------------------------------------------------ | +| the repo's fixtures | 51 across nine projects | 98% (50/51) — but 46 of the 51 carry a single category | +| a 200-page project from the repo's own bench generator | 413, 400 of them two-category | **52% (213/413)** | + +The rule behind both numbers: **a single-category key agrees by construction, and a multi-category key +disagrees routinely and systematically.** Minimal fixtures are mostly single-category, which is a property of +minimal fixtures; a real project's pages carry both `seo` and `performance` results, and then the +multi-category case dominates. At the field project's scale the mean disagrees on most URL routes. + +The disagreement is not a rounding artefact. Every page key in the synthetic project reads `score: 92` against +a mean of 86 — `{ seo: 97, performance: 75 }`, six points apart — because the union ratio weights by inventory +while the mean weights the categories equally. + +The single measured instance, worked through: | value | denominator | result | | ------------------------ | ------------------------------------- | ------- | @@ -102,11 +117,13 @@ The mean of 95 and 100 is 97.5, not 96. `routes[].score` is one ratio against ev measured against; each category score is a ratio against that category's own inventory. Both are correct and they answer different questions. -A first draft of this section stated the non-coincidence flatly and generalised it to "a mean of ratios with -different denominators is not the ratio of the sums", which is not a theorem — equal ratios coincide for any -denominators. The honest form is the one in this heading: not guaranteed, commonly true anyway, and the docs -must say it that way too, because a user comparing their own report against a flat "it does not average" would -find the report contradicting the documentation on most routes. +Two drafts of this section were wrong in opposite directions. The first stated the non-coincidence flatly and +generalised it to "a mean of ratios with different denominators is not the ratio of the sums", which is not a +theorem — equal ratios coincide for any denominators. The second called it a rare exception, on the strength of +a fixture corpus that turned out to be atypical. The honest form is the heading: **not guaranteed**, agreeing +by construction in one shape and disagreeing systematically in the other. The docs must use that wording in +both directions — a flat "it does not average" is contradicted by any single-category route, and a flat "it +does" by any page carrying two categories. This is the third level of the scoring model where an aggregate is not guaranteed to be re-derivable from the parts below it — after `computeHealth` over category scores, and a category score over its key scores. Same From a102129d23c920db3e59d3eb280feecc826ddf7f Mon Sep 17 00:00:00 2001 From: oekazuma Date: Tue, 4 Aug 2026 19:35:16 +0900 Subject: [PATCH 04/12] docs: key the mean rule to equal ratios, not category count MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The third draft's rule contradicted the table above it: 200 of the synthetic corpus's 400 multi-category keys agree, all of them clean, so a rule keyed on category count predicts 3% agreement for a corpus measured at 52%. A key agrees when every category on it scores the same ratio — single-category and clean multi-category are both that case, and clean keys dominate a healthy project. Three real exceptions to the disagreement are stated, including the k=2 theorem that equal inventories or equal ratios are the only ones. Also drop an unmeasured field-scale prediction. --- ...2026-08-04-route-category-scores-design.md | 62 ++++++++++++++----- 1 file changed, 48 insertions(+), 14 deletions(-) diff --git a/docs/superpowers/specs/2026-08-04-route-category-scores-design.md b/docs/superpowers/specs/2026-08-04-route-category-scores-design.md index 034edde4b..f14e7ba3a 100644 --- a/docs/superpowers/specs/2026-08-04-route-category-scores-design.md +++ b/docs/superpowers/specs/2026-08-04-route-category-scores-design.md @@ -96,14 +96,47 @@ ways** — so neither can be presented as the typical case. | the repo's fixtures | 51 across nine projects | 98% (50/51) — but 46 of the 51 carry a single category | | a 200-page project from the repo's own bench generator | 413, 400 of them two-category | **52% (213/413)** | -The rule behind both numbers: **a single-category key agrees by construction, and a multi-category key -disagrees routinely and systematically.** Minimal fixtures are mostly single-category, which is a property of -minimal fixtures; a real project's pages carry both `seo` and `performance` results, and then the -multi-category case dominates. At the field project's scale the mean disagrees on most URL routes. - -The disagreement is not a rounding artefact. Every page key in the synthetic project reads `score: 92` against -a mean of 86 — `{ seo: 97, performance: 75 }`, six points apart — because the union ratio weights by inventory -while the mean weights the categories equally. +**The rule is about ratios, not category counts.** A key agrees exactly when **every category on it scores the +same ratio**, because then the union of the parts and the mean of the parts are the same number. Two very +different-looking keys are both that case: + +- a **single-category** key — the bucket _is_ the whole result set, so both numbers come from the identical + `computeScore` call. Exact and unbreakable: 59 of 59 single-category keys across both corpora agree, and no + counterexample is constructible. +- a **clean multi-category** key — every ratio is zero, so `{ architecture: 100, performance: 100 }` agrees for + the same reason. **Clean keys are most keys on a healthy project.** + +That second case is what an earlier draft missed, and missing it made the rule contradict the table above it: +200 of the synthetic corpus's 400 multi-category keys agree, all of them clean, while only 13 of its 213 +agreements are single-category. A rule keyed on category count predicts 3% agreement for a corpus measured at +52%. + +Neither corpus predicts what a given project will show. Minimal fixtures are mostly single-category, which is a +property of minimal fixtures; the synthetic generator's pages are uniformly flawed, which is not how a real +project looks. With 276 findings across 585 keys, most of the field project's keys are clean — and clean keys +agree. + +**When the ratios do differ the mean usually disagrees, systematically once the gap exceeds a point** — every +dirty page in the synthetic corpus, six points apart. "Usually" is the right word, because three exceptions are +real and two of them are ordinary: + +1. **A sub-point gap collapses under flooring.** One `seo` `info` beside a clean `performance` gives + `{ seo: 99, performance: 100 }`; the raw mean is 99.545 and the raw union 99.275, and both display **99**. + That is the most common shape of a lightly-flawed page. +2. **Equal observed inventories force agreement whatever the ratios.** For two categories this is a theorem: + mean and union coincide exactly when `(i₂ − i₁)(f₁i₂ − f₂i₁) = 0`, so either the inventories match or the + ratios do, and nothing else. Verified at `i₁ = i₂ = 28` with `f = 5` and `f = 10` — both display **73**. + Reachable by ordinary configuration, since turning rules off shrinks an inventory. +3. **Sporadic exact coincidences exist under the pristine default registry.** An exhaustive integer search + found sixteen, two on realistic component-only keys: clean `performance` with `correctness` at 36/96 and + `security` at 25/35 gives a union and a mean that are bit-identical at 63.690476. + +So the user-facing wording stays **not guaranteed** in both directions, and the reason is now a rule rather +than a frequency. + +The disagreement, where it happens, is not a rounding artefact. Every dirty page key in the synthetic project +reads `score: 92` against a floored mean of 86 — `{ seo: 97, performance: 75 }` — because the union ratio +weights by inventory while the mean weights the categories equally. The single measured instance, worked through: @@ -117,13 +150,14 @@ The mean of 95 and 100 is 97.5, not 96. `routes[].score` is one ratio against ev measured against; each category score is a ratio against that category's own inventory. Both are correct and they answer different questions. -Two drafts of this section were wrong in opposite directions. The first stated the non-coincidence flatly and -generalised it to "a mean of ratios with different denominators is not the ratio of the sums", which is not a -theorem — equal ratios coincide for any denominators. The second called it a rare exception, on the strength of -a fixture corpus that turned out to be atypical. The honest form is the heading: **not guaranteed**, agreeing +Three drafts of this section were wrong in three different ways, recorded because the error moved each time. +The first stated the non-coincidence flatly and generalised it to "a mean of ratios with different denominators +is not the ratio of the sums", which is not a theorem. The second called it a rare exception, on the strength of +a fixture corpus that turned out to be atypical. The third keyed the rule to category count, which contradicted +its own measurements, because a clean multi-category key agrees. The honest form is the heading: **not guaranteed**, agreeing by construction in one shape and disagreeing systematically in the other. The docs must use that wording in -both directions — a flat "it does not average" is contradicted by any single-category route, and a flat "it -does" by any page carrying two categories. +both directions — a flat "it does not average" is contradicted by any single-category route and by any clean +one, and a flat "it does" by any page whose categories score differently. This is the third level of the scoring model where an aggregate is not guaranteed to be re-derivable from the parts below it — after `computeHealth` over category scores, and a category score over its key scores. Same From ef2d85589c6484cfbdba78653e3efe9889f18a07 Mon Sep 17 00:00:00 2001 From: oekazuma Date: Tue, 4 Aug 2026 19:41:17 +0900 Subject: [PATCH 05/12] docs: drop a count this design cannot reproduce, fix the coincidence example MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The sporadic-coincidence exception carried three errors inherited from a reviewer's first search: a count that search had silently truncated, an example attached to the wrong key shape (it needs performance observed at 37, not its component-only 9), and "bit-identical" for doubles that sit one ulp apart. The count is now omitted rather than replaced, since it cannot be reproduced here. Also generalise the equal-inventory exception to any k, scope its uniqueness proof to k=2, and add the flooring qualifier the docs clause needed — exception 1 falsified it as written. --- ...2026-08-04-route-category-scores-design.md | 30 ++++++++++++++----- 1 file changed, 22 insertions(+), 8 deletions(-) diff --git a/docs/superpowers/specs/2026-08-04-route-category-scores-design.md b/docs/superpowers/specs/2026-08-04-route-category-scores-design.md index f14e7ba3a..636b3174f 100644 --- a/docs/superpowers/specs/2026-08-04-route-category-scores-design.md +++ b/docs/superpowers/specs/2026-08-04-route-category-scores-design.md @@ -123,13 +123,27 @@ real and two of them are ordinary: 1. **A sub-point gap collapses under flooring.** One `seo` `info` beside a clean `performance` gives `{ seo: 99, performance: 100 }`; the raw mean is 99.545 and the raw union 99.275, and both display **99**. That is the most common shape of a lightly-flawed page. -2. **Equal observed inventories force agreement whatever the ratios.** For two categories this is a theorem: - mean and union coincide exactly when `(i₂ − i₁)(f₁i₂ − f₂i₁) = 0`, so either the inventories match or the - ratios do, and nothing else. Verified at `i₁ = i₂ = 28` with `f = 5` and `f = 10` — both display **73**. - Reachable by ordinary configuration, since turning rules off shrinks an inventory. -3. **Sporadic exact coincidences exist under the pristine default registry.** An exhaustive integer search - found sixteen, two on realistic component-only keys: clean `performance` with `correctness` at 36/96 and - `security` at 25/35 gives a union and a mean that are bit-identical at 63.690476. +2. **Equal observed inventories force agreement whatever the ratios**, for any number of categories — the + identity is `Σ(fⱼ/i)/k = (Σfⱼ)/(k·i)`. Verified at `k = 2`, `i = 28` each, `f = 5` and `f = 10` (both + display **73**) and at `k = 3`, `i = 30` each, `f = 3/12/24` (both exactly 56.666666666666664). For two + categories these are provably the **only** exceptions besides equal ratios: mean and union coincide exactly + when `(i₂ − i₁)(f₁i₂ − f₂i₁) = 0`. That proof is for `k = 2` only; exception 3 exhibits coincidences at + `k ≥ 3` that neither condition explains. Reachable by ordinary configuration, since turning rules off + shrinks an inventory. +3. **Sporadic exact coincidences exist under the pristine default registry**, with neither equal ratios nor + equal inventories. One verified here end to end, on a component-only key carrying all five categories — + `seo` 5/5, `performance` 0/9, `correctness` 56/96, `security` 35/35, `architecture` 6/8 — where the union + and the mean are both a deficit of exactly 200/3 and agree as doubles, not merely after flooring. + + No count is given. A reviewer's exhaustive search reported one, then corrected it on a second pass after + finding the first had silently skipped part of the space; the existence of such keys is what matters here, + and a number this design cannot reproduce is the kind of figure it has already got wrong twice. + + Two cautions for anyone who recomputes: the coincidence depends on the **observed** inventory, not the + category's nominal one — a key observing `performance` at 37 (its route and component pairs together) can + coincide where the same failures at its component-only 9 give 56.43 against 63.69 — and "coincide" means + equal as exact rationals. As doubles the two can sit one ulp apart, 63.69047619047619 against + 63.69047619047618, which is why the agreement metric throughout this section is the displayed value. So the user-facing wording stays **not guaranteed** in both directions, and the reason is now a rule rather than a frequency. @@ -157,7 +171,7 @@ a fixture corpus that turned out to be atypical. The third keyed the rule to cat its own measurements, because a clean multi-category key agrees. The honest form is the heading: **not guaranteed**, agreeing by construction in one shape and disagreeing systematically in the other. The docs must use that wording in both directions — a flat "it does not average" is contradicted by any single-category route and by any clean -one, and a flat "it does" by any page whose categories score differently. +one, and a flat "it does" by any page whose categories score differently by more than flooring absorbs. This is the third level of the scoring model where an aggregate is not guaranteed to be re-derivable from the parts below it — after `computeHealth` over category scores, and a category score over its key scores. Same From 7992d51498f4a278aad24a915b0295029b0e3e95 Mon Sep 17 00:00:00 2001 From: oekazuma Date: Tue, 4 Aug 2026 19:46:09 +0900 Subject: [PATCH 06/12] docs: say exact coincidences, since flooring is a third way k=2 agrees --- .../specs/2026-08-04-route-category-scores-design.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/superpowers/specs/2026-08-04-route-category-scores-design.md b/docs/superpowers/specs/2026-08-04-route-category-scores-design.md index 636b3174f..8981cf55c 100644 --- a/docs/superpowers/specs/2026-08-04-route-category-scores-design.md +++ b/docs/superpowers/specs/2026-08-04-route-category-scores-design.md @@ -126,7 +126,7 @@ real and two of them are ordinary: 2. **Equal observed inventories force agreement whatever the ratios**, for any number of categories — the identity is `Σ(fⱼ/i)/k = (Σfⱼ)/(k·i)`. Verified at `k = 2`, `i = 28` each, `f = 5` and `f = 10` (both display **73**) and at `k = 3`, `i = 30` each, `f = 3/12/24` (both exactly 56.666666666666664). For two - categories these are provably the **only** exceptions besides equal ratios: mean and union coincide exactly + categories these are provably the **only exact coincidences** besides equal ratios: mean and union coincide exactly when `(i₂ − i₁)(f₁i₂ − f₂i₁) = 0`. That proof is for `k = 2` only; exception 3 exhibits coincidences at `k ≥ 3` that neither condition explains. Reachable by ordinary configuration, since turning rules off shrinks an inventory. From 0aaa60dcd29a97f1fa1e0a33df201a03ee50f2e0 Mon Sep 17 00:00:00 2001 From: oekazuma Date: Tue, 4 Aug 2026 19:48:30 +0900 Subject: [PATCH 07/12] docs: plan the per-route category scores --- .../plans/2026-08-04-route-category-scores.md | 441 ++++++++++++++++++ 1 file changed, 441 insertions(+) create mode 100644 docs/superpowers/plans/2026-08-04-route-category-scores.md diff --git a/docs/superpowers/plans/2026-08-04-route-category-scores.md b/docs/superpowers/plans/2026-08-04-route-category-scores.md new file mode 100644 index 000000000..fab37e357 --- /dev/null +++ b/docs/superpowers/plans/2026-08-04-route-category-scores.md @@ -0,0 +1,441 @@ +# Per-route category scores Implementation Plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** Let a reader of the JSON report see what each route scored in each category, so a category's number +can be traced to the routes that produced it. + +**Architecture:** `scoresByCategory` gains an optional `ScoreOptions` parameter forwarded to `computeScore`; +the JSON reporter calls it per route with `{ applyCriticalCap: false }` and puts the resulting scores on each +`routes[]` entry. Nothing else in the scoring model changes. + +**Tech Stack:** TypeScript (ESM, `.js` import specifiers), vitest, oxlint + oxfmt, Astro Starlight docs. + +**Spec:** `docs/superpowers/specs/2026-08-04-route-category-scores-design.md`. **Read it before Task 1** — it +was rejected four times by adversarial review, and three of those rejections were claims that sounded +authoritative and were false. Every number in this plan is measured; if one disagrees with what you compute, +stop and report it rather than adjusting the test. + +## Global Constraints + +- **`packages/core/src/` is runtime-agnostic**: no `node:` imports, no I/O, no runtime-specific globals. +- **`scoresByCategory`'s existing behaviour must not change.** Its new third parameter defaults to `{}`, and + every current caller passes nothing. `computeHealth` in particular **depends on a category still being + capped at 79 by a `critical`** — `2026-07-31-score-honesty-design.md` requires a capped category to pull + Health down. +- **The per-route map carries scores only, never `scoreModel`.** +- **Only the categories that produced a result on that route appear.** A route with no `architecture` result + must **not** show `architecture: 100`; that would claim a measurement that never happened. +- **The critical cap is off on the per-route path**, matching `routes[].score`. +- **Comments and docs are for the next reader** (`AGENTS.md`): a line earns its place only when it says + something the code cannot. Prefer one line over three. +- **Never name another tool, linter, plugin or product** in any doc, comment or commit message. +- **en/ja docs ship together.** +- Conventional commits, scoped by package. **A changeset is required** — `minor` (a new report field), listing + `@svelte-vitals/core`, `svelte-vitals`, `@svelte-vitals/vite`. + +## File Structure + +| File | Responsibility | +| ---------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------ | +| `packages/core/src/scoring/score.ts` | `scoresByCategory` gains `options: ScoreOptions = {}`, forwarded to `computeScore`. Three lines. | +| `packages/core/src/reporter/json.ts` | `JsonReport['routes']` gains `categories`; `buildJsonReport` fills it. | +| `packages/core/test/score.test.ts` | the new parameter, and that omitting it still caps. | +| `packages/core/test/json-report.test.ts` | the field's contents, the absence rule, the cap, and the disagreement with `routes[].score`. | +| `packages/core/test/html-report.test.ts`, `packages/vite/test/app-shell-static.test.ts`, `packages/vite/test/ui-dashboard.test.ts` | eight literal route constructions that stop compiling. | +| `docs/src/content/docs/guides/(reporting)/reporters.md` + ja | the sample and the prose. | +| `.changeset/route-category-scores.md` | **new** | + +Two tasks. The scoring parameter is separable from the reporter that uses it — a reviewer can reject one while +approving the other — and the parameter is where the risk lives, because getting its default wrong silently +changes Health. + +--- + +## Task 1: `scoresByCategory` takes scoring options + +**Files:** + +- Modify: `packages/core/src/scoring/score.ts` (`scoresByCategory`, around line 120) +- Test: `packages/core/test/score.test.ts` (append) + +**Interfaces:** + +- Consumes: `ScoreOptions`, `computeScore`, `Category`, `ScoreResult` — all already in that file. +- Produces: + + ```ts + export function scoresByCategory( + results: Result[], + config: Config, + options: ScoreOptions = {} + ): Partial>; + ``` + + Task 2 calls it as `scoresByCategory(rs, config, { applyCriticalCap: false })`. + +- [ ] **Step 1: Write the failing tests** + +Append to `packages/core/test/score.test.ts`. The `fail` helper already exists at the top of that file; do not +redefine it. + +```ts +describe('scoresByCategory — scoring options', () => { + const config = defineConfig({}); + // One failing `seo` critical on one route: the ratio gives 100 − 1500/110 = 86.36, the cap gives 79. + const results = [fail('seo/title-presence', '/a', 'critical')]; + + it('caps a category when called without options, which computeHealth depends on', () => { + expect(scoresByCategory(results, config).seo!.score).toBe(79); + expect(scoresByCategory(results, config).seo!.scoreModel.criticalCap).toBe(79); + }); + + it('leaves the category uncapped when the cap is switched off', () => { + const sr = scoresByCategory(results, config, { applyCriticalCap: false }).seo!; + expect(sr.score).toBe(86); + expect(sr.scoreModel.criticalCap).toBeNull(); + }); +}); +``` + +- [ ] **Step 2: Run the tests to verify they fail** + +Run from `packages/core`: `../../node_modules/.bin/vitest run test/score.test.ts -t 'scoring options'` +Expected: FAIL — `scoresByCategory` accepts two arguments, so the third is a type error and the uncapped +expectation gets 79. + +Only two tests: `options` is forwarded as one object, so proving `applyCriticalCap` reaches `computeScore` +proves `rules` does too. A third test for `rules` would assert the same forwarding twice. + +- [ ] **Step 3: Add the parameter** + +In `packages/core/src/scoring/score.ts`, change the signature and the one call inside: + +```ts +/** Compute an independent score per category present in `results` (issue #10). */ +export function scoresByCategory( + results: Result[], + config: Config, + options: ScoreOptions = {} +): Partial> { +``` + +and the loop's final line from `computeScore(rs, config)` to `computeScore(rs, config, options)`. + +Change nothing else. In particular do **not** change `computeHealth`, which calls +`scoresByCategory(results, config)` and must keep capping. + +- [ ] **Step 4: Run the tests to verify they pass** + +Run: `../../node_modules/.bin/vitest run test/score.test.ts -t 'scoring options'` +Expected: PASS, 2 tests. + +- [ ] **Step 5: Confirm no existing behaviour moved** + +```bash +(cd packages/core && ../../node_modules/.bin/vitest run) +(cd packages/cli && ../../node_modules/.bin/vitest run) +(cd packages/vite && ../../node_modules/.bin/vitest run) +``` + +Expected: all pass, unchanged counts. An optional trailing parameter defaulting to `{}` is behaviour-identical +because `computeScore` already defaults its own `options` to `{}` — so **any** failure here means the default +was got wrong, and that is a Health regression, not a test to adjust. Report it rather than editing an +expectation. + +- [ ] **Step 6: Typecheck and lint** + +```bash +(cd packages/core && ../../node_modules/.bin/tsc --noEmit) +node_modules/.bin/oxlint . && node_modules/.bin/oxfmt --check . +``` + +Expected: clean. + +- [ ] **Step 7: Commit** + +```bash +git add packages/core/src/scoring/score.ts packages/core/test/score.test.ts +git commit -m "feat(core): let scoresByCategory take scoring options" +``` + +--- + +## Task 2: The report field, its consumers, and the docs + +**Files:** + +- Modify: `packages/core/src/reporter/json.ts` (the `JsonReport` interface around line 39, and + `buildJsonReport`'s `routes` assembly around line 82) +- Modify: `packages/core/test/html-report.test.ts` (4 literal route constructions), + `packages/vite/test/app-shell-static.test.ts` and `packages/vite/test/ui-dashboard.test.ts` (4 more) +- Test: `packages/core/test/json-report.test.ts` (append) +- Modify: `docs/src/content/docs/guides/(reporting)/reporters.md` and + `docs/src/content/docs/ja/guides/(reporting)/reporters.md` +- Create: `.changeset/route-category-scores.md` + +**Interfaces:** + +- Consumes: `scoresByCategory(results, config, options)` from Task 1. +- Produces: `JsonReport['routes']` becomes + `Array<{ route: string; score: number; categories: Record; issues: JsonIssue[] }>`. + +- [ ] **Step 1: Write the failing tests** + +Append to `packages/core/test/json-report.test.ts`. That file has a module-level `config` and a module-level +`results` array used by the existing suites — do **not** reuse or extend `results`, since these cases need +specific inventories. Reuse `config`, and declare each case's own results inline as below. + +```ts +describe('buildJsonReport — per-route category scores', () => { + const config = defineConfig({}); + const meta = { version: '0.0.0' }; + + it('scores each category present on the route', () => { + // seo::route inventory 110, one warning failing -> 100 − 500/110 = 95.45 -> 95. + // performance::route inventory 28, nothing failing -> 100. + const results: Result[] = [ + { + id: 'seo/canonical-url', + category: 'seo', + severity: 'warning', + detection: { presence: 'none', value: 'absent' }, + route: '/a', + message: 'x' + }, + { + id: 'performance/preconnect', + category: 'performance', + severity: 'warning', + detection: { presence: 'own', value: 'static' }, + route: '/a', + message: 'ok' + } + ]; + const report = buildJsonReport(results, config, meta); + expect(report.routes[0]!.categories).toEqual({ seo: 95, performance: 100 }); + }); + + it('omits a category that produced no result on the route', () => { + const results: Result[] = [ + { + id: 'seo/canonical-url', + category: 'seo', + severity: 'warning', + detection: { presence: 'none', value: 'absent' }, + route: '/a', + message: 'x' + } + ]; + const report = buildJsonReport(results, config, meta); + expect(Object.keys(report.routes[0]!.categories)).toEqual(['seo']); + expect(report.routes[0]!.categories.architecture).toBeUndefined(); + }); + + it('does not cap a category at 79 for a route carrying a critical', () => { + // The ratio gives 100 − 1500/110 = 86. Asserting routes[].score instead would pass on a + // capped implementation, because that path is already cap-free. + const results: Result[] = [ + { + id: 'seo/title-presence', + category: 'seo', + severity: 'critical', + detection: { presence: 'none', value: 'absent' }, + route: '/a', + message: 'x' + } + ]; + const report = buildJsonReport(results, config, meta); + expect(report.routes[0]!.categories.seo).toBe(86); + }); + + it('keeps routes[].score as the union ratio, which need not equal the category mean', () => { + // The union observes both pairs: inventory 138, failed 5 -> 100 − 500/138 = 96. + // The categories are 95 and 100, whose mean is 97.5. Both numbers are asserted on one input + // because their disagreement is the property the spec records. + const results: Result[] = [ + { + id: 'seo/canonical-url', + category: 'seo', + severity: 'warning', + detection: { presence: 'none', value: 'absent' }, + route: '/a', + message: 'x' + }, + { + id: 'performance/preconnect', + category: 'performance', + severity: 'warning', + detection: { presence: 'own', value: 'static' }, + route: '/a', + message: 'ok' + } + ]; + const report = buildJsonReport(results, config, meta); + expect(report.routes[0]!.score).toBe(96); + expect(report.routes[0]!.categories).toEqual({ seo: 95, performance: 100 }); + }); +}); +``` + +- [ ] **Step 2: Run the tests to verify they fail** + +Run from `packages/core`: `../../node_modules/.bin/vitest run test/json-report.test.ts -t 'per-route category'` +Expected: FAIL — `categories` is `undefined` on every route entry. + +- [ ] **Step 3: Add the field** + +In `packages/core/src/reporter/json.ts`, change the interface line: + +```ts +routes: Array<{ route: string; score: number; categories: Record; issues: JsonIssue[] }>; +``` + +and the `routes` assembly in `buildJsonReport`: + +```ts + .map(({ route, results: rs }) => ({ + route, + score: computeScore(rs, config, { applyCriticalCap: false }).score, + // Per category, scored against that category's own inventory — so this does not average to `score`, + // which is one ratio over the union of the pairs the route touched. + categories: Object.fromEntries( + Object.entries(scoresByCategory(rs, config, { applyCriticalCap: false })).map(([cat, sr]) => [cat, sr!.score]) + ), + issues: rs + .filter((r) => isPenalized(r.detection, config.treatDynamicAs)) + .map((r) => ({ ...issueOf(r), severity: effectiveSeverity(r, config) })) + })); +``` + +Add `scoresByCategory` to the existing import from `../scoring/score.js`. + +- [ ] **Step 4: Run the tests to verify they pass** + +Run: `../../node_modules/.bin/vitest run test/json-report.test.ts -t 'per-route category'` +Expected: PASS, 4 tests. + +- [ ] **Step 5: Fix the eight route constructions that stopped compiling** + +`categories` is required, so every literal `JsonReport['routes']` element needs it. Rebuild core first, then +typecheck each package — a grep for the type name will not find these, because several are typed through +another type: + +```bash +(cd packages/core && ../../node_modules/.bin/tsup) +for p in core cli vite; do (cd packages/$p && ../../node_modules/.bin/tsc --noEmit); done +``` + +Expected: 4 errors in `packages/core/test/html-report.test.ts`, and 4 across +`packages/vite/test/app-shell-static.test.ts` and `packages/vite/test/ui-dashboard.test.ts`. The CLI and the +markdown reporter compile clean. + +Add `categories: {}` to each — these fixtures exercise rendering, not scoring, and an empty map is the honest +value for a hand-built route with no scored results. Do not invent scores for them. + +- [ ] **Step 6: Assert the field reaches the HTML snapshot** + +Append to `packages/core/test/html-report.test.ts`. It already imports `formatHtmlReport` from `../src/index.js` +and builds a `JsonReport` literal named `report` with a `model()` helper for `ScoreModel` — that literal is one +of the four constructions you fixed in Step 5. + +```ts +it('carries per-route category scores into the embedded snapshot', () => { + // `sanitizeReport` spreads each route, so this passes today — it exists to fail if that spread is + // ever replaced by an explicit field list. + const results: Result[] = [ + { + id: 'seo/canonical-url', + category: 'seo', + severity: 'warning', + detection: { presence: 'none', value: 'absent' }, + route: '/a', + message: 'x' + } + ]; + const html = formatHtmlReport(results, defineConfig({}), { version: '0.0.0' }); + expect(html).toContain('"categories":{"seo":95}'); +}); +``` + +If the embedded JSON is formatted differently — check by logging a slice of the output rather than guessing — +assert on the parsed snapshot instead of a substring, and say in your report which form you used. + +- [ ] **Step 7: Run everything** + +```bash +for p in core cli vite; do (cd packages/$p && ../../node_modules/.bin/vitest run); done +(cd packages/core && ../../node_modules/.bin/tsup) +for p in core cli vite; do (cd packages/$p && ../../node_modules/.bin/tsc --noEmit); done +node_modules/.bin/oxlint . && node_modules/.bin/oxfmt --check . +``` + +Expected: all green. (`packages/mcp` has no `tsconfig.json`; skip it.) + +- [ ] **Step 8: Confirm the real report grows as the spec measured** + +```bash +BIN="$(pwd)/packages/cli/dist/bin.js" +(cd packages/cli/test/fixtures/basic-project && node "$BIN" --reporter json > "$TMPDIR/after.json") +wc -c < "$TMPDIR/after.json" +``` + +Expected: **69,003 bytes**, up from 67,656 — 1,347 bytes over 22 routes. Report the actual figure. A +materially different number means the field's shape is not what the spec measured; say so rather than +adjusting the spec. + +- [ ] **Step 9: Update both docs pages** + +`docs/src/content/docs/guides/(reporting)/reporters.md` shows the `routes[]` sample around line 53. Add +`categories` to the sample beside `score`, with a comment, and add prose saying what it is. + +**The prose must say the relationship is not guaranteed, in both directions.** A flat "it does not average to +`score`" is contradicted by any single-category route and by any clean one; a flat "it does" is contradicted by +any route whose categories differ by more than flooring absorbs. Keep it to one or two sentences. Do the same +in `docs/src/content/docs/ja/guides/(reporting)/reporters.md` — the Japanese must be idiomatic, not a +transliteration. + +- [ ] **Step 10: Write the changeset** + +Create `.changeset/route-category-scores.md`: + +```md +--- +'@svelte-vitals/core': minor +'svelte-vitals': minor +'@svelte-vitals/vite': minor +--- + +Each entry in the JSON report's `routes` array now carries a `categories` map of category name to score. + +A category's score is an average over its keys, so a category that looks wrong gives no clue which routes +produced it. The report listed each route's findings but not what each route scored per category, and since a +key's score became a ratio against the severity-weighted inventory of the checks it was measured against, that +number is no longer something a reader can reconstruct by hand. + +Only the categories that produced a result on a route appear, so an absent category means "not measured here" +rather than "perfect here". A route's `categories` values are **not guaranteed** to average to its `score`: +`score` is one ratio over everything the route was measured against, while each category score uses that +category's own inventory. They agree whenever every category on the route scores the same ratio — including +every route with no findings — and can differ by several points otherwise. +``` + +- [ ] **Step 11: Commit** + +```bash +git add packages/core/src/reporter/json.ts packages/core/test packages/vite/test docs .changeset +git commit -m "feat(core): report each route's score per category" +``` + +--- + +## Notes for whoever runs this + +- A full-workspace `pnpm` command fails in this sandbox for a known, pre-existing reason (the `docs` package's + dependencies), and `pnpm --filter build` fails on a reflink error. Use `../../node_modules/.bin/tsup` + from inside the package. +- The spec records three things as deliberately out of scope, so do not implement them: rendering the field in + the HTML report or the dev dashboard, making the aggregates re-derivable, and a per-route `scoreModel`. +- The spec's own history is worth one minute of your time before you start. Its claim about how often + `routes[].score` equals the category mean was wrong three times in three different ways. If you find + yourself about to write a test asserting that relationship in general, re-read that section first. From ec737cce3fb08260aca00061c52cc133dbf75e87 Mon Sep 17 00:00:00 2001 From: oekazuma Date: Tue, 4 Aug 2026 20:03:55 +0900 Subject: [PATCH 08/12] feat(core): let scoresByCategory take scoring options --- packages/core/src/scoring/score.ts | 8 ++++++-- packages/core/test/score.test.ts | 17 +++++++++++++++++ 2 files changed, 23 insertions(+), 2 deletions(-) diff --git a/packages/core/src/scoring/score.ts b/packages/core/src/scoring/score.ts index b39ccd3a7..25c13a7c7 100644 --- a/packages/core/src/scoring/score.ts +++ b/packages/core/src/scoring/score.ts @@ -118,7 +118,11 @@ export function computeScore(results: Result[], config: Config, options: ScoreOp } /** Compute an independent score per category present in `results` (issue #10). */ -export function scoresByCategory(results: Result[], config: Config): Partial> { +export function scoresByCategory( + results: Result[], + config: Config, + options: ScoreOptions = {} +): Partial> { const byCat = new Map(); for (const r of results) { const cat = r.category ?? 'seo'; @@ -127,7 +131,7 @@ export function scoresByCategory(results: Result[], config: Config): Partial> = {}; - for (const [cat, rs] of byCat) out[cat] = computeScore(rs, config); + for (const [cat, rs] of byCat) out[cat] = computeScore(rs, config, options); return out; } diff --git a/packages/core/test/score.test.ts b/packages/core/test/score.test.ts index 8f9822d39..54b844879 100644 --- a/packages/core/test/score.test.ts +++ b/packages/core/test/score.test.ts @@ -178,6 +178,23 @@ describe('scoresByCategory', () => { }); }); +describe('scoresByCategory — scoring options', () => { + const config = defineConfig({}); + // One failing `seo` critical on one route: the ratio gives 100 − 1500/110 = 86.36, the cap gives 79. + const results = [fail('seo/title-presence', '/a', 'critical')]; + + it('caps a category when called without options, which computeHealth depends on', () => { + expect(scoresByCategory(results, config).seo!.score).toBe(79); + expect(scoresByCategory(results, config).seo!.scoreModel.criticalCap).toBe(79); + }); + + it('leaves the category uncapped when the cap is switched off', () => { + const sr = scoresByCategory(results, config, { applyCriticalCap: false }).seo!; + expect(sr.score).toBe(86); + expect(sr.scoreModel.criticalCap).toBeNull(); + }); +}); + const CONFIG = defineConfig({}); /** `keys` route keys, the first `findings` of them carrying one `info` finding. */ From 56a7ec93c65fe19ff651051fe00bc2111f13b1ef Mon Sep 17 00:00:00 2001 From: oekazuma Date: Tue, 4 Aug 2026 20:08:26 +0900 Subject: [PATCH 09/12] test(core): name the cap test after its behaviour, not its reason --- packages/core/test/score.test.ts | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/packages/core/test/score.test.ts b/packages/core/test/score.test.ts index 54b844879..017a39986 100644 --- a/packages/core/test/score.test.ts +++ b/packages/core/test/score.test.ts @@ -183,7 +183,8 @@ describe('scoresByCategory — scoring options', () => { // One failing `seo` critical on one route: the ratio gives 100 − 1500/110 = 86.36, the cap gives 79. const results = [fail('seo/title-presence', '/a', 'critical')]; - it('caps a category when called without options, which computeHealth depends on', () => { + it('caps a category at 79 when called without options', () => { + // computeHealth calls it this way and depends on a capped category pulling Health down. expect(scoresByCategory(results, config).seo!.score).toBe(79); expect(scoresByCategory(results, config).seo!.scoreModel.criticalCap).toBe(79); }); From 76d9f68f81e77fc88a628f2eaed38c79a82a3e1e Mon Sep 17 00:00:00 2001 From: oekazuma Date: Tue, 4 Aug 2026 20:15:22 +0900 Subject: [PATCH 10/12] feat(core): report each route's score per category --- .changeset/route-category-scores.md | 18 ++++ .../docs/guides/(reporting)/reporters.md | 3 + .../docs/ja/guides/(reporting)/reporters.md | 3 + packages/core/src/reporter/json.ts | 9 +- packages/core/test/html-report.test.ts | 34 ++++++- packages/core/test/json-report.test.ts | 90 +++++++++++++++++++ packages/vite/test/app-shell-static.test.ts | 1 + packages/vite/test/ui-dashboard.test.ts | 3 + 8 files changed, 156 insertions(+), 5 deletions(-) create mode 100644 .changeset/route-category-scores.md diff --git a/.changeset/route-category-scores.md b/.changeset/route-category-scores.md new file mode 100644 index 000000000..5dbb39096 --- /dev/null +++ b/.changeset/route-category-scores.md @@ -0,0 +1,18 @@ +--- +'@svelte-vitals/core': minor +'svelte-vitals': minor +'@svelte-vitals/vite': minor +--- + +Each entry in the JSON report's `routes` array now carries a `categories` map of category name to score. + +A category's score is an average over its keys, so a category that looks wrong gives no clue which routes +produced it. The report listed each route's findings but not what each route scored per category, and since a +key's score became a ratio against the severity-weighted inventory of the checks it was measured against, that +number is no longer something a reader can reconstruct by hand. + +Only the categories that produced a result on a route appear, so an absent category means "not measured here" +rather than "perfect here". A route's `categories` values are **not guaranteed** to average to its `score`: +`score` is one ratio over everything the route was measured against, while each category score uses that +category's own inventory. They agree whenever every category on the route scores the same ratio — including +every route with no findings — and can differ by several points otherwise. diff --git a/docs/src/content/docs/guides/(reporting)/reporters.md b/docs/src/content/docs/guides/(reporting)/reporters.md index 170e3b4e1..8f88c36ce 100644 --- a/docs/src/content/docs/guides/(reporting)/reporters.md +++ b/docs/src/content/docs/guides/(reporting)/reporters.md @@ -54,6 +54,7 @@ svelte-vitals --reporter json { "route": "/about", // a route id, or a source file path for file-scoped rules "score": 95, // share of this route's rule inventory (by category/scope, weighted by severity) left intact + "categories": { "seo": 94 }, // per category present on this route, scored against that category's own inventory "issues": [ { "id": "seo/single-h1", // the rule id @@ -81,6 +82,8 @@ Two field names are worth pointing out, because guessing them wrongly fails sile `line`, `docsUrl` and `fix` are present only when the rule supplies them, and `location` only for a finding tied to a file. `issues` lists **failing** findings only — passing checks are counted in `summary.passed` but are not listed. A route with no failures still appears in `routes`, with an empty `issues` array and its own score. +`categories` holds only the categories that produced a result on that route — an absent category means "not measured here," not "perfect here." Its values are **not guaranteed** to average to the route's own `score`, in either direction: `score` is one ratio over everything the route was measured against, while each category score uses that category's own inventory. They agree whenever every category on the route scores the same ratio (including every route with no findings) and can differ by several points otherwise. + `rules` answers a question the rest of the report cannot: **whether a rule ran at all.** `issues` lists only failing findings, so a rule that found nothing leaves no trace there — and a rule disabled at the top level (`--ignore`, `--rules`, `--category`, or `rules: { id: 'off' }` in config) leaves the same absence. diff --git a/docs/src/content/docs/ja/guides/(reporting)/reporters.md b/docs/src/content/docs/ja/guides/(reporting)/reporters.md index 91e104fc1..d7bc6af0f 100644 --- a/docs/src/content/docs/ja/guides/(reporting)/reporters.md +++ b/docs/src/content/docs/ja/guides/(reporting)/reporters.md @@ -54,6 +54,7 @@ svelte-vitals --reporter json { "route": "/about", // ルート ID。ファイル単位のルールではソースファイルのパス "score": 95, // このルートが属するカテゴリ/スコープのルール一覧(重大度で重み付け)のうち、無傷で残った割合 + "categories": { "seo": 94 }, // このルートで結果が出たカテゴリごとのスコア。そのカテゴリ自身の一覧に対する割合 "issues": [ { "id": "seo/single-h1", // ルール ID @@ -81,6 +82,8 @@ svelte-vitals --reporter json `line`・`docsUrl`・`fix` はルールが提供した場合のみ、`location` はファイルに紐づく検出の場合のみ現れます。`issues` に並ぶのは**失敗した検出のみ**です。合格したチェックは `summary.passed` に数として計上されますが、一覧には出ません。失敗が1件も無いルートも `routes` には現れ、`issues` が空配列のまま自分のスコアを持ちます。 +`categories` にはそのルートで実際に結果が出たカテゴリだけが並びます。あるカテゴリが無いのは「ここでは測定していない」という意味であり、「ここは満点だった」という意味ではありません。この値がルート自身の `score` の平均になっているとは、どちらの向きにも**保証されません**。`score` はそのルートが測定対象とした全体に対する一つの割合であるのに対し、各カテゴリのスコアはそのカテゴリ自身の一覧に対する割合だからです。ルート上のすべてのカテゴリが同じ割合を示すとき(検出結果が無いルートも含む)のみ両者は一致し、それ以外では数ポイントずれることがあります。 + `rules` は、レポートの他の部分では答えられない問い、**そのルールがそもそも実行されたかどうか**に答えます。`issues` に載るのは失敗した検出結果だけなので、何も検出しなかったルールはそこに痕跡を残しません。トップレベルで無効化した(`--ignore`、`--rules`、`--category`、または設定ファイルの `rules: { id: 'off' }`)ルールも同じく現れません。判定には `rules` を見ます。存在すれば実行された、存在しなければトップレベルで除外された、という意味です。ただし一つだけ例外があり、それは次に述べます。 件数が表すのはツリーではなくレポートそのものです。baseline、抑制、`--diff` によるフィルタリングはレポートを組み立てる前に適用されるため、検出結果がすべて抑制されたルールも `findings: 0` のまま `rules` に残ります。`overrides` で無効化したルールも同様です。トップレベルとは異なり、`overrides` はルールが実行された後にその結果(合格分も含む)を取り除くため、`{ "findings": 0, "passed": 0 }` として現れ、選択されて何も検出しなかったルールと見分けがつきません。`rules` に存在することが保証するのは、`--ignore`・`--rules`・`--category`・設定ファイルのトップレベルの `rules` で除外されなかったことだけで、`overrides` が検出結果を何も残さなかったことまでは保証しません。 diff --git a/packages/core/src/reporter/json.ts b/packages/core/src/reporter/json.ts index fc489f854..7c818cfe5 100644 --- a/packages/core/src/reporter/json.ts +++ b/packages/core/src/reporter/json.ts @@ -1,5 +1,5 @@ import type { Category, Config, Result } from '../types.js'; -import { computeScore, computeHealth, type ScoreModel } from '../scoring/score.js'; +import { computeScore, computeHealth, scoresByCategory, type ScoreModel } from '../scoring/score.js'; import { summarize, effectiveSeverity, type Summary } from '../summary.js'; import { isPenalized } from '../rule.js'; @@ -36,7 +36,7 @@ export interface JsonReport { categories: Record; summary: Summary; rules: Record; - routes: Array<{ route: string; score: number; issues: JsonIssue[] }>; + routes: Array<{ route: string; score: number; categories: Record; issues: JsonIssue[] }>; siteIssues: JsonIssue[]; } @@ -84,6 +84,11 @@ export function buildJsonReport( .map(({ route, results: rs }) => ({ route, score: computeScore(rs, config, { applyCriticalCap: false }).score, + // Per category, scored against that category's own inventory — so this does not average to `score`, + // which is one ratio over the union of the pairs the route touched. + categories: Object.fromEntries( + Object.entries(scoresByCategory(rs, config, { applyCriticalCap: false })).map(([cat, sr]) => [cat, sr!.score]) + ), issues: rs .filter((r) => isPenalized(r.detection, config.treatDynamicAs)) .map((r) => ({ ...issueOf(r), severity: effectiveSeverity(r, config) })) diff --git a/packages/core/test/html-report.test.ts b/packages/core/test/html-report.test.ts index 0e039d127..ea70f2ab7 100644 --- a/packages/core/test/html-report.test.ts +++ b/packages/core/test/html-report.test.ts @@ -1,6 +1,14 @@ import { describe, it, expect } from 'vitest'; -import { buildHtmlDocument, formatHtmlReport, renderAppShell, escapeHtml, safeHref, scoreBand } from '../src/index.js'; -import type { AppSnapshot } from '../src/index.js'; +import { + buildHtmlDocument, + formatHtmlReport, + renderAppShell, + escapeHtml, + safeHref, + scoreBand, + defineConfig +} from '../src/index.js'; +import type { AppSnapshot, Result } from '../src/index.js'; import type { JsonReport } from '../src/reporter/json.js'; import type { ScoreModel } from '../src/scoring/score.js'; @@ -17,10 +25,11 @@ const report: JsonReport = { summary: { critical: 1, warning: 2, info: 1, passed: 37, dynamic: 3 }, rules: {}, routes: [ - { route: '/', score: 100, issues: [] }, + { route: '/', score: 100, categories: {}, issues: [] }, { route: '/products/[id]', score: 40, + categories: {}, issues: [ { id: 'seo/title-presence', @@ -149,6 +158,23 @@ describe('formatHtmlReport', () => { expect(out).toContain(''); expect(extractEmbeddedSnapshot(out).live).toBe(false); }); + + it('carries per-route category scores into the embedded snapshot', () => { + // `sanitizeReport` spreads each route, so this passes today — it exists to fail if that spread is + // ever replaced by an explicit field list. + const results: Result[] = [ + { + id: 'seo/canonical-url', + category: 'seo', + severity: 'warning', + detection: { presence: 'none', value: 'absent' }, + route: '/a', + message: 'x' + } + ]; + const html = formatHtmlReport(results, defineConfig({}), { version: '0.0.0' }); + expect(html).toContain('"categories":{"seo":95}'); + }); }); describe('parity with the live dashboard shell', () => { @@ -181,6 +207,7 @@ describe('safety hardening (buildHtmlDocument is a public API; JsonReport is loo { route: '/x', score: 0, + categories: {}, issues: [ { id: 'seo/title-presence', @@ -209,6 +236,7 @@ describe('safety hardening (buildHtmlDocument is a public API; JsonReport is loo { route: '/a', score: 80, + categories: {}, issues: [ { id: 'seo/title-presence', diff --git a/packages/core/test/json-report.test.ts b/packages/core/test/json-report.test.ts index 049d66ed5..46be9e267 100644 --- a/packages/core/test/json-report.test.ts +++ b/packages/core/test/json-report.test.ts @@ -145,3 +145,93 @@ describe('buildJsonReport — per-rule evidence', () => { expect(parsed.rules).toEqual({ 'seo/single-h1': { findings: 0, passed: 0 } }); }); }); + +describe('buildJsonReport — per-route category scores', () => { + const config = defineConfig({}); + const meta = { version: '0.0.0' }; + + it('scores each category present on the route', () => { + // seo::route inventory 110, one warning failing -> 100 − 500/110 = 95.45 -> 95. + // performance::route inventory 28, nothing failing -> 100. + const results: Result[] = [ + { + id: 'seo/canonical-url', + category: 'seo', + severity: 'warning', + detection: { presence: 'none', value: 'absent' }, + route: '/a', + message: 'x' + }, + { + id: 'performance/preconnect', + category: 'performance', + severity: 'warning', + detection: { presence: 'own', value: 'static' }, + route: '/a', + message: 'ok' + } + ]; + const report = buildJsonReport(results, config, meta); + expect(report.routes[0]!.categories).toEqual({ seo: 95, performance: 100 }); + }); + + it('omits a category that produced no result on the route', () => { + const results: Result[] = [ + { + id: 'seo/canonical-url', + category: 'seo', + severity: 'warning', + detection: { presence: 'none', value: 'absent' }, + route: '/a', + message: 'x' + } + ]; + const report = buildJsonReport(results, config, meta); + expect(Object.keys(report.routes[0]!.categories)).toEqual(['seo']); + expect(report.routes[0]!.categories.architecture).toBeUndefined(); + }); + + it('does not cap a category at 79 for a route carrying a critical', () => { + // The ratio gives 100 − 1500/110 = 86. Asserting routes[].score instead would pass on a + // capped implementation, because that path is already cap-free. + const results: Result[] = [ + { + id: 'seo/title-presence', + category: 'seo', + severity: 'critical', + detection: { presence: 'none', value: 'absent' }, + route: '/a', + message: 'x' + } + ]; + const report = buildJsonReport(results, config, meta); + expect(report.routes[0]!.categories.seo).toBe(86); + }); + + it('keeps routes[].score as the union ratio, which need not equal the category mean', () => { + // The union observes both pairs: inventory 138, failed 5 -> 100 − 500/138 = 96. + // The categories are 95 and 100, whose mean is 97.5. Both numbers are asserted on one input + // because their disagreement is the property the spec records. + const results: Result[] = [ + { + id: 'seo/canonical-url', + category: 'seo', + severity: 'warning', + detection: { presence: 'none', value: 'absent' }, + route: '/a', + message: 'x' + }, + { + id: 'performance/preconnect', + category: 'performance', + severity: 'warning', + detection: { presence: 'own', value: 'static' }, + route: '/a', + message: 'ok' + } + ]; + const report = buildJsonReport(results, config, meta); + expect(report.routes[0]!.score).toBe(96); + expect(report.routes[0]!.categories).toEqual({ seo: 95, performance: 100 }); + }); +}); diff --git a/packages/vite/test/app-shell-static.test.ts b/packages/vite/test/app-shell-static.test.ts index 7d0049378..649b545a1 100644 --- a/packages/vite/test/app-shell-static.test.ts +++ b/packages/vite/test/app-shell-static.test.ts @@ -21,6 +21,7 @@ const report: JsonReport = { { route: '/blog/hello', score: 50, + categories: {}, issues: [ { id: 'seo/title-presence', diff --git a/packages/vite/test/ui-dashboard.test.ts b/packages/vite/test/ui-dashboard.test.ts index 53aa1d5ea..1d74ff368 100644 --- a/packages/vite/test/ui-dashboard.test.ts +++ b/packages/vite/test/ui-dashboard.test.ts @@ -14,6 +14,7 @@ const baseSnapshot: DashboardSnapshot = { { route: '/a', score: 80, + categories: {}, issues: [ { id: 'seo/title-presence', @@ -65,6 +66,7 @@ describe('renderDashboardShell', () => { { route: '/a', score: 80, + categories: {}, issues: [ { id: 'seo/title-presence', @@ -111,6 +113,7 @@ describe('renderDashboardShell', () => { { route: '/a', score: 50, + categories: {}, issues: [ { id: 'seo/title-presence', From ff4366346add28e15ddcd791de7dbbb50b623fe9 Mon Sep 17 00:00:00 2001 From: oekazuma Date: Tue, 4 Aug 2026 20:39:40 +0900 Subject: [PATCH 11/12] docs: the 37 inventory is the naive registry sum, not a reachable one MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The caution presented a key observing performance at 37 as a case that can occur. It cannot: a key is either a route id or a file path, so no key observes both of a category's scopes, and performance contributes 28 or 9. 37 is precisely what a reader gets by summing the registry — which is the trap the caution is about, so saying so makes the example do its job instead of implying a state that does not exist. --- .../2026-08-04-route-category-scores-design.md | 18 +++++++++++++----- 1 file changed, 13 insertions(+), 5 deletions(-) diff --git a/docs/superpowers/specs/2026-08-04-route-category-scores-design.md b/docs/superpowers/specs/2026-08-04-route-category-scores-design.md index 8981cf55c..97502a9a3 100644 --- a/docs/superpowers/specs/2026-08-04-route-category-scores-design.md +++ b/docs/superpowers/specs/2026-08-04-route-category-scores-design.md @@ -139,11 +139,19 @@ real and two of them are ordinary: finding the first had silently skipped part of the space; the existence of such keys is what matters here, and a number this design cannot reproduce is the kind of figure it has already got wrong twice. - Two cautions for anyone who recomputes: the coincidence depends on the **observed** inventory, not the - category's nominal one — a key observing `performance` at 37 (its route and component pairs together) can - coincide where the same failures at its component-only 9 give 56.43 against 63.69 — and "coincide" means - equal as exact rationals. As doubles the two can sit one ulp apart, 63.69047619047619 against - 63.69047619047618, which is why the agreement metric throughout this section is the displayed value. + Two cautions for anyone who recomputes. + + **The inventory is the one observed on the key, never the category's nominal total.** `performance`'s + non-project rules sum to 37 — 28 route-scoped plus 9 component-scoped — and 37 is exactly the number a + reader gets by adding up the registry. It is also unreachable: a key is either a route id or a file path, + so a single key never observes both of a category's scopes, and `performance` contributes 28 or 9, never 37. Summing the registry produces a coincidence at 37 that looks like confirmation, while the reachable + component-only 9 gives 56.43 against 63.69 — no coincidence at all. That trap is why this caution exists, + and it is load-bearing for the implementation too: `computeScore` sums `inventory.get(p)` over the pairs + observed on the key. + + **"Coincide" means equal as exact rationals.** As doubles the two values can sit one ulp apart — + 63.69047619047619 against 63.69047619047618 — which is why the agreement metric throughout this section is + the displayed value rather than the raw one. So the user-facing wording stays **not guaranteed** in both directions, and the reason is now a rule rather than a frequency. From d673093b9b81efe0aa16e2a7d64b6e5f2eb7b3b1 Mon Sep 17 00:00:00 2001 From: oekazuma Date: Wed, 5 Aug 2026 09:42:51 +0900 Subject: [PATCH 12/12] docs: stop claiming the category scores do not average to the route score MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The design spent four review passes arriving at "not guaranteed", and both the code comment and the Japanese guide then restated it as a certainty — the comment flatly, and the Japanese with のみ, which makes equal ratios necessary as well as sufficient. The English said "whenever ... and can differ otherwise" and was right, so the two languages had also drifted apart. --- docs/src/content/docs/ja/guides/(reporting)/reporters.md | 2 +- packages/core/src/reporter/json.ts | 4 ++-- 2 files changed, 3 insertions(+), 3 deletions(-) diff --git a/docs/src/content/docs/ja/guides/(reporting)/reporters.md b/docs/src/content/docs/ja/guides/(reporting)/reporters.md index d7bc6af0f..a8c6542df 100644 --- a/docs/src/content/docs/ja/guides/(reporting)/reporters.md +++ b/docs/src/content/docs/ja/guides/(reporting)/reporters.md @@ -82,7 +82,7 @@ svelte-vitals --reporter json `line`・`docsUrl`・`fix` はルールが提供した場合のみ、`location` はファイルに紐づく検出の場合のみ現れます。`issues` に並ぶのは**失敗した検出のみ**です。合格したチェックは `summary.passed` に数として計上されますが、一覧には出ません。失敗が1件も無いルートも `routes` には現れ、`issues` が空配列のまま自分のスコアを持ちます。 -`categories` にはそのルートで実際に結果が出たカテゴリだけが並びます。あるカテゴリが無いのは「ここでは測定していない」という意味であり、「ここは満点だった」という意味ではありません。この値がルート自身の `score` の平均になっているとは、どちらの向きにも**保証されません**。`score` はそのルートが測定対象とした全体に対する一つの割合であるのに対し、各カテゴリのスコアはそのカテゴリ自身の一覧に対する割合だからです。ルート上のすべてのカテゴリが同じ割合を示すとき(検出結果が無いルートも含む)のみ両者は一致し、それ以外では数ポイントずれることがあります。 +`categories` にはそのルートで実際に結果が出たカテゴリだけが並びます。あるカテゴリが無いのは「ここでは測定していない」という意味であり、「ここは満点だった」という意味ではありません。この値がルート自身の `score` の平均になっているとは、どちらの向きにも**保証されません**。`score` はそのルートが測定対象とした全体に対する一つの割合であるのに対し、各カテゴリのスコアはそのカテゴリ自身の一覧に対する割合だからです。ルート上のすべてのカテゴリが同じ割合を示すとき(検出結果が無いルートも含む)は両者が一致し、それ以外では数ポイントずれることがあります。 `rules` は、レポートの他の部分では答えられない問い、**そのルールがそもそも実行されたかどうか**に答えます。`issues` に載るのは失敗した検出結果だけなので、何も検出しなかったルールはそこに痕跡を残しません。トップレベルで無効化した(`--ignore`、`--rules`、`--category`、または設定ファイルの `rules: { id: 'off' }`)ルールも同じく現れません。判定には `rules` を見ます。存在すれば実行された、存在しなければトップレベルで除外された、という意味です。ただし一つだけ例外があり、それは次に述べます。 diff --git a/packages/core/src/reporter/json.ts b/packages/core/src/reporter/json.ts index 7c818cfe5..1768fcdea 100644 --- a/packages/core/src/reporter/json.ts +++ b/packages/core/src/reporter/json.ts @@ -84,8 +84,8 @@ export function buildJsonReport( .map(({ route, results: rs }) => ({ route, score: computeScore(rs, config, { applyCriticalCap: false }).score, - // Per category, scored against that category's own inventory — so this does not average to `score`, - // which is one ratio over the union of the pairs the route touched. + // Per category, scored against that category's own inventory, so this is not guaranteed to average to + // `score` — which is one ratio over the union of the pairs the route touched. categories: Object.fromEntries( Object.entries(scoresByCategory(rs, config, { applyCriticalCap: false })).map(([cat, sr]) => [cat, sr!.score]) ),