Skip to content

Bench: Add the docgen perf methodology note, the per-engine harness and the shared bench plumbing - #35573

Closed
valentinpalkovic wants to merge 15 commits into
nextfrom
valentin/sb11-story-1-10-perf-methodology
Closed

valentinpalkovic wants to merge 15 commits into
nextfrom
valentin/sb11-story-1-10-perf-methodology

Conversation

@valentinpalkovic

@valentinpalkovic valentinpalkovic commented Jul 24, 2026 •

Copy link
Copy Markdown
Contributor

What I did

This PR carries the measurement contract for the docgen perf suite, the per-engine harness that implements it, and the shared plumbing both docgen bench suites now sit on.

The methodology note

scripts/bench/PERF-METHODOLOGY.md is the written agreement on what we measure and how.
Without it the follow-up baseline work measures the wrong things, or produces numbers that cannot be compared from run to run.

It pins the five gated metrics (cold extraction, warm extraction, whole-project scan, peak memory, leak detection over a save series).
It pins the determinism rules per metric: fixed synthetic projects, warmup, median-of-N for latency, series statistics for memory, and relative comparisons on one machine only.
It pins the budget shape — timing budgets are ratios, never absolute milliseconds, while memory budgets stay absolute with generous headroom.
It pins the CI tiers: per-PR report-only, gating only on the daily tier.

The three user-perceived metrics — dev-server startup delta, time-to-Controls-populated, and docs-page props-table render time — join as report-only extensions of the sandbox bench task and never gate.

The note also records where today's reality falls short of the target.
Generated sandboxes do not read the docgen-server flag yet.
The daily tier is only triggered by the ci:daily label; no scheduled pipeline exists in the repo, so gating on it currently means gating when someone applies a label.

The per-engine harness

yarn bench:docgen-perf (from scripts/) spawns one fresh child process per measurement.
Cold medians come from N=5 spawns, with N pinned in config.ts and recorded in the results JSON — numbers taken at different N are not comparable.
Warm latency and the memory metrics read one repetition's save series, and the one they read is the repetition whose cold sample is the median.
Repetition 1 pays for a cold module graph and a cold page cache and measures several times slower than the rest, so reading it would bias four of the five metrics.
An engine that fails part-way through its repetitions is reported as failed, never as measured at an N it did not reach.
--quick runs a smoke profile whose results are marked non-comparable.

Each framework gets a legacy-vs-new control pair, measured in the same invocation with spawn order alternating across repetitions:

  • React — react-docgen against ComponentMetaManager. react-docgen-typescript is measurable behind a flag but carries no budget row.
  • Vue — vue-docgen-api, still the default plugin, against vue-component-meta, the opt-in successor. Both run over the same generated project.

vue-component-meta runs over generated monorepo-shaped projects, because its failure mode is workspace shape rather than component count: a packages/* layout with per-package tsconfigs, cross-package prop-type chains with depth and fan-out levers, a fake heavy .d.ts library inside the generated tree, and a scenario that touches a widely-imported base type to measure warm re-extraction under wide checker invalidation.

Compodoc runs as the external CLI it is.
Cold and whole-project scan are the same full-project run, warm is a second run after touching one file, and peak memory is the child's RSS polled from outside.
@compodoc/compodoc is pinned exactly at 2.0.0 in scripts/package.json, the version the Angular docgen baselines already capture against (code/lib/docgen-harness/README.md), so both harnesses describe the same tool.
The engine skips with an explicit message if the binary is missing, which is what a partial install looks like.

Generated projects land in the sandbox scratch directory and are never checked in.
No CI wiring and no budget numbers here; that is the baseline story's job.

The shared plumbing

The two suites had grown six copies of the argument parser, four copies of the save loop, and two copies each of the memory sampler, the renderer-module loader and the Vue scenario setup.
The memory gate — the only one of the two that fails a build — imported from the new perf directory, which gates nothing.

Both suites now depend on scripts/bench/docgen-shared/ and neither depends on the other.

Engine children implement three methods — cold, applySave, reextract — and the shared series harness owns the stopwatch.
That matters more than the deduplication: which part of a save is timed is the measurement contract, and it used to be hand-placed in each child. applySave is never timed, reextract always is.

The orchestrator has no per-engine branches. An engine is a BenchEngine:

  • SeriesChildEngine spawns a harness child; aggregation and member counts are the same for every engine of that kind because they share the series harness.
  • CompodocEngine drives the CLI and holds the resolved binary in a field.

React and Vue are the same class with a different child file and different flags, which is data rather than behavior, so registry.ts is a table.
Adding an engine is one entry.
run.ts drops from 556 lines to 195, and the aggregation, ordering, ratio and rendering rules move into modules tests can reach.

Two caveats a reviewer should not miss

Both control ratios need reading with care, and the note says so next to the claim.

The React warm ratio is not a like-for-like save.
Both sides re-extract one changed component, which compares the engines on equal work.
Only the new engine has a production path shaped that way — buildDocgenPayload hardcodes react-component-meta, so the single-component docgen worker never runs a legacy parser.
The legacy parsers are reachable only through generator.ts, which invalidates every cache and re-extracts every component on each manifest build.
A real legacy save therefore costs the whole project, not one file.
react-legacy.ts --scope all measures that shape.

The Vue ratio measures resolution depth as much as speed.
vue-docgen-api does not resolve tsconfig paths aliases, and the Vite plugin calls parse(id) without an alias map, so on these generated projects it documents a fraction of what vue-component-meta documents — and nothing at all on the save it is timed on.
Both engines report their documented-member count, and the suite prints those counts beside every ratio, warm as well as cold, marking the pair not like-for-like when they disagree.
A ratio marked not like-for-like must not become a budget.

Smaller fixes

  • Member counts are read from the same repetition the warm and memory figures come from, rather than repetition 1 — the one the suite otherwise avoids on purpose.
  • vue-docgen-api was imported by a scripts/ harness but declared only in code/frameworks/vue3-vite. It resolved by hoisting. Now declared, at the range the framework uses.
  • The shared argument parser rejects what the old one accepted: --saves --json out.json used to parse as saves="--json", become NaN, and reach the generator as a silently wrong project size.

One finding worth flagging

The new entry points run on native Node type stripping, not jiti.
Under --import jiti/register, react-docgen's browserslist dependency fails its JSON data require (jsReleases.map is not a function), so the legacy engine is not measurable under the jiti loader.
Only the reused memory-harness.ts child keeps its jiti flags, and a test now holds that split in place.

Strip-only mode also rejects TypeScript parameter properties, which is why Args assigns its field in the constructor body.

Not addressed here

.oxfmtrc.json ignores a bare bench, which exempts all of scripts/bench from format checks.
Narrowing that pattern would reformat files outside this PR, so it is left as a deliberate decision for someone else to make.

Checklist for Contributors

Testing

The changes in this PR are covered in the following automated tests:

  • stories
  • unit tests
  • integration tests
  • end-to-end tests

72 unit tests over the rules the note relies on: the median-repetition choice, the pinned-N check, the pair alternation, the ratio guards and their like-for-like warning, the argument parser, and that both sides of a control pair stay registered and run by default.

Manual testing

From scripts/: yarn bench:docgen-perf --quick runs the whole suite in a few minutes.
All five engines measure: react-legacy, react-osa, vue-docgen-api and vue-component-meta across three Vue scenarios, plus compodoc.
Drop --quick for the pinned N=5 profile, which takes considerably longer.

yarn bench:docgen-memory verifies the existing gate. Run it on Node 22.22.3, the version CI uses — on Node 24 the live positive control OOMs under the 1536MB cap. That is not from this PR: the pre-refactor harness OOMs identically on the same config and the same machine. Somebody should confirm a pass on 22.22.3 before this merges.

The memory harness still writes every field gate.ts reads, and its budgeted config passes with unchanged behavior.

For the note itself, the review is reading it: please check the claims against the referenced files and flag anything that does not match reality.

Documentation

  • Add or update documentation reflecting your changes
  • If you are deprecating/removing a feature, make sure to update
    MIGRATION.MD

Checklist for Maintainers

  • When this PR is ready for testing, make sure to add ci:normal, ci:merged or ci:daily GH label to it to run a specific set of sandboxes. The particular set of sandboxes can be found in code/lib/cli-storybook/src/sandbox-templates.ts

  • Declare whether manual QA will be needed for this PR during the next release, through qa:needed or qa:skip

  • Make sure this PR contains one of the labels below:

    Available labels
    • bug: Internal changes that fixes incorrect behavior.
    • maintenance: User-facing maintenance tasks.
    • dependencies: Upgrading (sometimes downgrading) dependencies.
    • build: Internal-facing build tooling & test updates. Will not show up in release changelog.
    • cleanup: Minor cleanup style change. Will not show up in release changelog.
    • documentation: Documentation only changes. Will not show up in release changelog.
    • feature request: Introducing a new feature.
    • BREAKING CHANGE: Changes that break compatibility in some way with current major version.
    • other: Changes that don't fit in the above categories.

🦋 Canary release

This PR does not have a canary release associated. You can request a canary release of this pull request by mentioning the @storybookjs/core team here.

core team members can create a canary release here or locally with gh workflow run --repo storybookjs/storybook publish.yml --field pr=<PR_NUMBER>

Fixes the metric set, the determinism method, the budget shape, and the
CI tiers for the per-engine docgen perf suite, and adds the budgets
table skeleton. The three user-perceived metrics join as report-only
extensions of the sandbox bench task; the gated suite stays Node-only.
Resolves the ambiguities a Sol + Fable review pass confirmed: how
median-of-N samples map onto fresh processes per metric, what the
timing ratio divides, what Compodoc's per-save metrics mean for a
one-shot CLI, and what a negative control must name. Corrects two
claims the code contradicted: retained growth is a last-minus-baseline
delta, not a least-squares fit, and benchAutodocs has no real wait on
the description text. Drops a duplicated sentence, plain-words the CI
tier wording, and realigns the table.
@valentinpalkovic valentinpalkovic added build Internal-facing build tooling & test updates ci:normal Run our default set of CI jobs (choose this for most PRs). qa:skip Pull Requests that do not need any QA. (e.g. documentation) labels Jul 24, 2026
Implements the harness half of the perf methodology note. A new
orchestrator (yarn bench:docgen-perf) measures the five floor metrics
per engine in plain Node: one fresh child process per measurement,
cold medians from N spawns with the pinned N recorded in the results,
warm as the median of the first repetition's save series.

react-docgen and ComponentMetaManager form the calibration pair: both
run in one invocation, their spawn order alternates per repetition,
and the results carry the legacy-vs-osa ratio of medians.
react-docgen-typescript is measurable behind a flag but carries no
budget row. vue-component-meta runs over generated monorepo-shaped
projects: workspace packages with per-package tsconfigs and paths
aliases, cross-package prop-type chains, a fake heavy .d.ts library in
the generated tree, and a scenario that touches a widely-imported base
type to measure warm extraction under wide checker invalidation.
Compodoc runs as the external CLI it is, with peak RSS polled from
outside; it skips with a message when no binary is installed, since
compodoc is a user-project dependency.

The memory harness now persists the cold and per-save durations it
already computed, and the gate reads its budget values from the shared
budgets.ts. Gate behavior is unchanged (verified on Node 22.22.3; its
hand-tuned OOM pair is margin-sensitive on other Node majors with or
without this change). The new entry points run on native Node type
stripping because react-docgen's browserslist data require breaks
under the jiti loader.
@valentinpalkovic valentinpalkovic changed the title Bench: Add the docgen perf methodology note Bench: Add the docgen perf methodology note and per-engine harness Jul 24, 2026
Review of the harness turned up four ways it reported numbers that did not
mean what they claimed.

Warm latency and both memory metrics read repetition 1's save series. That
repetition pays for a cold module graph and a cold page cache and runs several
times slower than the rest: react-osa cold measured 2199ms on rep 1 and 552ms
on rep 2. Only cold latency was protected, by its median. The series now comes
from the repetition whose cold sample is the median.

An engine that failed part-way through its repetitions still landed in the
results as `measured`, carrying fewer samples than the pinned N the same file
recorded. It now stays failed.

The compodoc RSS poller called `.trim()` on `spawnSync().stdout`, which is
undefined when `ps` is missing. That throws from a timer callback, where no
per-engine catch can reach it, so one missing binary killed the whole run.

`--json` took the next argv entry without checking it, so `--json --quick`
wrote a file named `--quick`. Flag values are now validated.

The legacy React warm sample re-extracts one component, but the legacy parsers
are reachable only through generator.ts, which invalidates every cache and
re-extracts every component; buildDocgen.ts hardcodes react-component-meta, so
the single-file worker never runs them. The warm ratio therefore understates
legacy's real per-save cost by roughly the component count. react-legacy.ts
gains `--scope all` for the production-shaped number, and the note says what
the equal-work ratio does and does not describe.

Also: the note claimed as missing several things this PR ships, and left Vue's
entry point unpinned after Vue's harness landed. Retained/transient derivation
was written out three times and now lives in stats.ts. `vue` was resolved by
hoisting alone and is now declared. Dropped three fields nothing reads.

One thing left alone: oxfmt's ignore list contains a bare `bench`, which
exempts all of scripts/bench from format checks. Narrowing it would reformat
files outside this PR.
Vue had no legacy baseline. The budget shape asks every engine for a
legacy-vs-new ratio measured in the same run, React had one, and Vue did not
even though the second implementation was sitting there: vue-docgen-api is
still the default plugin, and vue-component-meta is the opt-in successor.

The new engine calls parse() per .vue file the way the Vite plugin does. The
parser re-reads from disk and keeps no cache, so a save costs one parse of one
file and the per-save sample needs no invalidation step. Both Vue engines run
over the same generated project and alternate spawn order across repetitions,
so their medians divide into a ratio.

The ratio needs reading with care, which is why both engines now report how
many members their cold pass documented. vue-docgen-api does not resolve
tsconfig paths aliases, and the plugin calls parse(id) without an alias map, so
on these projects it documents a fraction of what vue-component-meta does. At
the pinned profile it finished a cold pass in 49-71ms against 437-456ms while
documenting 50-80 members against 320-350. The suite prints both counts beside
the ratio and marks them not like-for-like when they disagree, so nobody reads
a 0.11 as a speed win over identical work.
@storybook-app-bot

storybook-app-bot Bot commented Jul 24, 2026 •

Copy link
Copy Markdown
Contributor

Package Benchmarks

Commit: dd13184, ran on 27 July 2026 at 09:21:12 UTC

No significant changes detected, all good. 👏

The compodoc engine was in the default engine list but always skipped, because
nothing in the repo declared the binary. Measuring Angular took a manual
install, so in practice nobody measured it.

scripts/package.json now pins @compodoc/compodoc at exactly 2.0.0 - the version
the Angular docgen baselines already capture against, per
code/lib/docgen-harness/README.md. Sharing one version keeps the perf suite and
the baseline harness describing the same tool. The pin is exact rather than a
range because compodoc's cost moves across versions, and a caret would drift
the numbers without anyone choosing to.

The resolved version is recorded with the results, read by walking up from the
binary that actually ran rather than from a fixed path, since the package
hoists to different node_modules depending on the install.

The skip path stays. It is still what a partial install looks like, so its
message now points at yarn install instead of calling compodoc a user-project
dependency.

Measured at the smoke profile: cold 574ms, warm 710ms, scan 574ms, peak 222MB.
Warm sits near cold, which is what the note predicts for a one-shot CLI.
The Vue legacy engine harness imports vue-docgen-api, but only
code/frameworks/vue3-vite declared it. It resolved by hoisting, which
would break the moment the framework dropped it or the install layout
changed. Declare it in scripts at the same range the framework uses.
The two docgen bench suites had grown six copies of the argument
parser, four copies of the save loop, and two copies each of the memory
sampler, the renderer-module loader and the Vue scenario setup. The
memory gate, which is the only one that fails a build, imported from
the new perf directory, which gates nothing.

Both suites now depend on scripts/bench/docgen-shared and neither
depends on the other. Engine children implement cold, applySave and
reextract; the shared series harness owns the stopwatch, so which part
of a save is measured is fixed in one place rather than hand-placed in
each child.

The orchestrator loses its per-engine branches. An engine is a
BenchEngine: SeriesChildEngine spawns a harness child, CompodocEngine
drives the CLI and holds the resolved binary. React and Vue are the
same class with a different child and different flags. run.ts drops
from 556 lines to 195, and the aggregation, ordering, ratio and
rendering rules move into modules that tests can reach.

Two reporting fixes come with that. Member counts now come from the
same repetition the warm and memory figures do, instead of repetition
1, which the suite otherwise avoids on purpose. And a warm ratio now
carries the same not-like-for-like warning the cold one did: on the
generated projects the legacy Vue parser documents nothing on the save
it is timed on, so a bare warm ratio read as a clean speed win.

The shared argument parser also rejects what the old one accepted.
"--saves --json out.json" used to parse as saves="--json", become NaN,
and reach the generator as a silently wrong project size.

Tests go from 12 to 72, covering the rules the note relies on: the
median-repetition choice, the pinned-N check, the pair alternation, the
ratio guards, and that both sides of a control pair stay registered and
run by default.

No measurement behavior changes. yarn bench:docgen-perf --quick reports
the same numbers, and the memory harness still writes every field the
gate reads.
@valentinpalkovic valentinpalkovic changed the title Bench: Add the docgen perf methodology note and per-engine harness Bench: Add the docgen perf methodology note, the per-engine harness and the shared bench plumbing Jul 28, 2026
The comments explained too much of what the code already says, and
several module docblocks argued for the design rather than describing
how to use it. Facts were repeated up to four times: the reason a Vue
ratio is not like-for-like appeared in two module blocks and two
functions.

Every comment that restates an identifier, narrates the code, or sells
the architecture is gone. What stays is the part a reader cannot infer:
that both legacy React parsers expose only global invalidation, so a
save that skips it measures a cache hit; that vue-component-meta never
re-stats a file, so a disk write without updateFile measures a stale
cache; that this child cannot run under jiti because react-docgen's
browserslist dependency fails its JSON data require; that ps is absent
on Windows and would throw from a timer callback no catch can reach;
that strip-only TypeScript rejects parameter properties. Each of those
now appears once, at the site a reader reaches first.

CLI invocation examples are kept everywhere they appeared - they are
the API usage.

83 comment lines net, comments only. No code, type, signature or test
assertion changed, verified by checking that every added and removed
line in the diff is a comment. yarn bench:docgen-perf --quick still
reports the same numbers.
The bench harnesses had a hand-rolled argument parser. node:util's
parseArgs already does the job, and scripts/eval/eval.ts already pairs
it with Zod, so this follows a convention that was there before.

It also closes a hole. The hand-rolled parser only validated flags it
was asked about, so a typo passed straight through: `--component 5`,
with the s missing, was ignored and the harness ran the default 300
components while the operator believed it ran 5. A benchmark that
reports a confident number for the wrong project size is the failure
this suite exists to prevent. parseArgs in strict mode rejects any
undeclared flag, which means every harness now has to declare the flags
the orchestrator sends it.

Zod covers what parseArgs does not: coercing the string values to
integers and rejecting fractions, negatives and NaN, and checking the
enum flags.

Options are typed by the harness interfaces rather than inferred from
the schemas. Under this TypeScript setup zod 3 types a `.default()` key
as optional even on parsed output, so inference would make every
defaulted option optional at the call site; args.test.ts covers the gap
by asserting a no-argument parse populates every field.

Verified: yarn bench:docgen-perf --quick measures all five engines with
no strict-mode mismatch, and the memory harness accepts both gate
configs, including the live one with --recycle, --heavy and
--no-force-gc.
parseHarnessOptions now camel-cases kebab flags, so each harness spells out only
the real renames such as --out onto outDir. The memory harness drops from fifteen
mapped keys to six, and a typo in one can no longer leave a default silently in
place.

forceGc and withNodeModules lose their double negative: the inversion sits where
the flag is renamed, not inside the Zod schema.

SeriesResult extends SeriesSummary instead of restating its four memory fields,
and runSeries spreads summarizeSeries. summarizeSeries walks the GC-sampled saves
once rather than twice.

ratioFor builds one object literal, with a likeForLike helper beside the existing
medianRatio. One outputTail helper replaces three copies of the last-N-lines
formatting. The react-legacy parser wiring moves into loadParser. run.ts drops
optional calls on methods BenchEngine always defines.

No behavior change. The bench unit tests pass, and the React, Vue and memory
harnesses were run end to end.
Camel-casing the flag names gave `maxRetainedGrowth`, but the schema key is
`maxRetainedGrowthMb`, so the value never landed and Zod fell back to the 400MB
default. parseArgs accepted the flag, nothing warned, and the harness compared
against a threshold nobody asked for. It needs the same explicit rename as
--out and --json.

args.test.ts now carries a flag with the same shape, so removing the rename
fails the suite instead of going quiet.

Also finish the react-renderer-module rename, which left generate-project.ts
importing a deleted file, and correct the SaveSample comment: saves count
from 1, not 0.
The engine guessed at two node_modules/.bin paths by counting `../` from its
own directory, then fell back to `which`. The version came from a separate walk
up to five directories from the resolved binary. One require.resolve of the
package.json replaces all of it and hands back the version for free, matching
how the React and Vue generators already find their packages.

Spawning `node <cli>` rather than the .bin shim drops the dependency on the
exec bit, and gives Windows a path that works at all: the shim is named
compodoc.cmd there, so neither candidate matched and `which` does not exist.

Dropping the PATH fallback is deliberate. @compodoc/compodoc is pinned in
scripts/package.json so the numbers describe one known version, and a global
compodoc of some other version would move them without anyone deciding to.
Missing package now means skipped, with a reason that names the fix.
The Angular row reported timings and nothing else, so there was no way to read
them. Compodoc writes documentation.json on every run and the engine never
opened it.

It now reports documented members, counting a component's own four member
arrays across components and directives. Classes are left out on purpose:
compodoc copies an ancestor's members into every descendant, so counting the
base class again would double every inherited member.

A member count alone would still mislead here. Compodoc never resolves a named
type, so `@Input() kind: ButtonKind` is stored as the name and an `@Output()`
is stored as `EventEmitter` with the payload dropped, while the same union
written inline is stored in full. Chain a type through twenty files and it
documents every member at the same speed, having looked through none of them.
So it also reports how many documented members carry a type it never resolved,
which is what keeps a future Angular pair from reading equal counts as equal
work.

Measured on the quick profile: 90 members cold, 91 warm, 10 of them opaque -
one EventEmitter per component.

The methodology note gains that reasoning, plus the signals finding: compodoc
matches input() and output() against the initializer's source text, so it
documents them at the same speed and count as decorators. The cost is fidelity,
not time, so it belongs in the harness fixtures rather than a bench scenario.
@valentinpalkovic

valentinpalkovic commented Jul 28, 2026 •

Copy link
Copy Markdown
Contributor Author

Closing this in favour of a stack of four smaller PRs. 41 files in one go was too much to ask of a reviewer, and the pieces are genuinely different kinds of change — a methodology note reads nothing like a lockfile bump.

Take them in order; each is based on the one below it:

  1. Bench: Add the docgen perf methodology note #35630 — the methodology note. Prose only. What we measure, and why a ratio is the only number we are allowed to gate on.
  2. Bench: Add the shared docgen bench plumbing and move the memory harness onto it #35631 — the shared bench plumbing, and the existing memory harness moving onto it. Two commits, kept together because the second is what justifies the first.
  3. Bench: Add the docgen perf engine framework and the compodoc engine #35636 — the perf engine framework and the Compodoc engine, plus the @compodoc/compodoc pin.
  4. Bench: Make the docgen perf suite runnable #35634 — the orchestration that makes it runnable, and the React and Vue engines.

The split follows the import graph, not the directory names. Splitting docgen-perf along its engines/ boundary does not typecheck standalone in either order, since registry.ts imports the Compodoc engine and the Compodoc engine imports back into the top-level framework files. Grouping by dependency layer instead is the smallest cut that works.

Each part declares only what it ships. The measurement profile and the EngineId list grow one engine at a time, alongside the code that reads them, rather than arriving complete in the first PR.

Every branch was verified on its own, not just the tip: yarn check, yarn lint, yarn install --immutable and the unit tests pass at each step (23 → 43 → 77 tests). yarn.lock is regenerated per branch with a real install, so the Compodoc tree lands in part 3 and the Vue pins in part 4.

One extra PR came out of this work and does not belong to the stack: #35628 records two Compodoc signal gaps as docgen-harness fixtures. It is independent and can be reviewed on its own.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

build Internal-facing build tooling & test updates ci:normal Run our default set of CI jobs (choose this for most PRs). qa:skip Pull Requests that do not need any QA. (e.g. documentation)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant