Repository navigation
Bench: Add the docgen perf methodology note, the per-engine harness and the shared bench plumbing - #35573
Bench: Add the docgen perf methodology note, the per-engine harness and the shared bench plumbing#35573valentinpalkovic wants to merge 15 commits into
Conversation
Fixes the metric set, the determinism method, the budget shape, and the CI tiers for the per-engine docgen perf suite, and adds the budgets table skeleton. The three user-perceived metrics join as report-only extensions of the sandbox bench task; the gated suite stays Node-only.
Resolves the ambiguities a Sol + Fable review pass confirmed: how median-of-N samples map onto fresh processes per metric, what the timing ratio divides, what Compodoc's per-save metrics mean for a one-shot CLI, and what a negative control must name. Corrects two claims the code contradicted: retained growth is a last-minus-baseline delta, not a least-squares fit, and benchAutodocs has no real wait on the description text. Drops a duplicated sentence, plain-words the CI tier wording, and realigns the table.
Implements the harness half of the perf methodology note. A new orchestrator (yarn bench:docgen-perf) measures the five floor metrics per engine in plain Node: one fresh child process per measurement, cold medians from N spawns with the pinned N recorded in the results, warm as the median of the first repetition's save series. react-docgen and ComponentMetaManager form the calibration pair: both run in one invocation, their spawn order alternates per repetition, and the results carry the legacy-vs-osa ratio of medians. react-docgen-typescript is measurable behind a flag but carries no budget row. vue-component-meta runs over generated monorepo-shaped projects: workspace packages with per-package tsconfigs and paths aliases, cross-package prop-type chains, a fake heavy .d.ts library in the generated tree, and a scenario that touches a widely-imported base type to measure warm extraction under wide checker invalidation. Compodoc runs as the external CLI it is, with peak RSS polled from outside; it skips with a message when no binary is installed, since compodoc is a user-project dependency. The memory harness now persists the cold and per-save durations it already computed, and the gate reads its budget values from the shared budgets.ts. Gate behavior is unchanged (verified on Node 22.22.3; its hand-tuned OOM pair is margin-sensitive on other Node majors with or without this change). The new entry points run on native Node type stripping because react-docgen's browserslist data require breaks under the jiti loader.
Review of the harness turned up four ways it reported numbers that did not mean what they claimed. Warm latency and both memory metrics read repetition 1's save series. That repetition pays for a cold module graph and a cold page cache and runs several times slower than the rest: react-osa cold measured 2199ms on rep 1 and 552ms on rep 2. Only cold latency was protected, by its median. The series now comes from the repetition whose cold sample is the median. An engine that failed part-way through its repetitions still landed in the results as `measured`, carrying fewer samples than the pinned N the same file recorded. It now stays failed. The compodoc RSS poller called `.trim()` on `spawnSync().stdout`, which is undefined when `ps` is missing. That throws from a timer callback, where no per-engine catch can reach it, so one missing binary killed the whole run. `--json` took the next argv entry without checking it, so `--json --quick` wrote a file named `--quick`. Flag values are now validated. The legacy React warm sample re-extracts one component, but the legacy parsers are reachable only through generator.ts, which invalidates every cache and re-extracts every component; buildDocgen.ts hardcodes react-component-meta, so the single-file worker never runs them. The warm ratio therefore understates legacy's real per-save cost by roughly the component count. react-legacy.ts gains `--scope all` for the production-shaped number, and the note says what the equal-work ratio does and does not describe. Also: the note claimed as missing several things this PR ships, and left Vue's entry point unpinned after Vue's harness landed. Retained/transient derivation was written out three times and now lives in stats.ts. `vue` was resolved by hoisting alone and is now declared. Dropped three fields nothing reads. One thing left alone: oxfmt's ignore list contains a bare `bench`, which exempts all of scripts/bench from format checks. Narrowing it would reformat files outside this PR.
Vue had no legacy baseline. The budget shape asks every engine for a legacy-vs-new ratio measured in the same run, React had one, and Vue did not even though the second implementation was sitting there: vue-docgen-api is still the default plugin, and vue-component-meta is the opt-in successor. The new engine calls parse() per .vue file the way the Vite plugin does. The parser re-reads from disk and keeps no cache, so a save costs one parse of one file and the per-save sample needs no invalidation step. Both Vue engines run over the same generated project and alternate spawn order across repetitions, so their medians divide into a ratio. The ratio needs reading with care, which is why both engines now report how many members their cold pass documented. vue-docgen-api does not resolve tsconfig paths aliases, and the plugin calls parse(id) without an alias map, so on these projects it documents a fraction of what vue-component-meta does. At the pinned profile it finished a cold pass in 49-71ms against 437-456ms while documenting 50-80 members against 320-350. The suite prints both counts beside the ratio and marks them not like-for-like when they disagree, so nobody reads a 0.11 as a speed win over identical work.
Package BenchmarksCommit: No significant changes detected, all good. 👏 |
The compodoc engine was in the default engine list but always skipped, because nothing in the repo declared the binary. Measuring Angular took a manual install, so in practice nobody measured it. scripts/package.json now pins @compodoc/compodoc at exactly 2.0.0 - the version the Angular docgen baselines already capture against, per code/lib/docgen-harness/README.md. Sharing one version keeps the perf suite and the baseline harness describing the same tool. The pin is exact rather than a range because compodoc's cost moves across versions, and a caret would drift the numbers without anyone choosing to. The resolved version is recorded with the results, read by walking up from the binary that actually ran rather than from a fixed path, since the package hoists to different node_modules depending on the install. The skip path stays. It is still what a partial install looks like, so its message now points at yarn install instead of calling compodoc a user-project dependency. Measured at the smoke profile: cold 574ms, warm 710ms, scan 574ms, peak 222MB. Warm sits near cold, which is what the note predicts for a one-shot CLI.
The Vue legacy engine harness imports vue-docgen-api, but only code/frameworks/vue3-vite declared it. It resolved by hoisting, which would break the moment the framework dropped it or the install layout changed. Declare it in scripts at the same range the framework uses.
The two docgen bench suites had grown six copies of the argument parser, four copies of the save loop, and two copies each of the memory sampler, the renderer-module loader and the Vue scenario setup. The memory gate, which is the only one that fails a build, imported from the new perf directory, which gates nothing. Both suites now depend on scripts/bench/docgen-shared and neither depends on the other. Engine children implement cold, applySave and reextract; the shared series harness owns the stopwatch, so which part of a save is measured is fixed in one place rather than hand-placed in each child. The orchestrator loses its per-engine branches. An engine is a BenchEngine: SeriesChildEngine spawns a harness child, CompodocEngine drives the CLI and holds the resolved binary. React and Vue are the same class with a different child and different flags. run.ts drops from 556 lines to 195, and the aggregation, ordering, ratio and rendering rules move into modules that tests can reach. Two reporting fixes come with that. Member counts now come from the same repetition the warm and memory figures do, instead of repetition 1, which the suite otherwise avoids on purpose. And a warm ratio now carries the same not-like-for-like warning the cold one did: on the generated projects the legacy Vue parser documents nothing on the save it is timed on, so a bare warm ratio read as a clean speed win. The shared argument parser also rejects what the old one accepted. "--saves --json out.json" used to parse as saves="--json", become NaN, and reach the generator as a silently wrong project size. Tests go from 12 to 72, covering the rules the note relies on: the median-repetition choice, the pinned-N check, the pair alternation, the ratio guards, and that both sides of a control pair stay registered and run by default. No measurement behavior changes. yarn bench:docgen-perf --quick reports the same numbers, and the memory harness still writes every field the gate reads.
The comments explained too much of what the code already says, and several module docblocks argued for the design rather than describing how to use it. Facts were repeated up to four times: the reason a Vue ratio is not like-for-like appeared in two module blocks and two functions. Every comment that restates an identifier, narrates the code, or sells the architecture is gone. What stays is the part a reader cannot infer: that both legacy React parsers expose only global invalidation, so a save that skips it measures a cache hit; that vue-component-meta never re-stats a file, so a disk write without updateFile measures a stale cache; that this child cannot run under jiti because react-docgen's browserslist dependency fails its JSON data require; that ps is absent on Windows and would throw from a timer callback no catch can reach; that strip-only TypeScript rejects parameter properties. Each of those now appears once, at the site a reader reaches first. CLI invocation examples are kept everywhere they appeared - they are the API usage. 83 comment lines net, comments only. No code, type, signature or test assertion changed, verified by checking that every added and removed line in the diff is a comment. yarn bench:docgen-perf --quick still reports the same numbers.
The bench harnesses had a hand-rolled argument parser. node:util's parseArgs already does the job, and scripts/eval/eval.ts already pairs it with Zod, so this follows a convention that was there before. It also closes a hole. The hand-rolled parser only validated flags it was asked about, so a typo passed straight through: `--component 5`, with the s missing, was ignored and the harness ran the default 300 components while the operator believed it ran 5. A benchmark that reports a confident number for the wrong project size is the failure this suite exists to prevent. parseArgs in strict mode rejects any undeclared flag, which means every harness now has to declare the flags the orchestrator sends it. Zod covers what parseArgs does not: coercing the string values to integers and rejecting fractions, negatives and NaN, and checking the enum flags. Options are typed by the harness interfaces rather than inferred from the schemas. Under this TypeScript setup zod 3 types a `.default()` key as optional even on parsed output, so inference would make every defaulted option optional at the call site; args.test.ts covers the gap by asserting a no-argument parse populates every field. Verified: yarn bench:docgen-perf --quick measures all five engines with no strict-mode mismatch, and the memory harness accepts both gate configs, including the live one with --recycle, --heavy and --no-force-gc.
parseHarnessOptions now camel-cases kebab flags, so each harness spells out only the real renames such as --out onto outDir. The memory harness drops from fifteen mapped keys to six, and a typo in one can no longer leave a default silently in place. forceGc and withNodeModules lose their double negative: the inversion sits where the flag is renamed, not inside the Zod schema. SeriesResult extends SeriesSummary instead of restating its four memory fields, and runSeries spreads summarizeSeries. summarizeSeries walks the GC-sampled saves once rather than twice. ratioFor builds one object literal, with a likeForLike helper beside the existing medianRatio. One outputTail helper replaces three copies of the last-N-lines formatting. The react-legacy parser wiring moves into loadParser. run.ts drops optional calls on methods BenchEngine always defines. No behavior change. The bench unit tests pass, and the React, Vue and memory harnesses were run end to end.
Camel-casing the flag names gave `maxRetainedGrowth`, but the schema key is `maxRetainedGrowthMb`, so the value never landed and Zod fell back to the 400MB default. parseArgs accepted the flag, nothing warned, and the harness compared against a threshold nobody asked for. It needs the same explicit rename as --out and --json. args.test.ts now carries a flag with the same shape, so removing the rename fails the suite instead of going quiet. Also finish the react-renderer-module rename, which left generate-project.ts importing a deleted file, and correct the SaveSample comment: saves count from 1, not 0.
The engine guessed at two node_modules/.bin paths by counting `../` from its own directory, then fell back to `which`. The version came from a separate walk up to five directories from the resolved binary. One require.resolve of the package.json replaces all of it and hands back the version for free, matching how the React and Vue generators already find their packages. Spawning `node <cli>` rather than the .bin shim drops the dependency on the exec bit, and gives Windows a path that works at all: the shim is named compodoc.cmd there, so neither candidate matched and `which` does not exist. Dropping the PATH fallback is deliberate. @compodoc/compodoc is pinned in scripts/package.json so the numbers describe one known version, and a global compodoc of some other version would move them without anyone deciding to. Missing package now means skipped, with a reason that names the fix.
The Angular row reported timings and nothing else, so there was no way to read them. Compodoc writes documentation.json on every run and the engine never opened it. It now reports documented members, counting a component's own four member arrays across components and directives. Classes are left out on purpose: compodoc copies an ancestor's members into every descendant, so counting the base class again would double every inherited member. A member count alone would still mislead here. Compodoc never resolves a named type, so `@Input() kind: ButtonKind` is stored as the name and an `@Output()` is stored as `EventEmitter` with the payload dropped, while the same union written inline is stored in full. Chain a type through twenty files and it documents every member at the same speed, having looked through none of them. So it also reports how many documented members carry a type it never resolved, which is what keeps a future Angular pair from reading equal counts as equal work. Measured on the quick profile: 90 members cold, 91 warm, 10 of them opaque - one EventEmitter per component. The methodology note gains that reasoning, plus the signals finding: compodoc matches input() and output() against the initializer's source text, so it documents them at the same speed and count as decorators. The cost is fidelity, not time, so it belongs in the harness fixtures rather than a bench scenario.
|
Closing this in favour of a stack of four smaller PRs. 41 files in one go was too much to ask of a reviewer, and the pieces are genuinely different kinds of change — a methodology note reads nothing like a lockfile bump. Take them in order; each is based on the one below it:
The split follows the import graph, not the directory names. Splitting Each part declares only what it ships. The measurement profile and the Every branch was verified on its own, not just the tip: One extra PR came out of this work and does not belong to the stack: #35628 records two Compodoc signal gaps as |
What I did
This PR carries the measurement contract for the docgen perf suite, the per-engine harness that implements it, and the shared plumbing both docgen bench suites now sit on.
The methodology note
scripts/bench/PERF-METHODOLOGY.mdis the written agreement on what we measure and how.Without it the follow-up baseline work measures the wrong things, or produces numbers that cannot be compared from run to run.
It pins the five gated metrics (cold extraction, warm extraction, whole-project scan, peak memory, leak detection over a save series).
It pins the determinism rules per metric: fixed synthetic projects, warmup, median-of-N for latency, series statistics for memory, and relative comparisons on one machine only.
It pins the budget shape — timing budgets are ratios, never absolute milliseconds, while memory budgets stay absolute with generous headroom.
It pins the CI tiers: per-PR report-only, gating only on the daily tier.
The three user-perceived metrics — dev-server startup delta, time-to-Controls-populated, and docs-page props-table render time — join as report-only extensions of the sandbox bench task and never gate.
The note also records where today's reality falls short of the target.
Generated sandboxes do not read the docgen-server flag yet.
The daily tier is only triggered by the
ci:dailylabel; no scheduled pipeline exists in the repo, so gating on it currently means gating when someone applies a label.The per-engine harness
yarn bench:docgen-perf(fromscripts/) spawns one fresh child process per measurement.Cold medians come from N=5 spawns, with N pinned in
config.tsand recorded in the results JSON — numbers taken at different N are not comparable.Warm latency and the memory metrics read one repetition's save series, and the one they read is the repetition whose cold sample is the median.
Repetition 1 pays for a cold module graph and a cold page cache and measures several times slower than the rest, so reading it would bias four of the five metrics.
An engine that fails part-way through its repetitions is reported as failed, never as measured at an N it did not reach.
--quickruns a smoke profile whose results are marked non-comparable.Each framework gets a legacy-vs-new control pair, measured in the same invocation with spawn order alternating across repetitions:
react-docgenagainstComponentMetaManager.react-docgen-typescriptis measurable behind a flag but carries no budget row.vue-docgen-api, still the default plugin, againstvue-component-meta, the opt-in successor. Both run over the same generated project.vue-component-metaruns over generated monorepo-shaped projects, because its failure mode is workspace shape rather than component count: apackages/*layout with per-package tsconfigs, cross-package prop-type chains with depth and fan-out levers, a fake heavy.d.tslibrary inside the generated tree, and a scenario that touches a widely-imported base type to measure warm re-extraction under wide checker invalidation.Compodoc runs as the external CLI it is.
Cold and whole-project scan are the same full-project run, warm is a second run after touching one file, and peak memory is the child's RSS polled from outside.
@compodoc/compodocis pinned exactly at 2.0.0 inscripts/package.json, the version the Angular docgen baselines already capture against (code/lib/docgen-harness/README.md), so both harnesses describe the same tool.The engine skips with an explicit message if the binary is missing, which is what a partial install looks like.
Generated projects land in the sandbox scratch directory and are never checked in.
No CI wiring and no budget numbers here; that is the baseline story's job.
The shared plumbing
The two suites had grown six copies of the argument parser, four copies of the save loop, and two copies each of the memory sampler, the renderer-module loader and the Vue scenario setup.
The memory gate — the only one of the two that fails a build — imported from the new perf directory, which gates nothing.
Both suites now depend on
scripts/bench/docgen-shared/and neither depends on the other.Engine children implement three methods —
cold,applySave,reextract— and the shared series harness owns the stopwatch.That matters more than the deduplication: which part of a save is timed is the measurement contract, and it used to be hand-placed in each child.
applySaveis never timed,reextractalways is.The orchestrator has no per-engine branches. An engine is a
BenchEngine:SeriesChildEnginespawns a harness child; aggregation and member counts are the same for every engine of that kind because they share the series harness.CompodocEnginedrives the CLI and holds the resolved binary in a field.React and Vue are the same class with a different child file and different flags, which is data rather than behavior, so
registry.tsis a table.Adding an engine is one entry.
run.tsdrops from 556 lines to 195, and the aggregation, ordering, ratio and rendering rules move into modules tests can reach.Two caveats a reviewer should not miss
Both control ratios need reading with care, and the note says so next to the claim.
The React warm ratio is not a like-for-like save.
Both sides re-extract one changed component, which compares the engines on equal work.
Only the new engine has a production path shaped that way —
buildDocgenPayloadhardcodesreact-component-meta, so the single-component docgen worker never runs a legacy parser.The legacy parsers are reachable only through
generator.ts, which invalidates every cache and re-extracts every component on each manifest build.A real legacy save therefore costs the whole project, not one file.
react-legacy.ts --scope allmeasures that shape.The Vue ratio measures resolution depth as much as speed.
vue-docgen-apidoes not resolve tsconfigpathsaliases, and the Vite plugin callsparse(id)without an alias map, so on these generated projects it documents a fraction of whatvue-component-metadocuments — and nothing at all on the save it is timed on.Both engines report their documented-member count, and the suite prints those counts beside every ratio, warm as well as cold, marking the pair not like-for-like when they disagree.
A ratio marked not like-for-like must not become a budget.
Smaller fixes
vue-docgen-apiwas imported by ascripts/harness but declared only incode/frameworks/vue3-vite. It resolved by hoisting. Now declared, at the range the framework uses.--saves --json out.jsonused to parse assaves="--json", becomeNaN, and reach the generator as a silently wrong project size.One finding worth flagging
The new entry points run on native Node type stripping, not jiti.
Under
--import jiti/register, react-docgen's browserslist dependency fails its JSON data require (jsReleases.map is not a function), so the legacy engine is not measurable under the jiti loader.Only the reused
memory-harness.tschild keeps its jiti flags, and a test now holds that split in place.Strip-only mode also rejects TypeScript parameter properties, which is why
Argsassigns its field in the constructor body.Not addressed here
.oxfmtrc.jsonignores a barebench, which exempts all ofscripts/benchfrom format checks.Narrowing that pattern would reformat files outside this PR, so it is left as a deliberate decision for someone else to make.
Checklist for Contributors
Testing
The changes in this PR are covered in the following automated tests:
72 unit tests over the rules the note relies on: the median-repetition choice, the pinned-N check, the pair alternation, the ratio guards and their like-for-like warning, the argument parser, and that both sides of a control pair stay registered and run by default.
Manual testing
From
scripts/:yarn bench:docgen-perf --quickruns the whole suite in a few minutes.All five engines measure: react-legacy, react-osa, vue-docgen-api and vue-component-meta across three Vue scenarios, plus compodoc.
Drop
--quickfor the pinned N=5 profile, which takes considerably longer.yarn bench:docgen-memoryverifies the existing gate. Run it on Node 22.22.3, the version CI uses — on Node 24 the live positive control OOMs under the 1536MB cap. That is not from this PR: the pre-refactor harness OOMs identically on the same config and the same machine. Somebody should confirm a pass on 22.22.3 before this merges.The memory harness still writes every field
gate.tsreads, and its budgeted config passes with unchanged behavior.For the note itself, the review is reading it: please check the claims against the referenced files and flag anything that does not match reality.
Documentation
MIGRATION.MD
Checklist for Maintainers
When this PR is ready for testing, make sure to add
ci:normal,ci:mergedorci:dailyGH label to it to run a specific set of sandboxes. The particular set of sandboxes can be found incode/lib/cli-storybook/src/sandbox-templates.tsDeclare whether manual QA will be needed for this PR during the next release, through
qa:neededorqa:skipMake sure this PR contains one of the labels below:
Available labels
bug: Internal changes that fixes incorrect behavior.maintenance: User-facing maintenance tasks.dependencies: Upgrading (sometimes downgrading) dependencies.build: Internal-facing build tooling & test updates. Will not show up in release changelog.cleanup: Minor cleanup style change. Will not show up in release changelog.documentation: Documentation only changes. Will not show up in release changelog.feature request: Introducing a new feature.BREAKING CHANGE: Changes that break compatibility in some way with current major version.other: Changes that don't fit in the above categories.🦋 Canary release
This PR does not have a canary release associated. You can request a canary release of this pull request by mentioning the
@storybookjs/coreteam here.core team members can create a canary release here or locally with
gh workflow run --repo storybookjs/storybook publish.yml --field pr=<PR_NUMBER>