fix(test): cap local unit-test concurrency at 4 to avoid exhausting commit charge - #13187
Conversation
|
Self-review note, and a pre-existing bug this change does not fix. Two things worth a maintainer'''s eye: 1. This PR also adds Worth noting the mild irony: 2. Pre-existing, untouched here: Either |
|
Thanks for capping local unit-test concurrency. I did not fold this into the local recut of Conflict was only Purpose of the recut: ship a production image, not change local test pacing. No objection to the change on developer laptops; it is not a production-image change, so it stays out of this recut. |
178d252
into
diegosouzapw:release/v3.8.51
The dedupe made `test` delegate to `test:unit`, but it also silently reverted the deliberate local concurrency cap from diegosouzapw#13187 ("cap local unit-test concurrency at 4 to avoid exhausting commit charge") back to 20. Measured on m1max (16 cores), same suite and same tree, only the flag differs: concurrency=20 -> 356 cancelled, 356 "event loop has already resolved" bailouts concurrency=4 -> 0 cancelled, 0 bailouts So 20 does not just slow the run down, it makes the runner abandon tests and still print a summary -- a false green. Restore 4; termination is preserved via delegation to test:unit, which already carries --test-force-exit.
The diegosouzapw#13187 batch commit (178d252) already gave `test` both `--test-force-exit` flags and the trailing `&& npm run test:unit:serial` step, so the delegation fix was redundant. It also broke tests/unit/test-serial-quarantine.test.ts, which asserts every parallel runner script ends with the serial step (base 4/4 -> head 3/4). package.json is now byte-identical to base; this PR is the stryker tap.testFiles fix only.
…iles (#13357) * fix(test): make npm run test terminate and restore RAYCAST env-doc sync Two independent defects, both in the test/dev entrypoint layer. 1. `npm run test` never terminated. It was a hand-maintained copy of `test:unit` that had drifted: it omitted `--test-force-exit` on BOTH node invocations and dropped the trailing `&& npm run test:unit:serial`. Per AGENTS.md ('Database Handles in Tests'), unreleased SQLite handles make Node's native runner hang indefinitely — every sibling script (`test:unit`, `test:unit:ci`, `test:unit:ci:shard`) already carried the flag; only `test` did not. Measured on m1max at 84c6ad7, same suite both arms: without the flag the runner was killed at the 420s ceiling (exit 137, no summary line, 23 orphaned node processes); with it the runner exited on its own in 419s leaving 1. `test` now delegates to `test:unit` so the two cannot drift again, which also makes the serial suite reachable from `npm run test` for the first time. 2. Removing the RAYCAST_* rows from ENVIRONMENT.md (#9) broke check-env-doc-sync. `parseEnvExampleVars` matches `^#?\s*(VAR)=`, so it counts COMMENTED-OUT vars: the four entries still sat at .env.example:1263-1266 and became `envMissingDoc` drift the moment their docs disappeared. The #9 verification only ran the fabricated-docs gate and missed this one. The block is dead either way — it documents open-sse/services/raycast.ts and scripts/raycast/usage-benchmark.mjs, both deleted with the GPL-derived provider in #11691, and no live code reads the vars — so it is removed rather than re-documented. envMissingDoc is now []. The remaining codeMissingEnv failure (CURSOR_AGENT_BINARY, CURSOR_MAX_FRAME_BYTES, OMNIROOT) is pre-existing drift on the base, absent from this diff, and left alone. * chore(stryker): register 3 covering unit tests missing from tap.testFiles check:mutation-test-coverage --strict fails identically on pristine release/v3.8.51 (f1e7148) with an empty diff — base debt blocking this PR. - combo-identical-error-streak.test.ts -> comboPredicates.ts - 13601-header-drop-count-surfaced.test.ts -> responseHeaders.ts - semantic-cache-no-truncated-writes.test.ts -> semanticCache.ts * fix(test): keep the #13187 concurrency-4 cap in test:unit The dedupe made `test` delegate to `test:unit`, but it also silently reverted the deliberate local concurrency cap from #13187 ("cap local unit-test concurrency at 4 to avoid exhausting commit charge") back to 20. Measured on m1max (16 cores), same suite and same tree, only the flag differs: concurrency=20 -> 356 cancelled, 356 "event loop has already resolved" bailouts concurrency=4 -> 0 cancelled, 0 bailouts So 20 does not just slow the run down, it makes the runner abandon tests and still print a summary -- a false green. Restore 4; termination is preserved via delegation to test:unit, which already carries --test-force-exit. * revert(test): drop redundant test-script delegation The #13187 batch commit (178d252) already gave `test` both `--test-force-exit` flags and the trailing `&& npm run test:unit:serial` step, so the delegation fix was redundant. It also broke tests/unit/test-serial-quarantine.test.ts, which asserts every parallel runner script ends with the serial step (base 4/4 -> head 3/4). package.json is now byte-identical to base; this PR is the stryker tap.testFiles fix only.
…ommit charge (diegosouzapw#13187) Aligns the two hand-typed local scripts with `test:unit:ci`, which already ran at concurrency 4; `--test-force-exit` was likewise the one flag `test` was missing. CI scripts are untouched. --- Validated in one consolidated worktree cut from `release/v3.8.51`, boarded together with the other 13 PRs of this batch — zero merge conflicts between them. - `typecheck:core` clean - complexity 2799 / baseline 3218 and cognitive-complexity 1265 / baseline 1437 — both under baseline - 71 focused assertions green across the 13 test files this batch adds or touches⚠️ base-red inherited: diegosouzapw#12732 — `Docs Gates (fast-path)`, `Merge integrity`, `No new ESLint warnings`, `Unit Tests fast-path` and `Fast Quality Gates` all reproduce on the pure `release/v3.8.51` tip (provider count 356 vs the 358 the modules define, SKILL.md drift, and `open-sse/utils/stream.ts` at 3115 > frozen 3098). None of them touch this diff. Thanks @anhtahaylove — the root-cause write-up, the measured before/after numbers and the red-before-green proof on every one of these made the batch reviewable as a unit.
…iles (diegosouzapw#13357) * fix(test): make npm run test terminate and restore RAYCAST env-doc sync Two independent defects, both in the test/dev entrypoint layer. 1. `npm run test` never terminated. It was a hand-maintained copy of `test:unit` that had drifted: it omitted `--test-force-exit` on BOTH node invocations and dropped the trailing `&& npm run test:unit:serial`. Per AGENTS.md ('Database Handles in Tests'), unreleased SQLite handles make Node's native runner hang indefinitely — every sibling script (`test:unit`, `test:unit:ci`, `test:unit:ci:shard`) already carried the flag; only `test` did not. Measured on m1max at 84c6ad7, same suite both arms: without the flag the runner was killed at the 420s ceiling (exit 137, no summary line, 23 orphaned node processes); with it the runner exited on its own in 419s leaving 1. `test` now delegates to `test:unit` so the two cannot drift again, which also makes the serial suite reachable from `npm run test` for the first time. 2. Removing the RAYCAST_* rows from ENVIRONMENT.md (diegosouzapw#9) broke check-env-doc-sync. `parseEnvExampleVars` matches `^#?\s*(VAR)=`, so it counts COMMENTED-OUT vars: the four entries still sat at .env.example:1263-1266 and became `envMissingDoc` drift the moment their docs disappeared. The diegosouzapw#9 verification only ran the fabricated-docs gate and missed this one. The block is dead either way — it documents open-sse/services/raycast.ts and scripts/raycast/usage-benchmark.mjs, both deleted with the GPL-derived provider in diegosouzapw#11691, and no live code reads the vars — so it is removed rather than re-documented. envMissingDoc is now []. The remaining codeMissingEnv failure (CURSOR_AGENT_BINARY, CURSOR_MAX_FRAME_BYTES, OMNIROOT) is pre-existing drift on the base, absent from this diff, and left alone. * chore(stryker): register 3 covering unit tests missing from tap.testFiles check:mutation-test-coverage --strict fails identically on pristine release/v3.8.51 (7339a9f) with an empty diff — base debt blocking this PR. - combo-identical-error-streak.test.ts -> comboPredicates.ts - 13601-header-drop-count-surfaced.test.ts -> responseHeaders.ts - semantic-cache-no-truncated-writes.test.ts -> semanticCache.ts * fix(test): keep the diegosouzapw#13187 concurrency-4 cap in test:unit The dedupe made `test` delegate to `test:unit`, but it also silently reverted the deliberate local concurrency cap from diegosouzapw#13187 ("cap local unit-test concurrency at 4 to avoid exhausting commit charge") back to 20. Measured on m1max (16 cores), same suite and same tree, only the flag differs: concurrency=20 -> 356 cancelled, 356 "event loop has already resolved" bailouts concurrency=4 -> 0 cancelled, 0 bailouts So 20 does not just slow the run down, it makes the runner abandon tests and still print a summary -- a false green. Restore 4; termination is preserved via delegation to test:unit, which already carries --test-force-exit. * revert(test): drop redundant test-script delegation The diegosouzapw#13187 batch commit (5d3c7d6) already gave `test` both `--test-force-exit` flags and the trailing `&& npm run test:unit:serial` step, so the delegation fix was redundant. It also broke tests/unit/test-serial-quarantine.test.ts, which asserts every parallel runner script ends with the serial step (base 4/4 -> head 3/4). package.json is now byte-identical to base; this PR is the stryker tap.testFiles fix only.
Problem
The local
testandtest:unitscripts run with--test-concurrency=20and--max-old-space-size=8192. Those multiply: up to 20 workers × 8 GB of heap reservation each. On a 32 GB developer machine with a ~53.8 GB commit limit, a full-suite run drives commit charge to the ceiling, and Windows starts refusing process creation — unrelated long-running Node processes die with0xC0000142(STATUS_DLL_INIT_FAILED, the usual symptom of a failed allocation at process start).This is a local-only footgun. CI never hits it:
ci.ymlandquality.ymlruntest:unit:ci:shard, which already uses--test-concurrency=4and spreads the suite across 8 GitHub-hosted runners with 16 GB each. The unshardedtest/test:unitscripts are the ones a developer types by hand, and they are the only ones tuned for a machine that doesn't exist.Fix
Set
--test-concurrency=4on both, matchingtest:unit:ci.Also added the missing
--test-force-exittotest. Without it the runner finishes every test and then hangs indefinitely on a live handle instead of exiting —test:unitand all the CI variants already pass this flag;testwas the odd one out.Verification
Full suite before and after, same machine, same base commit — the only difference is this
package.jsonchange:--test-concurrency=20--test-concurrency=4Sampled commit charge every 9 s across the run. No unrelated process died, and the local gateway on port 20128 stayed up throughout.
Results are unchanged: 413 of 415 distinct failing test names are identical across the two runs, differing by 2 in each direction — ordinary flakiness, not a behaviour change. (Those 791 failures are pre-existing on this base on Windows, mostly
spawn/ENOENTfrom tooling this machine doesn't have installed. They are unrelated to this change.)Wall time is not meaningfully worse: 1418 s here versus 1499 s for the pre-change baseline run. The 20-way setting was buying nothing — the suite is I/O-bound long before it is CPU-bound, and the oversubscription was pure memory pressure.