fix(docker): size the Next build worker pool for a 16 GB runner - #11419
Merged
Merged
Conversation
Every "Publish to Docker Hub" run has failed since 2026-08-22 23:14 UTC — 96 of
the last 100. The builder stage dies with:
ERROR: failed to solve: ResourceExhausted: process "/bin/sh -c ... npm run
build ..." did not complete successfully: cannot allocate memory
That is the kernel, not V8. The log puts it precisely: the compile phase always
finishes ("✓ Compiled successfully in 4.2min") and the build is killed right
after "Collecting page data using 7 workers".
Each page-data worker is its own process and inherits NODE_OPTIONS, so the
--max-old-space-size ceiling is per PROCESS, not per build. CIRCLE_NODE_TOTAL=8
means 7 workers, and 7 of them alongside the parent no longer fit the 16 GB /
4 vCPU GitHub-hosted runners the pipeline builds on. It was intermittent for a
while before going 100%, which is what a threshold crossed by ordinary codebase
growth looks like — 7 was also oversubscribing a 4 vCPU runner.
Lower the pool to 3 (2 workers) and make it a build arg, so a big builder can
raise it back with `--build-arg OMNIROUTE_BUILD_WORKERS=8`.
tests/unit/docker-build-memory-budget.test.ts pins the budget: it reads the two
ARG defaults out of the Dockerfile and fails if `parent heap + workers × peak`
outgrows the runner, or if the pool oversubscribes its CPUs. Red on the base
(3/3), green here (3/3). The per-worker peak it budgets with is documented as an
inference from this failure, not a measurement.
DOCKER_GUIDE's build-arg table was stale (it still listed the pre-#10060 4096 MB
default); updated and given the new knob plus the symptom to recognize.
CIRCLE_NODE_TOTAL and OMNIROUTE_BUILD_WORKERS are allowlisted in the
fabricated-docs gate with the reason: neither is read via process.env here — one
is a Dockerfile ARG, the other is read by Next itself.
Note: the real proof is the next publish run. This failure mode only reproduces
on a memory-constrained host, so it cannot be reproduced by the unit suite; the
test guards the arithmetic, not the outcome.
raheemuddin786
pushed a commit
to raheemuddin786/OmniRoute
that referenced
this pull request
Aug 24, 2026
…osouzapw#11419) Every "Publish to Docker Hub" run has failed since 2026-08-22 23:14 UTC — 96 of the last 100. The builder stage dies with: ERROR: failed to solve: ResourceExhausted: process "/bin/sh -c ... npm run build ..." did not complete successfully: cannot allocate memory That is the kernel, not V8. The log puts it precisely: the compile phase always finishes ("✓ Compiled successfully in 4.2min") and the build is killed right after "Collecting page data using 7 workers". Each page-data worker is its own process and inherits NODE_OPTIONS, so the --max-old-space-size ceiling is per PROCESS, not per build. CIRCLE_NODE_TOTAL=8 means 7 workers, and 7 of them alongside the parent no longer fit the 16 GB / 4 vCPU GitHub-hosted runners the pipeline builds on. It was intermittent for a while before going 100%, which is what a threshold crossed by ordinary codebase growth looks like — 7 was also oversubscribing a 4 vCPU runner. Lower the pool to 3 (2 workers) and make it a build arg, so a big builder can raise it back with `--build-arg OMNIROUTE_BUILD_WORKERS=8`. tests/unit/docker-build-memory-budget.test.ts pins the budget: it reads the two ARG defaults out of the Dockerfile and fails if `parent heap + workers × peak` outgrows the runner, or if the pool oversubscribes its CPUs. Red on the base (3/3), green here (3/3). The per-worker peak it budgets with is documented as an inference from this failure, not a measurement. DOCKER_GUIDE's build-arg table was stale (it still listed the pre-diegosouzapw#10060 4096 MB default); updated and given the new knob plus the symptom to recognize. CIRCLE_NODE_TOTAL and OMNIROUTE_BUILD_WORKERS are allowlisted in the fabricated-docs gate with the reason: neither is read via process.env here — one is a Dockerfile ARG, the other is read by Next itself. Note: the real proof is the next publish run. This failure mode only reproduces on a memory-constrained host, so it cannot be reproduced by the unit suite; the test guards the arithmetic, not the outcome. Co-authored-by: Xiangzhe <bakryun0718@proton.me>
muhamadgalihsaputra
pushed a commit
to niyatna/NiyatnaRoute
that referenced
this pull request
Sep 27, 2026
…osouzapw#11419) Every "Publish to Docker Hub" run has failed since 2026-08-22 23:14 UTC — 96 of the last 100. The builder stage dies with: ERROR: failed to solve: ResourceExhausted: process "/bin/sh -c ... npm run build ..." did not complete successfully: cannot allocate memory That is the kernel, not V8. The log puts it precisely: the compile phase always finishes ("✓ Compiled successfully in 4.2min") and the build is killed right after "Collecting page data using 7 workers". Each page-data worker is its own process and inherits NODE_OPTIONS, so the --max-old-space-size ceiling is per PROCESS, not per build. CIRCLE_NODE_TOTAL=8 means 7 workers, and 7 of them alongside the parent no longer fit the 16 GB / 4 vCPU GitHub-hosted runners the pipeline builds on. It was intermittent for a while before going 100%, which is what a threshold crossed by ordinary codebase growth looks like — 7 was also oversubscribing a 4 vCPU runner. Lower the pool to 3 (2 workers) and make it a build arg, so a big builder can raise it back with `--build-arg OMNIROUTE_BUILD_WORKERS=8`. tests/unit/docker-build-memory-budget.test.ts pins the budget: it reads the two ARG defaults out of the Dockerfile and fails if `parent heap + workers × peak` outgrows the runner, or if the pool oversubscribes its CPUs. Red on the base (3/3), green here (3/3). The per-worker peak it budgets with is documented as an inference from this failure, not a measurement. DOCKER_GUIDE's build-arg table was stale (it still listed the pre-diegosouzapw#10060 4096 MB default); updated and given the new knob plus the symptom to recognize. CIRCLE_NODE_TOTAL and OMNIROUTE_BUILD_WORKERS are allowlisted in the fabricated-docs gate with the reason: neither is read via process.env here — one is a Dockerfile ARG, the other is read by Next itself. Note: the real proof is the next publish run. This failure mode only reproduces on a memory-constrained host, so it cannot be reproduced by the unit suite; the test guards the arithmetic, not the outcome. Co-authored-by: Xiangzhe <bakryun0718@proton.me>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes the "Publish to Docker Hub" pipeline, red on 96 of the last 100 runs — every run since 2026-08-22 23:14 UTC. Example: run 32666809903, both
linux/amd64andlinux/arm64.Diagnosis
That is the kernel, not a V8 heap OOM — and the log says exactly where:
The compile phase always finishes. The build dies in page-data collection.
Each page-data worker is its own process and inherits
NODE_OPTIONS, so--max-old-space-size=6144is a ceiling per process, not per build.CIRCLE_NODE_TOTAL=8gives 7 workers, and 7 of them alongside the parent no longer fit the 16 GB / 4 vCPU GitHub-hosted runners (ubuntu-24.04,ubuntu-24.04-arm). It was intermittent before going 100% — success and failure alternated on 08-22 — which is what a threshold crossed by ordinary codebase growth looks like. 7 workers was also oversubscribing a 4 vCPU runner.Fix
Lower the pool to
3(2 workers) and make it a build arg, so a large builder can raise it back:docker build --build-arg OMNIROUTE_BUILD_WORKERS=8 .The existing
#10060comment already capped this at 8 for the same class of failure on 32-core builders; the value simply is not low enough for a 4 vCPU / 16 GB runner any more.Validation
tests/unit/docker-build-memory-budget.test.ts— 3 tests, red on the base (3/3), green here (3/3). It reads the two ARG defaults straight out of the Dockerfile and fails ifparent heap + workers × per-worker peakoutgrows the runner, or if the pool oversubscribes its CPUs. So a future one-line bump to either knob has to re-do the arithmetic instead of silently reddening the publish pipeline again.The per-worker peak the budget uses (2560 MB) is documented as an inference from this failure — 7 workers did not fit in 16 GB, which puts the peak north of ~1.8 GB — not as a measurement. That is called out in the test's comment.
Other checks:
dockerfile-dashboard-embed-arg-10273,dockerfile-npm-bundled-cve-patch,dockerfile-better-sqlite3-node-gyp-6700,tls-client-node-docker-binary-7802,dockerfile-base-path-arg,docker-llmlingua-optionals-9166,assemble-standalone-onnxruntime-native-asset,docker-healthcheck-base-path) — 26/26check:docs-all(doc-links + fabricated-docs),check:tracked-artifacts— PASSprettier --check,eslint— cleandocs/guides/DOCKER_GUIDE.md's build-arg table was stale (still listing the pre-#100604096default); updated, plus the new knob and the symptom to recognize — a build that dies after✓ Compiled successfully.CIRCLE_NODE_TOTALandOMNIROUTE_BUILD_WORKERSare added to the fabricated-docs env allowlist with the reason inline: neither is read viaprocess.envin this repo — one is a Dockerfile ARG, the other is read by Next itself.The failure only reproduces on a memory-constrained host, so no unit test can reproduce the OOM — the test guards the arithmetic, not the outcome. The real validation is the next publish run after merge. If it still OOMs, the knob to turn is
WORKER_PEAK_MBin the test plus a lowerOMNIROUTE_BUILD_WORKERS; the diagnosis (page-data workers, not the compile pass) holds either way.