Skip to content

fix(docker): size the Next build worker pool for a 16 GB runner - #11419

Merged
diegosouzapw merged 1 commit into
release/v3.8.50from
fix/docker-build-oom-page-data
Aug 24, 2026
Merged

diegosouzapw merged 1 commit into
release/v3.8.50from
fix/docker-build-oom-page-data

Conversation

@diegosouzapw

Copy link
Copy Markdown
Owner

Fixes the "Publish to Docker Hub" pipeline, red on 96 of the last 100 runs — every run since 2026-08-22 23:14 UTC. Example: run 32666809903, both linux/amd64 and linux/arm64.

Diagnosis

ERROR: failed to solve: ResourceExhausted: process "/bin/sh -c ... npm run build ..."
did not complete successfully: cannot allocate memory

That is the kernel, not a V8 heap OOM — and the log says exactly where:

#21 256.1 ✓ Compiled successfully in 4.2min
#21 256.1   Collecting page data using 7 workers ...
...
#21 ERROR: ... cannot allocate memory

The compile phase always finishes. The build dies in page-data collection.

Each page-data worker is its own process and inherits NODE_OPTIONS, so --max-old-space-size=6144 is a ceiling per process, not per build. CIRCLE_NODE_TOTAL=8 gives 7 workers, and 7 of them alongside the parent no longer fit the 16 GB / 4 vCPU GitHub-hosted runners (ubuntu-24.04, ubuntu-24.04-arm). It was intermittent before going 100% — success and failure alternated on 08-22 — which is what a threshold crossed by ordinary codebase growth looks like. 7 workers was also oversubscribing a 4 vCPU runner.

Fix

Lower the pool to 3 (2 workers) and make it a build arg, so a large builder can raise it back:

docker build --build-arg OMNIROUTE_BUILD_WORKERS=8 .

The existing #10060 comment already capped this at 8 for the same class of failure on 32-core builders; the value simply is not low enough for a 4 vCPU / 16 GB runner any more.

Validation

tests/unit/docker-build-memory-budget.test.ts — 3 tests, red on the base (3/3), green here (3/3). It reads the two ARG defaults straight out of the Dockerfile and fails if parent heap + workers × per-worker peak outgrows the runner, or if the pool oversubscribes its CPUs. So a future one-line bump to either knob has to re-do the arithmetic instead of silently reddening the publish pipeline again.

The per-worker peak the budget uses (2560 MB) is documented as an inference from this failure — 7 workers did not fit in 16 GB, which puts the peak north of ~1.8 GB — not as a measurement. That is called out in the test's comment.

Other checks:

  • Dockerfile sibling tests (dockerfile-dashboard-embed-arg-10273, dockerfile-npm-bundled-cve-patch, dockerfile-better-sqlite3-node-gyp-6700, tls-client-node-docker-binary-7802, dockerfile-base-path-arg, docker-llmlingua-optionals-9166, assemble-standalone-onnxruntime-native-asset, docker-healthcheck-base-path) — 26/26
  • check:docs-all (doc-links + fabricated-docs), check:tracked-artifacts — PASS
  • prettier --check, eslint — clean

docs/guides/DOCKER_GUIDE.md's build-arg table was stale (still listing the pre-#10060 4096 default); updated, plus the new knob and the symptom to recognize — a build that dies after ✓ Compiled successfully.

CIRCLE_NODE_TOTAL and OMNIROUTE_BUILD_WORKERS are added to the fabricated-docs env allowlist with the reason inline: neither is read via process.env in this repo — one is a Dockerfile ARG, the other is read by Next itself.

⚠️ What this PR cannot prove

The failure only reproduces on a memory-constrained host, so no unit test can reproduce the OOM — the test guards the arithmetic, not the outcome. The real validation is the next publish run after merge. If it still OOMs, the knob to turn is WORKER_PEAK_MB in the test plus a lower OMNIROUTE_BUILD_WORKERS; the diagnosis (page-data workers, not the compile pass) holds either way.

⚠️ base-red inherited: #9985

Every "Publish to Docker Hub" run has failed since 2026-08-22 23:14 UTC — 96 of
the last 100. The builder stage dies with:

  ERROR: failed to solve: ResourceExhausted: process "/bin/sh -c ... npm run
  build ..." did not complete successfully: cannot allocate memory

That is the kernel, not V8. The log puts it precisely: the compile phase always
finishes ("✓ Compiled successfully in 4.2min") and the build is killed right
after "Collecting page data using 7 workers".

Each page-data worker is its own process and inherits NODE_OPTIONS, so the
--max-old-space-size ceiling is per PROCESS, not per build. CIRCLE_NODE_TOTAL=8
means 7 workers, and 7 of them alongside the parent no longer fit the 16 GB /
4 vCPU GitHub-hosted runners the pipeline builds on. It was intermittent for a
while before going 100%, which is what a threshold crossed by ordinary codebase
growth looks like — 7 was also oversubscribing a 4 vCPU runner.

Lower the pool to 3 (2 workers) and make it a build arg, so a big builder can
raise it back with `--build-arg OMNIROUTE_BUILD_WORKERS=8`.

tests/unit/docker-build-memory-budget.test.ts pins the budget: it reads the two
ARG defaults out of the Dockerfile and fails if `parent heap + workers × peak`
outgrows the runner, or if the pool oversubscribes its CPUs. Red on the base
(3/3), green here (3/3). The per-worker peak it budgets with is documented as an
inference from this failure, not a measurement.

DOCKER_GUIDE's build-arg table was stale (it still listed the pre-#10060 4096 MB
default); updated and given the new knob plus the symptom to recognize.
CIRCLE_NODE_TOTAL and OMNIROUTE_BUILD_WORKERS are allowlisted in the
fabricated-docs gate with the reason: neither is read via process.env here — one
is a Dockerfile ARG, the other is read by Next itself.

Note: the real proof is the next publish run. This failure mode only reproduces
on a memory-constrained host, so it cannot be reproduced by the unit suite; the
test guards the arithmetic, not the outcome.
@diegosouzapw
diegosouzapw merged commit 8bbe92c into release/v3.8.50 Aug 24, 2026
14 of 22 checks passed
raheemuddin786 pushed a commit to raheemuddin786/OmniRoute that referenced this pull request Aug 24, 2026
…osouzapw#11419)

Every "Publish to Docker Hub" run has failed since 2026-08-22 23:14 UTC — 96 of
the last 100. The builder stage dies with:

  ERROR: failed to solve: ResourceExhausted: process "/bin/sh -c ... npm run
  build ..." did not complete successfully: cannot allocate memory

That is the kernel, not V8. The log puts it precisely: the compile phase always
finishes ("✓ Compiled successfully in 4.2min") and the build is killed right
after "Collecting page data using 7 workers".

Each page-data worker is its own process and inherits NODE_OPTIONS, so the
--max-old-space-size ceiling is per PROCESS, not per build. CIRCLE_NODE_TOTAL=8
means 7 workers, and 7 of them alongside the parent no longer fit the 16 GB /
4 vCPU GitHub-hosted runners the pipeline builds on. It was intermittent for a
while before going 100%, which is what a threshold crossed by ordinary codebase
growth looks like — 7 was also oversubscribing a 4 vCPU runner.

Lower the pool to 3 (2 workers) and make it a build arg, so a big builder can
raise it back with `--build-arg OMNIROUTE_BUILD_WORKERS=8`.

tests/unit/docker-build-memory-budget.test.ts pins the budget: it reads the two
ARG defaults out of the Dockerfile and fails if `parent heap + workers × peak`
outgrows the runner, or if the pool oversubscribes its CPUs. Red on the base
(3/3), green here (3/3). The per-worker peak it budgets with is documented as an
inference from this failure, not a measurement.

DOCKER_GUIDE's build-arg table was stale (it still listed the pre-diegosouzapw#10060 4096 MB
default); updated and given the new knob plus the symptom to recognize.
CIRCLE_NODE_TOTAL and OMNIROUTE_BUILD_WORKERS are allowlisted in the
fabricated-docs gate with the reason: neither is read via process.env here — one
is a Dockerfile ARG, the other is read by Next itself.

Note: the real proof is the next publish run. This failure mode only reproduces
on a memory-constrained host, so it cannot be reproduced by the unit suite; the
test guards the arithmetic, not the outcome.

Co-authored-by: Xiangzhe <bakryun0718@proton.me>
@diegosouzapw
diegosouzapw deleted the fix/docker-build-oom-page-data branch August 25, 2026 02:37
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…osouzapw#11419)

Every "Publish to Docker Hub" run has failed since 2026-08-22 23:14 UTC — 96 of
the last 100. The builder stage dies with:

  ERROR: failed to solve: ResourceExhausted: process "/bin/sh -c ... npm run
  build ..." did not complete successfully: cannot allocate memory

That is the kernel, not V8. The log puts it precisely: the compile phase always
finishes ("✓ Compiled successfully in 4.2min") and the build is killed right
after "Collecting page data using 7 workers".

Each page-data worker is its own process and inherits NODE_OPTIONS, so the
--max-old-space-size ceiling is per PROCESS, not per build. CIRCLE_NODE_TOTAL=8
means 7 workers, and 7 of them alongside the parent no longer fit the 16 GB /
4 vCPU GitHub-hosted runners the pipeline builds on. It was intermittent for a
while before going 100%, which is what a threshold crossed by ordinary codebase
growth looks like — 7 was also oversubscribing a 4 vCPU runner.

Lower the pool to 3 (2 workers) and make it a build arg, so a big builder can
raise it back with `--build-arg OMNIROUTE_BUILD_WORKERS=8`.

tests/unit/docker-build-memory-budget.test.ts pins the budget: it reads the two
ARG defaults out of the Dockerfile and fails if `parent heap + workers × peak`
outgrows the runner, or if the pool oversubscribes its CPUs. Red on the base
(3/3), green here (3/3). The per-worker peak it budgets with is documented as an
inference from this failure, not a measurement.

DOCKER_GUIDE's build-arg table was stale (it still listed the pre-diegosouzapw#10060 4096 MB
default); updated and given the new knob plus the symptom to recognize.
CIRCLE_NODE_TOTAL and OMNIROUTE_BUILD_WORKERS are allowlisted in the
fabricated-docs gate with the reason: neither is read via process.env here — one
is a Dockerfile ARG, the other is read by Next itself.

Note: the real proof is the next publish run. This failure mode only reproduces
on a memory-constrained host, so it cannot be reproduced by the unit suite; the
test guards the arithmetic, not the outcome.

Co-authored-by: Xiangzhe <bakryun0718@proton.me>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants