Repository navigation
fix(ci): give main's build its own lane on the self-hosted pool - #11901
Merged
Merged
Conversation
The .113 box has 31 GB and a single next-build peaks at 14–16 GB RSS: one build fits with room, two sit at the edge, three take the box down. On 2026-08-28 13:50Z the kernel OOM-killed main's next-build (15.7 GB) while a PR build ran beside it — five Build jobs had been queued by a burst of PRs — and the publish lost its artefact, which sends it into the 40-minute rebuild that OOMs on its own (attempt 5 of this release). Job-level concurrency on `build`, two lanes: heavy-build-main pushes to main — never contended, never behind PR traffic heavy-build-pr pull requests — serialize among themselves cancel-in-progress stays false: a running build is never killed by a newer one. GitHub's own rule for a group is one running + one pending, older pendings cancelled — so under a burst the third PR build shows "cancelled" and needs a re-run. That is the trade-off, stated: a cancelled PR check is re-runnable; a dead main build costs a release. The proper fix remains a label split (omni-build on two runners, omni-light on the rest) so the queue lives on the runner side without cancellations — an operator decision recorded in docs/ops/RUNNER_BOX.md.
Contributor
CI Coverage Report
Coverage artifact was not available for this run. |
diegosouzapw
added a commit
that referenced
this pull request
Aug 28, 2026
…e pipeline fixes Brings e4683cd (#11867 Alibaba allowlist time bomb), 09de69e (#11891 config expiry detector), e71be03 (#11893 runner janitor), 9dc8eab (#11895 provenance × self-hosted lint) and f564b64 (#11901 heavy-build lanes). main is already an ancestor of this branch (v3.8.50 sync-back), so the merge is exactly these five commits. # Conflicts: # tests/unit/alibaba-free-tier-allowlist.test.ts
This was referenced Aug 28, 2026
diegosouzapw
added a commit
that referenced
this pull request
Aug 28, 2026
…11932) The .113 box (31 GB) holds one next-build (14–16 GB RSS) comfortably and two at the edge; on 2026-08-28 the kernel killed main's build twice while PR builds ran beside it. Labels are the runner-side cap: only omniroute-113-5 and omniroute-113-6 carry omni-build (added through the runners API, no re-registration), and every job that runs a next build — ci.yml build, npm-publish.yml publish, both nightly-release-green validations — now asks for that label. A third heavy job queues on GitHub instead of racing for memory. The six other runners keep omni-release and no longer take builds. Pairs with the heavy-build-* concurrency lanes (#11901); documented in docs/ops/RUNNER_BOX.md.
muhamadgalihsaputra
pushed a commit
to niyatna/NiyatnaRoute
that referenced
this pull request
Sep 27, 2026
…osouzapw#11901) The .113 box has 31 GB and a single next-build peaks at 14–16 GB RSS: one build fits with room, two sit at the edge, three take the box down. On 2026-08-28 13:50Z the kernel OOM-killed main's next-build (15.7 GB) while a PR build ran beside it — five Build jobs had been queued by a burst of PRs — and the publish lost its artefact, which sends it into the 40-minute rebuild that OOMs on its own (attempt 5 of this release). Job-level concurrency on `build`, two lanes: heavy-build-main pushes to main — never contended, never behind PR traffic heavy-build-pr pull requests — serialize among themselves cancel-in-progress stays false: a running build is never killed by a newer one. GitHub's own rule for a group is one running + one pending, older pendings cancelled — so under a burst the third PR build shows "cancelled" and needs a re-run. That is the trade-off, stated: a cancelled PR check is re-runnable; a dead main build costs a release. The proper fix remains a label split (omni-build on two runners, omni-light on the rest) so the queue lives on the runner side without cancellations — an operator decision recorded in docs/ops/RUNNER_BOX.md.
muhamadgalihsaputra
pushed a commit
to niyatna/NiyatnaRoute
that referenced
this pull request
Sep 27, 2026
…e pipeline fixes Brings 342ca95 (diegosouzapw#11867 Alibaba allowlist time bomb), ce53544 (diegosouzapw#11891 config expiry detector), 8b8d84d (diegosouzapw#11893 runner janitor), 9e87941 (diegosouzapw#11895 provenance × self-hosted lint) and 2fe1c30 (diegosouzapw#11901 heavy-build lanes). main is already an ancestor of this branch (v3.8.50 sync-back), so the merge is exactly these five commits. # Conflicts: # tests/unit/alibaba-free-tier-allowlist.test.ts
muhamadgalihsaputra
pushed a commit
to niyatna/NiyatnaRoute
that referenced
this pull request
Sep 27, 2026
…iegosouzapw#11932) The .113 box (31 GB) holds one next-build (14–16 GB RSS) comfortably and two at the edge; on 2026-08-28 the kernel killed main's build twice while PR builds ran beside it. Labels are the runner-side cap: only omniroute-113-5 and omniroute-113-6 carry omni-build (added through the runners API, no re-registration), and every job that runs a next build — ci.yml build, npm-publish.yml publish, both nightly-release-green validations — now asks for that label. A third heavy job queues on GitHub instead of racing for memory. The six other runners keep omni-release and no longer take builds. Pairs with the heavy-build-* concurrency lanes (diegosouzapw#11901); documented in docs/ops/RUNNER_BOX.md.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problema
A
.113tem 31 GB e umnext-buildsozinho chega a 14–16 GB de RSS: um build cabe com folga, dois ficam no limite, três derrubam a caixa. Em 2026-08-28 13:50Z o kernel matou por OOM onext-builddemain(15,7 GB) enquanto um build de PR rodava ao lado — cinco jobsBuildtinham sido enfileirados por uma rajada de PRs. Sem artefato, o publish cai no rebuild de 40 min que estoura memória sozinho (tentativa nº 5 desta release).O que muda
concurrencyno jobbuild, em duas faixas:heavy-build-mainmainheavy-build-prcancel-in-progress: false— um build em execução nunca é morto por um mais novo.O trade-off, dito claramente
A regra do GitHub por grupo é 1 em execução + 1 pendente; pendentes mais antigos são cancelados. Numa rajada, o 3º build de PR aparece como "cancelled" e precisa de re-run. Um check de PR cancelado é re-executável; um build de
mainmorto custa uma release.O conserto definitivo continua sendo o split por label (
omni-buildem 2 runners,omni-lightnos demais), que move a fila para o lado do runner sem cancelamentos — decisão de operador, registrada emdocs/ops/RUNNER_BOX.md(#11893).Validação
check:workflows --ratchet: zizmor inalterado (194)heavy-build-prUnit Tests (1/8)— allowlist da Alibaba, consertada em #11867)