Skip to content

fix(ci): give main's build its own lane on the self-hosted pool - #11901

Merged
diegosouzapw merged 1 commit into
mainfrom
fix/ci-heavy-build-concurrency-lane
Aug 28, 2026
Merged

diegosouzapw merged 1 commit into
mainfrom
fix/ci-heavy-build-concurrency-lane

Conversation

@diegosouzapw

Copy link
Copy Markdown
Owner

Problema

A .113 tem 31 GB e um next-build sozinho chega a 14–16 GB de RSS: um build cabe com folga, dois ficam no limite, três derrubam a caixa. Em 2026-08-28 13:50Z o kernel matou por OOM o next-build de main (15,7 GB) enquanto um build de PR rodava ao lado — cinco jobs Build tinham sido enfileirados por uma rajada de PRs. Sem artefato, o publish cai no rebuild de 40 min que estoura memória sozinho (tentativa nº 5 desta release).

[13:50:41] Out of memory: Killed process 540001 (next-build) anon-rss:15766796kB
           task_memcg=/system.slice/actions.runner.diegosouzapw-OmniRoute.omniroute-113-6.service
[13:50:56] Worker: Step result: Canceled
[13:51:40] systemd: Started actions.runner...omniroute-113-6.service   ← a fome levou o Listener junto

O que muda

concurrency no job build, em duas faixas:

grupo quem efeito
heavy-build-main pushes em main nunca disputa com PR; uma release nunca fica atrás de tráfego de PR
heavy-build-pr pull requests serializam entre si

cancel-in-progress: false — um build em execução nunca é morto por um mais novo.

O trade-off, dito claramente

A regra do GitHub por grupo é 1 em execução + 1 pendente; pendentes mais antigos são cancelados. Numa rajada, o 3º build de PR aparece como "cancelled" e precisa de re-run. Um check de PR cancelado é re-executável; um build de main morto custa uma release.

O conserto definitivo continua sendo o split por label (omni-build em 2 runners, omni-light nos demais), que move a fila para o lado do runner sem cancelamentos — decisão de operador, registrada em docs/ops/RUNNER_BOX.md (#11893).

Validação

  • YAML válido; check:workflows --ratchet: zizmor inalterado (194)
  • O próprio CI desta PR roda na faixa heavy-build-pr

⚠️ base-red inherited: #11866 (Unit Tests (1/8) — allowlist da Alibaba, consertada em #11867)

The .113 box has 31 GB and a single next-build peaks at 14–16 GB RSS: one
build fits with room, two sit at the edge, three take the box down. On
2026-08-28 13:50Z the kernel OOM-killed main's next-build (15.7 GB) while a PR
build ran beside it — five Build jobs had been queued by a burst of PRs — and
the publish lost its artefact, which sends it into the 40-minute rebuild that
OOMs on its own (attempt 5 of this release).

Job-level concurrency on `build`, two lanes:

  heavy-build-main   pushes to main — never contended, never behind PR traffic
  heavy-build-pr     pull requests — serialize among themselves

cancel-in-progress stays false: a running build is never killed by a newer
one. GitHub's own rule for a group is one running + one pending, older pendings
cancelled — so under a burst the third PR build shows "cancelled" and needs a
re-run. That is the trade-off, stated: a cancelled PR check is re-runnable; a
dead main build costs a release.

The proper fix remains a label split (omni-build on two runners, omni-light on
the rest) so the queue lives on the runner side without cancellations — an
operator decision recorded in docs/ops/RUNNER_BOX.md.
@github-actions

Copy link
Copy Markdown
Contributor

CI Coverage Report

  • Coverage job: skipped
  • PR test policy: success

Coverage artifact was not available for this run.

@diegosouzapw
diegosouzapw merged commit f564b64 into main Aug 28, 2026
47 of 49 checks passed
diegosouzapw added a commit that referenced this pull request Aug 28, 2026
…e pipeline fixes

Brings e4683cd (#11867 Alibaba allowlist time bomb), 09de69e (#11891
config expiry detector), e71be03 (#11893 runner janitor), 9dc8eab
(#11895 provenance × self-hosted lint) and f564b64 (#11901 heavy-build
lanes). main is already an ancestor of this branch (v3.8.50 sync-back), so the
merge is exactly these five commits.

# Conflicts:
#	tests/unit/alibaba-free-tier-allowlist.test.ts
diegosouzapw added a commit that referenced this pull request Aug 28, 2026
…11932)

The .113 box (31 GB) holds one next-build (14–16 GB RSS) comfortably and two
at the edge; on 2026-08-28 the kernel killed main's build twice while PR
builds ran beside it. Labels are the runner-side cap: only omniroute-113-5
and omniroute-113-6 carry omni-build (added through the runners API, no
re-registration), and every job that runs a next build — ci.yml build,
npm-publish.yml publish, both nightly-release-green validations — now asks
for that label. A third heavy job queues on GitHub instead of racing for
memory. The six other runners keep omni-release and no longer take builds.
Pairs with the heavy-build-* concurrency lanes (#11901); documented in
docs/ops/RUNNER_BOX.md.
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…osouzapw#11901)

The .113 box has 31 GB and a single next-build peaks at 14–16 GB RSS: one
build fits with room, two sit at the edge, three take the box down. On
2026-08-28 13:50Z the kernel OOM-killed main's next-build (15.7 GB) while a PR
build ran beside it — five Build jobs had been queued by a burst of PRs — and
the publish lost its artefact, which sends it into the 40-minute rebuild that
OOMs on its own (attempt 5 of this release).

Job-level concurrency on `build`, two lanes:

  heavy-build-main   pushes to main — never contended, never behind PR traffic
  heavy-build-pr     pull requests — serialize among themselves

cancel-in-progress stays false: a running build is never killed by a newer
one. GitHub's own rule for a group is one running + one pending, older pendings
cancelled — so under a burst the third PR build shows "cancelled" and needs a
re-run. That is the trade-off, stated: a cancelled PR check is re-runnable; a
dead main build costs a release.

The proper fix remains a label split (omni-build on two runners, omni-light on
the rest) so the queue lives on the runner side without cancellations — an
operator decision recorded in docs/ops/RUNNER_BOX.md.
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…e pipeline fixes

Brings 342ca95 (diegosouzapw#11867 Alibaba allowlist time bomb), ce53544 (diegosouzapw#11891
config expiry detector), 8b8d84d (diegosouzapw#11893 runner janitor), 9e87941
(diegosouzapw#11895 provenance × self-hosted lint) and 2fe1c30 (diegosouzapw#11901 heavy-build
lanes). main is already an ancestor of this branch (v3.8.50 sync-back), so the
merge is exactly these five commits.

# Conflicts:
#	tests/unit/alibaba-free-tier-allowlist.test.ts
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…iegosouzapw#11932)

The .113 box (31 GB) holds one next-build (14–16 GB RSS) comfortably and two
at the edge; on 2026-08-28 the kernel killed main's build twice while PR
builds ran beside it. Labels are the runner-side cap: only omniroute-113-5
and omniroute-113-6 carry omni-build (added through the runners API, no
re-registration), and every job that runs a next build — ci.yml build,
npm-publish.yml publish, both nightly-release-green validations — now asks
for that label. A third heavy job queues on GitHub instead of racing for
memory. The six other runners keep omni-release and no longer take builds.
Pairs with the heavy-build-* concurrency lanes (diegosouzapw#11901); documented in
docs/ops/RUNNER_BOX.md.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant