Skip to content

fix(ci): run the build-bearing nightly jobs on the box's light pool (#11965) - #11967

Merged
diegosouzapw merged 1 commit into
release/v3.8.51from
fix/ci-11965-nightly-jobs-omni-light
Aug 29, 2026
Merged

diegosouzapw merged 1 commit into
release/v3.8.51from
fix/ci-11965-nightly-jobs-omni-light

Conversation

@diegosouzapw

Copy link
Copy Markdown
Owner

O que muda

job antes depois
nightly-schemathesis → Schemathesis ubuntu-latest (passa por minutos) [self-hosted, omni-light]
nightly-llm-security → promptfoo + garak ubuntu-latest (2 jobs mortos em 28/08, sem step falho = VM derrubada) [self-hosted, omni-light]
nightly-resilience → axe a11y (webServer auto-builda o Next) ubuntu-latest (4/4 vermelho) [self-hosted, omni-light]

Todos com fallback para hospedado quando USE_VPS_RUNNER está desligado (mesmo padrão do nightly-release-green). Rodam 1×/dia entre 04:00 e 06:00 UTC, com a caixa ociosa. k6, heap-growth e chaos continuam hospedados (passam).

Frota (feito hoje, reversível)

  • Label omni-light em omniroute-113 e omniroute-113-2 (API, sem re-registro): jobs de ~6 GB.
  • omniroute-113-3/-4/-7/-8 desligados (systemctl disable --now; enable --now traz de volta) — só o Build da main e os nightlies usam a caixa, 8 listeners estavam ociosos e cada um é um inquilino potencial de 14 GB.
  • Janitor com teto MAX_ACTIVE_RUNNERS=4.
  • Pico teórico 2 pesados + 2 leves ≈ 42 GB (31 GB RAM + 16 GB swap). Mais RAM na VM Proxmox é a alavanca real — vira 3 pesados + 2 leves só mudando labels.

Docs: docs/ops/RUNNER_BOX.md atualizado. check:workflows --ratchet 194/194; check-workflows.test.ts 32/32; backend-only-smoke-workflows.test.ts 6/6; docs-sync PASS.

Closes #11965

…11965)

Four nightly jobs run a backend-only `next build` on ubuntu-latest (7 GB):
Schemathesis, promptfoo injection guard, garak probes and the axe a11y suite
(self-building webServer). On release/v3.8.51 three of them died with the hosted
VM shutdown signature and nobody saw it — nightlies have no audience — and the
fourth passes by a margin of minutes. They now target [self-hosted, omni-light]
(hosted fallback when USE_VPS_RUNNER is off), a new two-listener label on the .113
box for jobs that need ~6 GB, not the 14-16 GB of a full build; they run once a day
in the 04:00-06:00 UTC window, when the box is idle.

Fleet reshaped the same day and documented in docs/ops/RUNNER_BOX.md: 4 active
OmniRoute listeners (omniroute-113-5/-6 omni-build, omniroute-113/-2 omni-light),
omniroute-113-3/-4/-7/-8 disabled (systemctl enable --now brings one back), janitor
ceiling MAX_ACTIVE_RUNNERS=4. The remaining headroom limit is the VM's 31 GB of RAM
(2 heavy + 2 light ≈ 42 GB peak, inside the 16 GB swap); more RAM on the Proxmox VM
is the lever that turns the label ceilings into 3 heavy + 2 light.

check:workflows --ratchet unchanged (194/194); check-workflows and
backend-only-smoke-workflows suites pass; docs-sync PASS.
@diegosouzapw
diegosouzapw merged commit 7517106 into release/v3.8.51 Aug 29, 2026
20 checks passed
@diegosouzapw
diegosouzapw deleted the fix/ci-11965-nightly-jobs-omni-light branch August 29, 2026 02:26
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…iegosouzapw#11965) (diegosouzapw#11967)

Four nightly jobs run a backend-only `next build` on ubuntu-latest (7 GB):
Schemathesis, promptfoo injection guard, garak probes and the axe a11y suite
(self-building webServer). On release/v3.8.51 three of them died with the hosted
VM shutdown signature and nobody saw it — nightlies have no audience — and the
fourth passes by a margin of minutes. They now target [self-hosted, omni-light]
(hosted fallback when USE_VPS_RUNNER is off), a new two-listener label on the .113
box for jobs that need ~6 GB, not the 14-16 GB of a full build; they run once a day
in the 04:00-06:00 UTC window, when the box is idle.

Fleet reshaped the same day and documented in docs/ops/RUNNER_BOX.md: 4 active
OmniRoute listeners (omniroute-113-5/-6 omni-build, omniroute-113/-2 omni-light),
omniroute-113-3/-4/-7/-8 disabled (systemctl enable --now brings one back), janitor
ceiling MAX_ACTIVE_RUNNERS=4. The remaining headroom limit is the VM's 31 GB of RAM
(2 heavy + 2 light ≈ 42 GB peak, inside the 16 GB swap); more RAM on the Proxmox VM
is the lever that turns the label ceilings into 3 heavy + 2 light.

check:workflows --ratchet unchanged (194/194); check-workflows and
backend-only-smoke-workflows suites pass; docs-sync PASS.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

1 participant