From 47052837e94c8b72c184de8f33695b110b04967e Mon Sep 17 00:00:00 2001 From: zodyp Date: Sun, 20 Sep 2026 05:22:34 -0300 Subject: [PATCH] =?UTF-8?q?docs:=20`timeout-minutes`=20was=20not=20what=20?= =?UTF-8?q?kills=20the=20sweep=20=E2=80=94=20I=20was=20wrong?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit #76 gave both sweep jobs `timeout-minutes: 180` on the theory that the job, declaring no budget, inherited one. Run 35496465106 disproved it. That run was dispatched on 450d6586f, which ALREADY carried the 180, and it died at 08:17:37 — 60m11s after the job started, the same mark as the three attempts before it: 07:19:01 ✅ [hard] Typecheck (core) 07:22:03 ✅ [hard] ESLint errors — 0 07:22:39 ✅ [drift] Cognitive complexity — 1234 vs baseline 1437 07:23:23 ✅ [hard] Docs sync + fabricated-docs (strict) ▶ the four serial suites start 08:17:37 ##[error] Process completed with exit code 143 Every static and drift gate passes. The suites then ran 54 minutes with no output and the process took SIGTERM. Run 35468579833 had recorded the reason in words all along: "The runner has received a shutdown signal". So the hosted runner stops at ~60 minutes, and the suites' own ceilings (80 + 15 + 40 + 20 = 155 min worst case, serial) cannot fit inside it. Two real options, neither reachable by editing a workflow: - set USE_VPS_RUNNER for a release window — the workflow already honours it and it moves the sweep off the hosted runner; - shard the slow suites so no single job needs more than an hour. The `timeout-minutes` lines stay: an explicit budget is still better than an inherited one, and a sweep that overruns THAT reports a timeout instead of an opaque 143. But the comment claiming it was the cause is now a comment saying it is not, with the measurement, so nobody repeats my attempt. audit/FINAL_THREE_AGENT_REVIEW.md carries the same correction — the HIGH stays open, and what changed is that the reason is known. YAML parses, both sweep jobs still at 180; prettier clean [doc-links] PASS Co-Authored-By: Claude Opus 5 --- .github/workflows/nightly-release-green.yml | 42 ++++++++++++++++++--- audit/FINAL_THREE_AGENT_REVIEW.md | 23 ++++++++++- 2 files changed, 57 insertions(+), 8 deletions(-) diff --git a/.github/workflows/nightly-release-green.yml b/.github/workflows/nightly-release-green.yml index ae6342499be..dc0bb7c5a0a 100644 --- a/.github/workflows/nightly-release-green.yml +++ b/.github/workflows/nightly-release-green.yml @@ -76,9 +76,24 @@ jobs: # says "~30-50min idle, up to ~80min under load", so it alone can exceed what was # left. Three earlier attempts died the same way. # - # The job declared no budget, so it inherited one. An explicit ceiling makes the - # budget a decision instead of an accident, and a sweep that overruns THIS says so - # as a timeout rather than as an opaque exit 143. + # An explicit ceiling makes the budget a decision instead of an accident, and a + # sweep that overruns THIS says so as a timeout rather than as an opaque exit 143. + # + # It is NOT what kills the sweep. That was the theory when this line was added, and + # run 35496465106 disproved it: dispatched on 450d6586f, which already carried + # `timeout-minutes: 180`, and killed at 08:17:37 — 60m11s after the job started, the + # same mark as the three attempts before it. Every static and drift gate had passed + # by 07:23:23; the four serial suites then ran for 54 minutes with no output and the + # process took SIGTERM. Run 35468579833 recorded the reason in words: "The runner has + # received a shutdown signal". + # + # So the hosted runner stops at ~60 minutes, and the suites' own ceilings + # (80 + 15 + 40 + 20 = 155 min worst case, serial) cannot fit inside that. Two real + # options, neither of which a workflow edit can reach: + # - set the USE_VPS_RUNNER repository variable for a release window, which this + # workflow already honours and which moves the sweep off the hosted runner; + # - shard the slow suites across jobs so no single job needs more than an hour. + # Raising this number again will not help; measure before trying. timeout-minutes: 180 env: JWT_SECRET: ci-nightly-secret-with-sufficient-length-for-validation @@ -293,9 +308,24 @@ jobs: # says "~30-50min idle, up to ~80min under load", so it alone can exceed what was # left. Three earlier attempts died the same way. # - # The job declared no budget, so it inherited one. An explicit ceiling makes the - # budget a decision instead of an accident, and a sweep that overruns THIS says so - # as a timeout rather than as an opaque exit 143. + # An explicit ceiling makes the budget a decision instead of an accident, and a + # sweep that overruns THIS says so as a timeout rather than as an opaque exit 143. + # + # It is NOT what kills the sweep. That was the theory when this line was added, and + # run 35496465106 disproved it: dispatched on 450d6586f, which already carried + # `timeout-minutes: 180`, and killed at 08:17:37 — 60m11s after the job started, the + # same mark as the three attempts before it. Every static and drift gate had passed + # by 07:23:23; the four serial suites then ran for 54 minutes with no output and the + # process took SIGTERM. Run 35468579833 recorded the reason in words: "The runner has + # received a shutdown signal". + # + # So the hosted runner stops at ~60 minutes, and the suites' own ceilings + # (80 + 15 + 40 + 20 = 155 min worst case, serial) cannot fit inside that. Two real + # options, neither of which a workflow edit can reach: + # - set the USE_VPS_RUNNER repository variable for a release window, which this + # workflow already honours and which moves the sweep off the hosted runner; + # - shard the slow suites across jobs so no single job needs more than an hour. + # Raising this number again will not help; measure before trying. timeout-minutes: 180 env: JWT_SECRET: ci-nightly-secret-with-sufficient-length-for-validation diff --git a/audit/FINAL_THREE_AGENT_REVIEW.md b/audit/FINAL_THREE_AGENT_REVIEW.md index fb20b3a3eef..46357f18d7e 100644 --- a/audit/FINAL_THREE_AGENT_REVIEW.md +++ b/audit/FINAL_THREE_AGENT_REVIEW.md @@ -61,8 +61,27 @@ existe para cobrir isso, mas em `push` roda `--quick`, que pula as suítes. Três disparos completos morreram idênticos: `exit 143`, aos 60 minutos, **sem nenhuma saída** — o passo redirecionava o log para arquivo e só o imprimia com um `cat` final que nunca era alcançado. **[#70](https://github.com/LMPrado-DZ23/OmniRoute/pull/70)** faz o log sair ao vivo; isso não conserta a -morte, conserta a cegueira. **Este HIGH permanece em aberto** e é o único item que separa esta linha do -portão CRITICAL 0 / HIGH 0. +morte, conserta a cegueira. + +Com o log visível, a causa ficou clara — e ela **não** era a que eu supus. A +**[#76](https://github.com/LMPrado-DZ23/OmniRoute/pull/76)** deu `timeout-minutes: 180` aos dois jobs, +partindo da hipótese de que o job herdava um teto por não declarar nenhum. A execução +[35496465106](https://github.com/LMPrado-DZ23/OmniRoute/actions/runs/35496465106) refutou isso: foi +disparada em `450d6586f`, que **já continha** os 180 minutos, e morreu às 08:17:37 — **60m11s** depois +de começar, a mesma marca das três anteriores. Todos os gates estáticos e de deriva passaram até +07:23:23; as quatro suítes seriais então rodaram 54 minutos sem uma linha de saída e o processo levou +SIGTERM. A execução 35468579833 já tinha registrado o motivo em palavras: _"The runner has received a +shutdown signal"_. + +Ou seja: **o runner hospedado para em ~60 minutos**, e os tetos das próprias suítes somam 155 minutos +no pior caso em modo serial. Duas saídas reais, nenhuma alcançável por edição de workflow: + +- ligar a variável `USE_VPS_RUNNER` numa janela de release — o workflow já a honra e tira a varredura + do runner hospedado; +- fatiar as suítes lentas em jobs separados, de modo que nenhum precise de mais de uma hora. + +**Este HIGH permanece em aberto** e é o único item que separa esta linha do portão CRITICAL 0 / HIGH 0. +O que mudou é que agora se sabe por quê, e que aumentar o número de novo não resolve. **B-H1 — PR de fork executava no runner LAN persistente do mantenedor.** `quality.yml:553` selecionava o pool `self-hosted` sem a cláusula de origem própria que `ci.yml:650`