Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
42 changes: 36 additions & 6 deletions .github/workflows/nightly-release-green.yml
Original file line number Diff line number Diff line change
Expand Up @@ -76,9 +76,24 @@ jobs:
# says "~30-50min idle, up to ~80min under load", so it alone can exceed what was
# left. Three earlier attempts died the same way.
#
# The job declared no budget, so it inherited one. An explicit ceiling makes the
# budget a decision instead of an accident, and a sweep that overruns THIS says so
# as a timeout rather than as an opaque exit 143.
# An explicit ceiling makes the budget a decision instead of an accident, and a
# sweep that overruns THIS says so as a timeout rather than as an opaque exit 143.
#
# It is NOT what kills the sweep. That was the theory when this line was added, and
# run 35496465106 disproved it: dispatched on 450d6586f, which already carried
# `timeout-minutes: 180`, and killed at 08:17:37 — 60m11s after the job started, the
# same mark as the three attempts before it. Every static and drift gate had passed
# by 07:23:23; the four serial suites then ran for 54 minutes with no output and the
# process took SIGTERM. Run 35468579833 recorded the reason in words: "The runner has
# received a shutdown signal".
#
# So the hosted runner stops at ~60 minutes, and the suites' own ceilings
# (80 + 15 + 40 + 20 = 155 min worst case, serial) cannot fit inside that. Two real
# options, neither of which a workflow edit can reach:
# - set the USE_VPS_RUNNER repository variable for a release window, which this
# workflow already honours and which moves the sweep off the hosted runner;
# - shard the slow suites across jobs so no single job needs more than an hour.
# Raising this number again will not help; measure before trying.
timeout-minutes: 180
env:
JWT_SECRET: ci-nightly-secret-with-sufficient-length-for-validation
Expand Down Expand Up @@ -293,9 +308,24 @@ jobs:
# says "~30-50min idle, up to ~80min under load", so it alone can exceed what was
# left. Three earlier attempts died the same way.
#
# The job declared no budget, so it inherited one. An explicit ceiling makes the
# budget a decision instead of an accident, and a sweep that overruns THIS says so
# as a timeout rather than as an opaque exit 143.
# An explicit ceiling makes the budget a decision instead of an accident, and a
# sweep that overruns THIS says so as a timeout rather than as an opaque exit 143.
#
# It is NOT what kills the sweep. That was the theory when this line was added, and
# run 35496465106 disproved it: dispatched on 450d6586f, which already carried
# `timeout-minutes: 180`, and killed at 08:17:37 — 60m11s after the job started, the
# same mark as the three attempts before it. Every static and drift gate had passed
# by 07:23:23; the four serial suites then ran for 54 minutes with no output and the
# process took SIGTERM. Run 35468579833 recorded the reason in words: "The runner has
# received a shutdown signal".
#
# So the hosted runner stops at ~60 minutes, and the suites' own ceilings
# (80 + 15 + 40 + 20 = 155 min worst case, serial) cannot fit inside that. Two real
# options, neither of which a workflow edit can reach:
# - set the USE_VPS_RUNNER repository variable for a release window, which this
# workflow already honours and which moves the sweep off the hosted runner;
# - shard the slow suites across jobs so no single job needs more than an hour.
# Raising this number again will not help; measure before trying.
timeout-minutes: 180
env:
JWT_SECRET: ci-nightly-secret-with-sufficient-length-for-validation
Expand Down
23 changes: 21 additions & 2 deletions audit/FINAL_THREE_AGENT_REVIEW.md
Original file line number Diff line number Diff line change
Expand Up @@ -61,8 +61,27 @@ existe para cobrir isso, mas em `push` roda `--quick`, que pula as suítes.
Três disparos completos morreram idênticos: `exit 143`, aos 60 minutos, **sem nenhuma saída** — o passo
redirecionava o log para arquivo e só o imprimia com um `cat` final que nunca era alcançado.
**[#70](https://github.com/LMPrado-DZ23/OmniRoute/pull/70)** faz o log sair ao vivo; isso não conserta a
morte, conserta a cegueira. **Este HIGH permanece em aberto** e é o único item que separa esta linha do
portão CRITICAL 0 / HIGH 0.
morte, conserta a cegueira.

Com o log visível, a causa ficou clara — e ela **não** era a que eu supus. A
**[#76](https://github.com/LMPrado-DZ23/OmniRoute/pull/76)** deu `timeout-minutes: 180` aos dois jobs,
partindo da hipótese de que o job herdava um teto por não declarar nenhum. A execução
[35496465106](https://github.com/LMPrado-DZ23/OmniRoute/actions/runs/35496465106) refutou isso: foi
disparada em `450d6586f`, que **já continha** os 180 minutos, e morreu às 08:17:37 — **60m11s** depois
de começar, a mesma marca das três anteriores. Todos os gates estáticos e de deriva passaram até
07:23:23; as quatro suítes seriais então rodaram 54 minutos sem uma linha de saída e o processo levou
SIGTERM. A execução 35468579833 já tinha registrado o motivo em palavras: _"The runner has received a
shutdown signal"_.

Ou seja: **o runner hospedado para em ~60 minutos**, e os tetos das próprias suítes somam 155 minutos
no pior caso em modo serial. Duas saídas reais, nenhuma alcançável por edição de workflow:

- ligar a variável `USE_VPS_RUNNER` numa janela de release — o workflow já a honra e tira a varredura
do runner hospedado;
- fatiar as suítes lentas em jobs separados, de modo que nenhum precise de mais de uma hora.

**Este HIGH permanece em aberto** e é o único item que separa esta linha do portão CRITICAL 0 / HIGH 0.
O que mudou é que agora se sabe por quê, e que aumentar o número de novo não resolve.

**B-H1 — PR de fork executava no runner LAN persistente do mantenedor.**
`quality.yml:553` selecionava o pool `self-hosted` sem a cláusula de origem própria que `ci.yml:650`
Expand Down
Loading