fix(cli): supervisor restarts on spontaneous exit-0 (OOM cgroup) + waits for port before respawn (#4425) - #4578
Merged
Conversation
Contributor
|
Warning You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again! |
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
…its for port before respawn (#4425)
diegosouzapw
force-pushed
the
fix/4425-supervisor
branch
from
June 21, 2026 22:50
3f64417 to
99250c7
Compare
KooshaPari
added a commit
to KooshaPari/OmniRoute
that referenced
this pull request
Jun 21, 2026
The Fast Quality Gates lint check was failing on every PR with '4 arquivos cresceram alem do cap': src/lib/db/core.ts, src/lib/usage/providerLimits.ts, src/shared/constants/providers.ts, open-sse/services/usage.ts. The frozen baselines in config/quality/file-size-baseline.json were last set at v3.8.30 and had drifted past the cap=800 due to legitimate feature growth from PR diegosouzapw#4381 (combos split), PR diegosouzapw#4433 (cluster opt-in profiles), and PR diegosouzapw#4480 (vacuum scheduler). This commit rebaselines those 4 frozen entries to their current actual line count (+2 buffer to cover wc -l's off-by-one and any stray edits during review). It does NOT change the cap=800 for new files, nor does it shrink any of the 4 monoliths. Structural shrink of these files is tracked separately in diegosouzapw#3501 (QG v2 chatCore split continuation). This rebaseline just restores green CI until those structural refactors land. Files changed: 1 (config/quality/file-size-baseline.json) - src/lib/db/core.ts: frozen 624 -> 781 (was +157 past cap=800...wait) Actually frozen was 624 vs cap=800, so core.ts was 157 lines UNDER cap. The drift is in the 4 files whose actuals grew past their frozen values. Verification: - node scripts/check/check-file-size.mjs -> '[file-size] OK -- 103 arquivos congelados, cap 800 para novos (2710 arquivos verificados)' - node scripts/check/check-env-doc-sync.mjs -> 'Env / docs contract is in sync' - node scripts/check/check-db-rules.mjs -> 'OK (85 modulos db/, 57 re-exportados, 28 intencionalmente-internos; 2 leituras de DB externo permitidas)' Unblocks every open PR currently stuck on Fast Quality Gates (diegosouzapw#4571, diegosouzapw#4576, diegosouzapw#4577, diegosouzapw#4578 + this PR's own branch).
This was referenced Jun 22, 2026
Merged
tkgo11
pushed a commit
to tkgo11/OmniRoute
that referenced
this pull request
Sep 23, 2026
…its for port before respawn (diegosouzapw#4425) (diegosouzapw#4578) Co-authored-by: Diego Rodrigues de Sa e Souza <diego.souza@cdwasolutions.com.br>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #4425 (partial — the recovery slice)
Problem
Under load the gateway crash-looped: (1) a systemd
MemoryMaxcgroup kill reports a clean exit (code 0), which the supervisor treated as an intentional stop and exited — leaving the gateway dead withRestart=on-failure; (2) it respawned immediately after a crash, before the OS released the listen socket → an EADDRINUSE cascade that exhausted the restart budget; (3) the 30s reset window dropped the crash counter too fast.Fix
New
bin/cli/runtime/supervisorPolicy.mjs(pure + unit-testable):shouldExitInsteadOfRestart(only an intentional shutdown exits; a spontaneous code-0 now restarts),RESTART_RESET_MS30s→60s,DEFAULT_MAX_RESTARTS2→3,computeRestartDelayMs, andisPortFree/waitUntilPortFree.processSupervisor.handleExitnow restarts on a spontaneous code-0 and waits (bounded) for the port to free up before respawning.This is the recovery slice — it does not fix the underlying memory growth (tracked under the #4041/#4380 OOM work); it stops a crashed/OOM'd process from staying down or cascading on EADDRINUSE.
Validation (TDD)
New
tests/unit/supervisor-policy-4425.test.ts(5/5: restart-on-code-0, constants, backoff, port-free detection + wait). Updated the existingcli-process-supervisor.test.tsto match the corrected behavior (code-0 restarts; 60s reset window). 16/16 green; lint clean.