fix(codex): lift nested child cooldowns on parent clear and on snapshot headroom - #12951
Conversation
…lear (#12788) Validado numa worktree combinada com a onda de persistência desta leva sobre `release/v3.8.51`: check-file-size e check-changelog-integrity OK, typecheck:core limpo, check-api-typecheck OK (289, dentro da baseline), 70 testes focados no runner Node e 1 no vitest, todos verdes. Persistir `NaN`/`Infinity` numa coluna TEXT envenena toda leitura futura, e um timestamp já expirado sobrescrevendo uma linha viva é pior que não escrever nada. O guard fundido na cabeça da função cobre os dois sem tocar leitores nem o caminho de clear. Os 6 casos do teste incluem o que mais importa: `null` continua limpando, e uma escrita expirada não derruba um cooldown ativo. Nota: 3 deles falham antes do guard, como você registrou. **Integração:** este arquivo colidiu com o #12951, que também guarda a cabeça de `setConnectionRateLimitUntil` — lá o `null` vira caminho de clear que também remove os cooldowns filhos do Codex. Os dois compõem: trata-se o `null` primeiro (clear + return), e o seu guard de finitude/expiração passa a valer para os não-nulos. Ambos preservados.
46e59d0 to
d84f6fc
Compare
|
Thanks for addressing both the manual-clear and restored-quota paths. I added an independent affected-instance observation to #12860: fresh weekly usage showed 100% remaining while a legacy Could the regression coverage also exercise the reset-credit button's actual service path, rather than only the generic parent clear and direct snapshot writes? At A useful regression would seed a future legacy That would directly cover the operator-visible symptom: “reset succeeded and the quota bar refilled, but requests still cannot use the account.” Related alternative reconciliation work: #13073. |
…ared Dashboard "clear cooldown" and the CAS recovery path only null rate_limited_until. Codex still stores per-scope timestamps in provider_specific_data.codexScopeRateLimitedUntil, so dispatch keeps skipping the account after quota is back. Clear those nested maps on the same write as the parent column: PUT null or empty string, clearConnectionRateLimit, and the CAS UPDATE (json_remove in the same statement). A 200 quota observation only drops that child's leftover, not Spark's. Fixes diegosouzapw#12817 Signed-off-by: Minxi Hou <houminxi@gmail.com>
…headroom Quota preflight parks Codex scope cooldowns into provider_specific_data.codexScopeRateLimitedUntil with source 'fallback'. When a background quota poll or account test receives fresh usage data demonstrating headroom on that scope, the fallback-sourced cooldown was not lifted, leaving the scope blocked until the parked duration elapsed. Lift fallback-sourced scope cooldowns in saveQuotaSnapshot when all active windows for that scope report positive remaining percentage and are not exhausted. Authoritative 429 quota_reset timestamps remain protected until their reset time arrives. Fixes diegosouzapw#12860 Addresses diegosouzapw#12817 Signed-off-by: Minxi Hou <houminxi@gmail.com>
…ldown Operator-visible path is consumeCodexResetCredit after redeem, not a direct snapshot write. Seed a leftover fallback Codex child, mock a successful redeem plus healthy weekly usage, and assert the child becomes eligible. Spark stays parked. Failed redeem, incomplete refresh, and alreadyRedeemed without headroom keep the nested maps. Signed-off-by: Minxi Hou <houminxi@gmail.com>
…m lift Signed-off-by: Minxi Hou <houminxi@gmail.com>
d84f6fc to
af2ed01
Compare
|
Thanks — that was the operator-visible hole.
Added four cases on that path:
No extra production branch: the button already rode the snapshot path. Coverage now pins it so a filled quota bar after reset cannot leave the leftover fallback child parked. |
bfbd090
into
diegosouzapw:release/v3.8.51
…lear (diegosouzapw#12788) Validado numa worktree combinada com a onda de persistência desta leva sobre `release/v3.8.51`: check-file-size e check-changelog-integrity OK, typecheck:core limpo, check-api-typecheck OK (289, dentro da baseline), 70 testes focados no runner Node e 1 no vitest, todos verdes. Persistir `NaN`/`Infinity` numa coluna TEXT envenena toda leitura futura, e um timestamp já expirado sobrescrevendo uma linha viva é pior que não escrever nada. O guard fundido na cabeça da função cobre os dois sem tocar leitores nem o caminho de clear. Os 6 casos do teste incluem o que mais importa: `null` continua limpando, e uma escrita expirada não derruba um cooldown ativo. Nota: 3 deles falham antes do guard, como você registrou. **Integração:** este arquivo colidiu com o diegosouzapw#12951, que também guarda a cabeça de `setConnectionRateLimitUntil` — lá o `null` vira caminho de clear que também remove os cooldowns filhos do Codex. Os dois compõem: trata-se o `null` primeiro (clear + return), e o seu guard de finitude/expiração passa a valer para os não-nulos. Ambos preservados.
…ot headroom (diegosouzapw#12951) Validado numa worktree combinada com a onda de persistência desta leva sobre `release/v3.8.51`: check-file-size e check-changelog-integrity OK, typecheck:core limpo, check-api-typecheck OK (289). Revalidei o head atual mergeado com o tip: typecheck:core limpo e **36/36** entre `db-rate-limit-guard` e as suítes desta PR. Levantar o cooldown do escopo pai sem deixar os filhos aninhados presos é o miolo — um cooldown órfão em filho é invisível no dashboard e mantém a conexão fora de rota sem explicação. **Nota de coordenação:** o `setConnectionRateLimitUntil` colidiu com o diegosouzapw#12788 (guard contra timestamp não-finito ou já expirado), que mergeei nesta mesma onda. Eu tinha resolvido a integração na minha worktree, mas ao empurrar o push foi rejeitado — você já tinha empurrado `441fd44`, `f853ba5` e `af2ed01` com a integração feita, e a sua ordenação é equivalente à minha. Descartei a minha e mantive a sua; o crédito é seu inteiro. Fica o registro de que push rejeitado não é erro leve: se eu tivesse mergeado sem reler, teria levado a branch errada.
Bug Description
Two ways a Codex account stays skipped after its quota is actually back.
#12817 — Clearing a Codex cooldown from the dashboard (or the CAS recovery path) only nulls
rate_limited_until. Dispatch still readsproviderSpecificData.codexScopeRateLimitedUntil, so the account stays skipped.#12860 — Quota preflight parks a cooldown at the reported
next_reset_at. When a later snapshot shows the window recovered early, nothing removes that timestamp, so the account keeps getting skipped until the original deadline passes.Fixes #12817
Fixes #12860
Root Cause
Codex writes per-scope timestamps into JSON (
codex/spark) without always mirroring them onto the parent column. Every "clear cooldown" path only touched that column, and no path ran in the other direction: a fresh snapshot proving headroom had no way to retire a cooldown the preflight had written.Fix
Parent column → nested maps (#12817):
rateLimitedUntil: nullor""also drops the nested Codex mapsclearConnectionRateLimitnulls the parent column and the nested maps in one transactionclearConnectionErrorIfUnchanged) usesjson_removein the same UPDATESnapshot → nested maps (#12860):
liftCodexScopeCooldownOnHeadroom(id, scope)retires a scope cooldown when a snapshot proves headroom, inside one transactionsaveQuotaSnapshotcalls it after persisting a Codex row withremaining_percentage > 0andis_exhausted !== 1fallback-sourced cooldowns are eligible. A cooldown sourced from an upstream 429quota_resetis authoritative and is held until it expires, with a 30s grace window so local clock skew cannot release it earlyhasCodexScopeCooldownshort-circuits before the scope-wide snapshot read, so the common (clean) connection pays one indexed lookup and no backupNon-Codex connections are unchanged. Crash recovery still keeps future nested timestamps (they are live quota windows).
How to Verify
Parent column path:
{ "rateLimitedUntil": null }(or click Clear cooldown).getCodexChildCooldownfor both models returns null;unrelatedJSON is still there.gpt-5.5clears only the Codex child; Spark remains limited.Snapshot path:
fallbackcooldown on thecodexscope with a reset an hour out.remaining_percentage: 100,is_exhausted: 0.codexentry is gone; a Spark cooldown on the same connection is untouched.quota_reset— it survives.Test Plan
tests/unit/codex-scope-cooldown-clear-12817.test.ts— 14 tests covering both issues (PUT null, PUT empty string,clearConnectionRateLimit, CAS, 200 observation, non-Codex, snapshot lift,quota_resetprotection, skew grace boundary on both sides, partial-window hold, Spark isolation, short-circuit probe)codex-account-cooldown-write,codex-account-helpers,quota-connection-recoverystill pass (40/40 combined)typecheck:core,check:cycles(427 files),check-docs-synccleanRisk Assessment
Low. Nested maps are Codex-only, and the snapshot lift is gated on provider
codexplus afallbacksource. Spark sibling state is preserved on a Codex-only 200 and on a Codex-only snapshot. The write runs in one transaction that re-reads the row under lock, so a concurrent cooldown write cannot be lost.providers.tsshrank 1186 → 1137 (CAS moved torateLimit.ts).