Skip to content

fix(codex): lift nested child cooldowns on parent clear and on snapshot headroom - #12951

Merged
diegosouzapw merged 5 commits into
diegosouzapw:release/v3.8.51from
HouMinXi:fix/codex-scope-cooldown-12817
Sep 10, 2026
Merged

diegosouzapw merged 5 commits into
diegosouzapw:release/v3.8.51from
HouMinXi:fix/codex-scope-cooldown-12817

Conversation

@HouMinXi

@HouMinXi HouMinXi commented Sep 7, 2026 •

Copy link
Copy Markdown
Contributor

Bug Description

Two ways a Codex account stays skipped after its quota is actually back.

#12817 — Clearing a Codex cooldown from the dashboard (or the CAS recovery path) only nulls rate_limited_until. Dispatch still reads providerSpecificData.codexScopeRateLimitedUntil, so the account stays skipped.

#12860 — Quota preflight parks a cooldown at the reported next_reset_at. When a later snapshot shows the window recovered early, nothing removes that timestamp, so the account keeps getting skipped until the original deadline passes.

Fixes #12817
Fixes #12860

Root Cause

Codex writes per-scope timestamps into JSON (codex / spark) without always mirroring them onto the parent column. Every "clear cooldown" path only touched that column, and no path ran in the other direction: a fresh snapshot proving headroom had no way to retire a cooldown the preflight had written.

Fix

Parent column → nested maps (#12817):

  • PUT rateLimitedUntil: null or "" also drops the nested Codex maps
  • clearConnectionRateLimit nulls the parent column and the nested maps in one transaction
  • CAS recovery (clearConnectionErrorIfUnchanged) uses json_remove in the same UPDATE
  • A 200 quota observation drops that child's leftover only; Spark stays

Snapshot → nested maps (#12860):

  • liftCodexScopeCooldownOnHeadroom(id, scope) retires a scope cooldown when a snapshot proves headroom, inside one transaction
  • saveQuotaSnapshot calls it after persisting a Codex row with remaining_percentage > 0 and is_exhausted !== 1
  • Only fallback-sourced cooldowns are eligible. A cooldown sourced from an upstream 429 quota_reset is authoritative and is held until it expires, with a 30s grace window so local clock skew cannot release it early
  • Every window in the scope must report headroom: one exhausted window still justifies the cooldown even when a sibling reports full
  • hasCodexScopeCooldown short-circuits before the scope-wide snapshot read, so the common (clean) connection pays one indexed lookup and no backup

Non-Codex connections are unchanged. Crash recovery still keeps future nested timestamps (they are live quota windows).

How to Verify

Parent column path:

  1. Persist Codex + Spark child cooldowns on one connection.
  2. PUT { "rateLimitedUntil": null } (or click Clear cooldown).
  3. getCodexChildCooldown for both models returns null; unrelated JSON is still there.
  4. A 200 quota observation for gpt-5.5 clears only the Codex child; Spark remains limited.

Snapshot path:

  1. Park a fallback cooldown on the codex scope with a reset an hour out.
  2. Save a snapshot for that connection with remaining_percentage: 100, is_exhausted: 0.
  3. The codex entry is gone; a Spark cooldown on the same connection is untouched.
  4. Repeat with the cooldown sourced from quota_reset — it survives.

Test Plan

  • tests/unit/codex-scope-cooldown-clear-12817.test.ts — 14 tests covering both issues (PUT null, PUT empty string, clearConnectionRateLimit, CAS, 200 observation, non-Codex, snapshot lift, quota_reset protection, skew grace boundary on both sides, partial-window hold, Spark isolation, short-circuit probe)
  • Existing codex-account-cooldown-write, codex-account-helpers, quota-connection-recovery still pass (40/40 combined)
  • Injection, each behavior separately: drop the PUT strip → red; drop the snapshot call → red; remove the skew grace → red; make the short-circuit probe ignore scope → red. Restore → green.
  • typecheck:core, check:cycles (427 files), check-docs-sync clean

Risk Assessment

Low. Nested maps are Codex-only, and the snapshot lift is gated on provider codex plus a fallback source. Spark sibling state is preserved on a Codex-only 200 and on a Codex-only snapshot. The write runs in one transaction that re-reads the row under lock, so a concurrent cooldown write cannot be lost. providers.ts shrank 1186 → 1137 (CAS moved to rateLimit.ts).

HouMinXi added a commit to HouMinXi/OmniRoute that referenced this pull request Sep 7, 2026


Signed-off-by: Minxi Hou <houminxi@gmail.com>
@HouMinXi HouMinXi changed the title fix(codex): lift nested child cooldowns when the parent column is cleared fix(codex): lift nested child cooldowns on parent clear and on snapshot headroom Sep 8, 2026
diegosouzapw pushed a commit that referenced this pull request Sep 8, 2026
…lear (#12788)

Validado numa worktree combinada com a onda de persistência desta leva sobre `release/v3.8.51`: check-file-size e check-changelog-integrity OK, typecheck:core limpo, check-api-typecheck OK (289, dentro da baseline), 70 testes focados no runner Node e 1 no vitest, todos verdes.

Persistir `NaN`/`Infinity` numa coluna TEXT envenena toda leitura futura, e um timestamp já expirado sobrescrevendo uma linha viva é pior que não escrever nada. O guard fundido na cabeça da função cobre os dois sem tocar leitores nem o caminho de clear.

Os 6 casos do teste incluem o que mais importa: `null` continua limpando, e uma escrita expirada não derruba um cooldown ativo. Nota: 3 deles falham antes do guard, como você registrou.

**Integração:** este arquivo colidiu com o #12951, que também guarda a cabeça de `setConnectionRateLimitUntil` — lá o `null` vira caminho de clear que também remove os cooldowns filhos do Codex. Os dois compõem: trata-se o `null` primeiro (clear + return), e o seu guard de finitude/expiração passa a valer para os não-nulos. Ambos preservados.
HouMinXi added a commit to HouMinXi/OmniRoute that referenced this pull request Sep 8, 2026


Signed-off-by: Minxi Hou <houminxi@gmail.com>
@HouMinXi
HouMinXi force-pushed the fix/codex-scope-cooldown-12817 branch from 46e59d0 to d84f6fc Compare September 8, 2026 17:52
@xiaoyaner0201

Copy link
Copy Markdown
Contributor

Thanks for addressing both the manual-clear and restored-quota paths. I added an independent affected-instance observation to #12860: fresh weekly usage showed 100% remaining while a legacy fallback Codex child cooldown kept the account excluded; removing only the stale child-map entries restored real traffic without a restart.

Could the regression coverage also exercise the reset-credit button's actual service path, rather than only the generic parent clear and direct snapshot writes?

At consumeCodexResetCredit on the current release, a recognized consume outcome invalidates the quota cache and calls fetchAndPersistProviderLimits. It does not directly invoke the parent clear path. The snapshot reconciliation in this patch looks relevant, but I have not run this PR and cannot claim that the whole button path passes.

A useful regression would seed a future legacy fallback child cooldown, mock a successful reset redemption plus fresh healthy weekly usage, invoke consumeCodexResetCredit, then assert the persisted Codex child is eligible without another manual clear or restart. It should preserve an unrelated Spark cooldown and retain protection on failed redemption, incomplete refresh, or a newer concurrent cooldown. An alreadyRedeemed retry should likewise rely on verified fresh quota evidence rather than unconditionally clearing current limits.

That would directly cover the operator-visible symptom: “reset succeeded and the quota bar refilled, but requests still cannot use the account.” Related alternative reconciliation work: #13073.

…ared

Dashboard "clear cooldown" and the CAS recovery path only null
rate_limited_until. Codex still stores per-scope timestamps in
provider_specific_data.codexScopeRateLimitedUntil, so dispatch keeps
skipping the account after quota is back.

Clear those nested maps on the same write as the parent column: PUT
null or empty string, clearConnectionRateLimit, and the CAS UPDATE
(json_remove in the same statement). A 200 quota observation only
drops that child's leftover, not Spark's.

Fixes diegosouzapw#12817

Signed-off-by: Minxi Hou <houminxi@gmail.com>


Signed-off-by: Minxi Hou <houminxi@gmail.com>
…headroom

Quota preflight parks Codex scope cooldowns into
provider_specific_data.codexScopeRateLimitedUntil with source 'fallback'.
When a background quota poll or account test receives fresh usage data
demonstrating headroom on that scope, the fallback-sourced cooldown was
not lifted, leaving the scope blocked until the parked duration elapsed.

Lift fallback-sourced scope cooldowns in saveQuotaSnapshot when all
active windows for that scope report positive remaining percentage and
are not exhausted. Authoritative 429 quota_reset timestamps remain
protected until their reset time arrives.

Fixes diegosouzapw#12860
Addresses diegosouzapw#12817

Signed-off-by: Minxi Hou <houminxi@gmail.com>
…ldown

Operator-visible path is consumeCodexResetCredit after redeem, not a
direct snapshot write. Seed a leftover fallback Codex child, mock a
successful redeem plus healthy weekly usage, and assert the child
becomes eligible. Spark stays parked. Failed redeem, incomplete
refresh, and alreadyRedeemed without headroom keep the nested maps.

Signed-off-by: Minxi Hou <houminxi@gmail.com>
…m lift

Signed-off-by: Minxi Hou <houminxi@gmail.com>
@HouMinXi
HouMinXi force-pushed the fix/codex-scope-cooldown-12817 branch from d84f6fc to af2ed01 Compare September 10, 2026 09:19
@HouMinXi

Copy link
Copy Markdown
Contributor Author

Thanks — that was the operator-visible hole.

consumeCodexResetCredit already invalidated the quota cache and called fetchAndPersistProviderLimits after a recognized redeem. That live-fetch writes setQuotaCache, which persists snapshots and is the same headroom lift as a dashboard usage refresh. The missing piece was a regression that actually invoked the button function instead of writing a snapshot by hand.

Added four cases on that path:

  • leftover fallback Codex child + successful redeem + healthy weekly usage → Codex child eligible, Spark child still parked
  • alreadyRedeemed without headroom (used_percent: 100) → nested maps stay
  • failed consume (HTTP 409 noCredit) → nested maps stay
  • incomplete usage refresh (usage endpoint 500; getCodexUsage swallows that into {message} and still returns from consume) → nested maps stay

No extra production branch: the button already rode the snapshot path. Coverage now pins it so a filled quota bar after reset cannot leave the leftover fallback child parked.

Pushed on this PR (af2ed0194). Related: #12860, #13073.

@diegosouzapw
diegosouzapw merged commit bfbd090 into diegosouzapw:release/v3.8.51 Sep 10, 2026
8 of 16 checks passed
@HouMinXi
HouMinXi deleted the fix/codex-scope-cooldown-12817 branch September 16, 2026 14:00
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…lear (diegosouzapw#12788)

Validado numa worktree combinada com a onda de persistência desta leva sobre `release/v3.8.51`: check-file-size e check-changelog-integrity OK, typecheck:core limpo, check-api-typecheck OK (289, dentro da baseline), 70 testes focados no runner Node e 1 no vitest, todos verdes.

Persistir `NaN`/`Infinity` numa coluna TEXT envenena toda leitura futura, e um timestamp já expirado sobrescrevendo uma linha viva é pior que não escrever nada. O guard fundido na cabeça da função cobre os dois sem tocar leitores nem o caminho de clear.

Os 6 casos do teste incluem o que mais importa: `null` continua limpando, e uma escrita expirada não derruba um cooldown ativo. Nota: 3 deles falham antes do guard, como você registrou.

**Integração:** este arquivo colidiu com o diegosouzapw#12951, que também guarda a cabeça de `setConnectionRateLimitUntil` — lá o `null` vira caminho de clear que também remove os cooldowns filhos do Codex. Os dois compõem: trata-se o `null` primeiro (clear + return), e o seu guard de finitude/expiração passa a valer para os não-nulos. Ambos preservados.
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…ot headroom (diegosouzapw#12951)

Validado numa worktree combinada com a onda de persistência desta leva sobre `release/v3.8.51`: check-file-size e check-changelog-integrity OK, typecheck:core limpo, check-api-typecheck OK (289). Revalidei o head atual mergeado com o tip: typecheck:core limpo e **36/36** entre `db-rate-limit-guard` e as suítes desta PR.

Levantar o cooldown do escopo pai sem deixar os filhos aninhados presos é o miolo — um cooldown órfão em filho é invisível no dashboard e mantém a conexão fora de rota sem explicação.

**Nota de coordenação:** o `setConnectionRateLimitUntil` colidiu com o diegosouzapw#12788 (guard contra timestamp não-finito ou já expirado), que mergeei nesta mesma onda. Eu tinha resolvido a integração na minha worktree, mas ao empurrar o push foi rejeitado — você já tinha empurrado `441fd44`, `f853ba5` e `af2ed01` com a integração feita, e a sua ordenação é equivalente à minha. Descartei a minha e mantive a sua; o crédito é seu inteiro. Fica o registro de que push rejeitado não é erro leve: se eu tivesse mergeado sem reler, teria levado a branch errada.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

3 participants