Skip to content

feat(quota): opt-in auto-ping to keep Codex quota windows warm (#6977) - #6995

Merged
diegosouzapw merged 3 commits into
release/v3.8.49from
feat/6977-quota-auto-ping
Jul 17, 2026
Merged

diegosouzapw merged 3 commits into
release/v3.8.49from
feat/6977-quota-auto-ping

Conversation

@diegosouzapw

Copy link
Copy Markdown
Owner

Summary

Ref #6977 — Codex's rolling "session" quota window only starts counting down once a request lands inside it, so an idle connection's window keeps sliding forward and the first real request after a long idle period pays for the whole warm-up latency. This adds a strictly opt-in, per-connection scheduler (default OFF) that watches an enabled connection's reported resetAt and, once it slides forward, fires one tiny non-billed-model request through the real Codex executor to keep the window warm.

  • New in-process scheduler src/lib/services/quotaAutoPing.ts, fully dependency-injected (settings, DB, credential refresh, usage fetch, executor, circuit breaker) and clock-injectable for deterministic tests. Reimplemented in TS from the shipped 9router src/shared/services/quotaAutoPing.js (Codex half only — Antigravity has no upstream reference for its 2-bucket reset shape).
  • Migration 123 (idempotent): last_ping_at / last_pinged_reset_key on provider_connections, so the scheduler never re-pings the same reset window twice.
  • Settings: codexAutoPing.connections map, default {} (nobody opted in — every ping burns a small amount of real quota), validated by the shared Zod schema (src/shared/validation/settingsSchemas.ts).
  • Respects the existing 3-layer resilience model: skips a connection whose provider circuit breaker is open (getCircuitBreaker("codex").canExecute()) or whose rateLimitedUntil cooldown is active, and applies its own 15-minute failure cooldown after a failed ping.
  • Wired into server boot (src/instrumentation-node.ts, inside the existing background-services gate) alongside the other schedulers.
  • UI: new per-connection toggle in Settings → AI → "Codex Quota Auto-Ping", with an explicit codexAutoPingWarning tooltip ("consumes real quota"). Lists OAuth Codex connections from /api/providers and PATCHes /api/settings. i18n keys added to all 43 locales (EN authored by hand; others filled from the EN fallback string as a starting point, pending real translation via the normal i18n pipeline).

Scope note (Phase 1 of #6977)

This PR ships the Codex half only, per the analyzed plan. Antigravity auto-ping (2 quota buckets, no shipped upstream reference) is intentionally out of scope and tracked as a follow-up against #6977 — not closing the issue from this PR.

Test plan

  • 15 new deterministic node:test unit tests (tests/unit/quota-auto-ping.test.ts): setting absent (no-op), first resetAt observation (cache-only, no ping), reset-slide ping, stable-reset no-op, min-ping-interval dedupe, same-resetKey dedupe across clock drift, session-quota-exhausted skip, blocking-quota-exhausted skip, non-OAuth skip, circuit-breaker-open skip, connection-cooldown skip, failure-cooldown skip, failed-ping bookkeeping (no DB write, failure cached), the real executor call shape (model/body/credentials), and credential-refresh-failure handling.
  • npm run typecheck:core — clean
  • npx eslint --suppressions-location config/quality/eslint-suppressions.json <changed files> — 0 errors, 0 warnings
  • node scripts/check/check-file-size.mjs — no new/growing violations (3 pre-existing base-red violations on this branch tip are unrelated files I never touched: ProxyRegistryManager.tsx, tokenHealthCheck.ts, open-sse/utils/stream.ts)
  • node scripts/check/check-complexity.mjs — 2056 == baseline 2056 (clean after extracting isPingCandidateBlocked / shouldSendPing / refreshConnectionForPing / pingProviderConnections helpers to keep every function ≤15 cyclomatic complexity)
  • node scripts/check/check-cognitive-complexity.mjs — 890 == baseline 890
  • node scripts/check/check-changelog-integrity.mjs — OK, no base bullets lost vs origin/release/v3.8.47
  • node --import tsx/esm --test tests/unit/quota-auto-ping.test.ts tests/unit/db-providers-crud.test.ts tests/unit/settings-transform-schema.test.ts tests/unit/settings-schema-routing-strategies.test.ts — all green, confirms the new provider_connections columns and settings schema entry don't regress existing CRUD/settings coverage

Codex's rolling "session" quota window only starts counting down once a
request lands inside it, so an idle connection's window keeps sliding
forward and the first real request after a long idle period pays for the
whole warm-up latency. This adds a strictly opt-in, per-connection
scheduler (default OFF) that watches an enabled connection's reported
resetAt and, once it slides forward, fires one tiny non-billed-model
request through the real Codex executor to keep the window warm.

- New in-process scheduler (src/lib/services/quotaAutoPing.ts), fully
  dependency-injected (settings, DB, credential refresh, usage fetch,
  executor, circuit breaker) and clock-injectable for deterministic tests.
  Reimplemented in TS from the shipped 9router
  src/shared/services/quotaAutoPing.js (Codex half only).
- Migration 123: last_ping_at / last_pinged_reset_key on
  provider_connections, so the scheduler never re-pings the same reset
  window twice.
- Settings: codexAutoPing.connections map, default {} (nobody opted in),
  validated by the shared Zod schema.
- Respects the existing resilience layers: skips a connection whose
  provider circuit breaker is open or whose rateLimitedUntil cooldown is
  active, and applies its own 15-minute failure cooldown after a failed
  ping.
- UI: new per-connection toggle (Settings -> AI -> Codex Quota Auto-Ping)
  with an explicit "consumes real quota" tooltip, wired to the settings
  PATCH route; i18n keys added to all 43 locales (EN authored, others
  filled from the EN fallback pending real translation).
- 15 new deterministic unit tests covering enable/disable, first-reset
  observation (cache-only, no ping), reset-slide ping, stable-reset
  no-op, min-ping-interval, same-resetKey dedupe, session/weekly quota
  exhaustion, non-OAuth skip, circuit-breaker-open skip, cooldown skip,
  failure-cooldown skip, failed-ping bookkeeping, the real executor call
  shape, and credential-refresh failure handling.

Antigravity's 2-bucket auto-ping is out of scope for this PR (no upstream
reference exists for its reset shape) and is tracked as a follow-up.

Ref #6977 (backend + settings + minimal UI toggle; Antigravity follow-up
tracked separately — not closing the issue from this PR)
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Warning

You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again!

@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@diegosouzapw diegosouzapw added the hold-vps PR verde, merge aguardando validação live (release-drain) label Jul 12, 2026
@diegosouzapw
diegosouzapw changed the base branch from release/v3.8.47 to release/v3.8.48 July 13, 2026 05:04
@diegosouzapw
diegosouzapw changed the base branch from release/v3.8.48 to release/v3.8.49 July 13, 2026 21:57
@diegosouzapw

Copy link
Copy Markdown
Owner Author

Thanks for building out Phase 1 of #6977 — this is a clean, well-scoped implementation: strictly opt-in (default off, per-connection), fully dependency-injected/deterministic scheduler, respects the existing circuit-breaker/cooldown resilience layers, and ships 15 solid unit tests. I ran them locally (all green) along with the related db-providers-crud and settings-schema suites (41/41 green, no regression). ESLint is clean on every touched file, and the call graph (getExecutor, getCodexUsage, refreshAndUpdateCredentials, getProviderConnections) all lines up with current signatures.

One blocker before merge: GitHub reports a conflict, and I scoped it with git merge-tree — it's exactly one file, src/i18n/messages/pt-BR.json (two key blocks landing at the same anchor line; trivial to keep both). Everything else that looks "changed in both" in a 3-way diff resolves cleanly on its own — the settings.ts divergence against the current release tip is just because this branch predates the #6678/#7149 fixes that landed since; a normal merge/rebase will preserve those, not revert them.

We'll resolve that i18n conflict and rebase onto the current release tip, then merge. Per the parked plan-file for #6977, the last step is the live_check on the VPS: enable auto-ping for one real Codex OAuth connection, wait for a session-quota reset, and confirm exactly one minimal ping fires (last_pinged_reset_key updates) without extra burn — will report back once that's done.

Small nit for later, not blocking: the module docstring calls the ping request "non-billed-model", which reads a little inconsistent with the UI's own "consumes a small amount of real Codex quota" warning — worth aligning the wording so it doesn't undersell what's actually happening.

…quota-auto-ping

# Conflicts:
#	src/i18n/messages/pt-BR.json
@diegosouzapw
diegosouzapw merged commit ad7692b into release/v3.8.49 Jul 17, 2026
5 of 8 checks passed
@diegosouzapw
diegosouzapw deleted the feat/6977-quota-auto-ping branch July 19, 2026 21:00
HouMinXi pushed a commit to HouMinXi/OmniRoute that referenced this pull request Aug 2, 2026
…souzapw#6977) (diegosouzapw#6995)

Codex's rolling "session" quota window only starts counting down once a
request lands inside it, so an idle connection's window keeps sliding
forward and the first real request after a long idle period pays for the
whole warm-up latency. This adds a strictly opt-in, per-connection
scheduler (default OFF) that watches an enabled connection's reported
resetAt and, once it slides forward, fires one tiny non-billed-model
request through the real Codex executor to keep the window warm.

- New in-process scheduler (src/lib/services/quotaAutoPing.ts), fully
  dependency-injected (settings, DB, credential refresh, usage fetch,
  executor, circuit breaker) and clock-injectable for deterministic tests.
  Reimplemented in TS from the shipped 9router
  src/shared/services/quotaAutoPing.js (Codex half only).
- Migration 123: last_ping_at / last_pinged_reset_key on
  provider_connections, so the scheduler never re-pings the same reset
  window twice.
- Settings: codexAutoPing.connections map, default {} (nobody opted in),
  validated by the shared Zod schema.
- Respects the existing resilience layers: skips a connection whose
  provider circuit breaker is open or whose rateLimitedUntil cooldown is
  active, and applies its own 15-minute failure cooldown after a failed
  ping.
- UI: new per-connection toggle (Settings -> AI -> Codex Quota Auto-Ping)
  with an explicit "consumes real quota" tooltip, wired to the settings
  PATCH route; i18n keys added to all 43 locales (EN authored, others
  filled from the EN fallback pending real translation).
- 15 new deterministic unit tests covering enable/disable, first-reset
  observation (cache-only, no ping), reset-slide ping, stable-reset
  no-op, min-ping-interval, same-resetKey dedupe, session/weekly quota
  exhaustion, non-OAuth skip, circuit-breaker-open skip, cooldown skip,
  failure-cooldown skip, failed-ping bookkeeping, the real executor call
  shape, and credential-refresh failure handling.

Antigravity's 2-bucket auto-ping is out of scope for this PR (no upstream
reference exists for its reset shape) and is tracked as a follow-up.

Ref diegosouzapw#6977 (backend + settings + minimal UI toggle; Antigravity follow-up
tracked separately — not closing the issue from this PR)
pacocartones added a commit to pacocartones/OmniRoute that referenced this pull request Sep 1, 2026
The opt-in Codex quota auto-ping (diegosouzapw#6977/diegosouzapw#6995) pinned its ping model to
`gpt-5.1-codex-mini` in src/shared/constants/quotaAutoPing.ts. OpenAI shut
that model down on 2026-07-23 and the repo's own lifecycle registry
(open-sse/services/modelLifecycle.ts + config/quality/model-lifecycle.json)
already rejects it on the request path, but the scheduler never consulted
that gate: every window slide sent the dead id through the real executor,
the failure landed in the 15-minute cooldown, and the same id was retried
forever. The feature could no longer warm a window, and any future
retirement of a pinned id would regress it the same way.

The ping model is now resolved per tick from the provider catalog and the
lifecycle registry: the first base-model entry (effort-suffixed variants are
skipped because the ping sets `reasoning.effort` itself) that
`isModelSelectable("codex", id)` allows, which is the same provider-scoped
gate chatCore applies. The registry import stays lazy, like the executor
import, because this module is on the instrumentation boot path (diegosouzapw#12074).
When nothing in the catalog is selectable, the provider is paused before
any throttle slot, usage read, or executor call, with one warning per state
change instead of a blind retry loop.

Closes diegosouzapw#11905

Co-authored-by: Leon Marcos <leonaniagomez@gmail.com>
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…souzapw#6977) (diegosouzapw#6995)

Codex's rolling "session" quota window only starts counting down once a
request lands inside it, so an idle connection's window keeps sliding
forward and the first real request after a long idle period pays for the
whole warm-up latency. This adds a strictly opt-in, per-connection
scheduler (default OFF) that watches an enabled connection's reported
resetAt and, once it slides forward, fires one tiny non-billed-model
request through the real Codex executor to keep the window warm.

- New in-process scheduler (src/lib/services/quotaAutoPing.ts), fully
  dependency-injected (settings, DB, credential refresh, usage fetch,
  executor, circuit breaker) and clock-injectable for deterministic tests.
  Reimplemented in TS from the shipped 9router
  src/shared/services/quotaAutoPing.js (Codex half only).
- Migration 123: last_ping_at / last_pinged_reset_key on
  provider_connections, so the scheduler never re-pings the same reset
  window twice.
- Settings: codexAutoPing.connections map, default {} (nobody opted in),
  validated by the shared Zod schema.
- Respects the existing resilience layers: skips a connection whose
  provider circuit breaker is open or whose rateLimitedUntil cooldown is
  active, and applies its own 15-minute failure cooldown after a failed
  ping.
- UI: new per-connection toggle (Settings -> AI -> Codex Quota Auto-Ping)
  with an explicit "consumes real quota" tooltip, wired to the settings
  PATCH route; i18n keys added to all 43 locales (EN authored, others
  filled from the EN fallback pending real translation).
- 15 new deterministic unit tests covering enable/disable, first-reset
  observation (cache-only, no ping), reset-slide ping, stable-reset
  no-op, min-ping-interval, same-resetKey dedupe, session/weekly quota
  exhaustion, non-OAuth skip, circuit-breaker-open skip, cooldown skip,
  failure-cooldown skip, failed-ping bookkeeping, the real executor call
  shape, and credential-refresh failure handling.

Antigravity's 2-bucket auto-ping is out of scope for this PR (no upstream
reference exists for its reset shape) and is tracked as a follow-up.

Ref diegosouzapw#6977 (backend + settings + minimal UI toggle; Antigravity follow-up
tracked separately — not closing the issue from this PR)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

hold-vps PR verde, merge aguardando validação live (release-drain)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant