Skip to content

feat(resilience): throttle concurrent upstream quota fetches — closes #6009 - #6058

Merged
diegosouzapw merged 6 commits into
release/v3.8.44from
feat/quota-fetch-throttle-6009
Jul 4, 2026
Merged

diegosouzapw merged 6 commits into
release/v3.8.44from
feat/quota-fetch-throttle-6009

Conversation

@diegosouzapw

Copy link
Copy Markdown
Owner

What

Adds a global min-interval throttle for upstream quota fetches on the per-request preflight/monitor path, so many accounts on one IP no longer fetch provider quota in the same second.

Why

Per the reporter and router-for-me/CLIProxyAPI#2385, a burst of simultaneous quota requests from one IP looks like automation and can get a Codex OAuth token revoked.

OmniRoute already spaces the periodic bulk provider-limits sync (PROVIDER_LIMITS_SYNC_SPACING_MS, sequential 1500ms) — but the concurrent combo/preflight path (combo.ts Promise.all(...fetchCodexQuota)) was not throttled. This PR fills exactly that gap.

How

  • New reusable gate open-sse/services/quotaFetchThrottle.ts — MinIntervalThrottle serializes callers so each network call starts ≥ minIntervalMs (+ jitter) after the previous. Injectable clock/RNG → deterministic tests.
  • codexQuotaFetcher.ts awaits throttleQuotaFetch() immediately before the fetch(/wham/usage) call — after the cache check, so cache hits are never delayed.
  • Fail-open: acquire() only ever awaits a timer; it cannot throw the fetcher off its fail-open path. minIntervalMs=0 disables it (byte-identical to before).
  • Configurable: OMNIROUTE_QUOTA_FETCH_MIN_INTERVAL_MS (default 250ms, clamped 0..5000). Documented in .env.example + ENVIRONMENT.md.

The gate is process-wide, so any other provider quota fetcher can adopt it with a one-line await throttleQuotaFetch() later (kept surgical here — Codex is the provider in the linked revocation report).

Note on the UI-lazy part of the issue

The issue also asks that UI quota fetching be lazy. The bulk sync is already interval-driven + spaced; and because this gate is global, even a UI-triggered refresh-all now gets spaced rather than bursting. A dedicated "no eager refresh-all in the dashboard" change (if still needed) is left as a separate follow-up to avoid speculative UI regressions.

Tests / validation (Hard Rule #18)

  • TDD: tests/unit/quota-fetch-throttle-6009.test.ts (5) — spacing of N concurrent calls (fake clock), 0=disabled, single call never delayed, jitter bound, env clamp/default. Green.
  • typecheck:core · check:cycles · check:docs-sync · check:env-doc-sync · eslint(touched) all clean.
  • No regression: existing codex-quota-fetcher, chatcore-codex-quota, codex-banked-reset-credits-5199, quota-preflight, resilience-settings-quota-preflight, eviction-guards-codexQuotaFetcher suites all green.

Many accounts on one IP fetching provider quota in the same second looks like
automation and — per router-for-me/CLIProxyAPI#2385 — can get a Codex OAuth token
revoked. Add a global min-interval gate (open-sse/services/quotaFetchThrottle.ts)
that serializes the actual network calls made by the Codex quota fetcher so their
starts are spaced >= OMNIROUTE_QUOTA_FETCH_MIN_INTERVAL_MS (default 250ms, 0=off).

Complements the existing bulk-sync spacing (PROVIDER_LIMITS_SYNC_SPACING_MS) which
already serialized the periodic provider-limits sync; this covers the concurrent
combo/preflight path (combo.ts Promise.all -> fetchCodexQuota) it did not. Cache
hits never reach the gate; fail-open (only ever awaits a timer). Reusable so other
provider fetchers can adopt throttleQuotaFetch() in one line.

Closes #6009.
Regression guard: tests/unit/quota-fetch-throttle-6009.test.ts (5).
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Warning

You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again!

@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@diegosouzapw
diegosouzapw merged commit dc2efa5 into release/v3.8.44 Jul 4, 2026
1 check was pending
@diegosouzapw
diegosouzapw deleted the feat/quota-fetch-throttle-6009 branch July 4, 2026 07:38
abhisheksharma2411 added a commit to abhisheksharma2411/OmniRoute that referenced this pull request Aug 31, 2026
…ta throttle

Every Codex quota read goes through throttleQuotaFetch() — the diegosouzapw#6009/diegosouzapw#6058
gate that spaces genuine upstream calls so many accounts behind one IP do not
fire in the same second, which is the pattern documented to have got a Codex
OAuth token revoked. The auto-ping scheduler called getCodexUsage() directly,
so the one Codex path that runs unattended every 60s per connection was the
one skipping the mitigation written for Codex.

The tick walks connections sequentially but without spacing, so N enabled
connections still produce N upstream usage requests within a few hundred ms.

Gate the read on the same throttle, injected through deps like every other
effect in this module. Placed after the skip checks so a connection filtered
out by the circuit breaker, a cooldown or the failure cache does not consume
a slot and delay the connections that do reach the network.

This does not change the polling cadence. Codex sets pingWhenResetAtSlides
because its resetAt slides forward while the window is idle, so the per-tick
re-fetch is deliberate and is left alone.

Closes diegosouzapw#11904
diegosouzapw pushed a commit that referenced this pull request Sep 1, 2026
…ta throttle (#12209)

Every Codex quota read goes through throttleQuotaFetch() — the #6009/#6058
gate that spaces genuine upstream calls so many accounts behind one IP do not
fire in the same second, which is the pattern documented to have got a Codex
OAuth token revoked. The auto-ping scheduler called getCodexUsage() directly,
so the one Codex path that runs unattended every 60s per connection was the
one skipping the mitigation written for Codex.

The tick walks connections sequentially but without spacing, so N enabled
connections still produce N upstream usage requests within a few hundred ms.

Gate the read on the same throttle, injected through deps like every other
effect in this module. Placed after the skip checks so a connection filtered
out by the circuit breaker, a cooldown or the failure cache does not consume
a slot and delay the connections that do reach the network.

This does not change the polling cadence. Codex sets pingWhenResetAtSlides
because its resetAt slides forward while the window is idle, so the per-tick
re-fetch is deliberate and is left alone.

Closes #11904
tkgo11 pushed a commit to tkgo11/OmniRoute that referenced this pull request Sep 23, 2026
…iegosouzapw#6009 (diegosouzapw#6058)

Throttle concurrent upstream quota fetches (diegosouzapw#6009). Integrated into release/v3.8.44.
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…ta throttle (diegosouzapw#12209)

Every Codex quota read goes through throttleQuotaFetch() — the diegosouzapw#6009/diegosouzapw#6058
gate that spaces genuine upstream calls so many accounts behind one IP do not
fire in the same second, which is the pattern documented to have got a Codex
OAuth token revoked. The auto-ping scheduler called getCodexUsage() directly,
so the one Codex path that runs unattended every 60s per connection was the
one skipping the mitigation written for Codex.

The tick walks connections sequentially but without spacing, so N enabled
connections still produce N upstream usage requests within a few hundred ms.

Gate the read on the same throttle, injected through deps like every other
effect in this module. Placed after the skip checks so a connection filtered
out by the circuit breaker, a cooldown or the failure cache does not consume
a slot and delay the connections that do reach the network.

This does not change the polling cadence. Codex sets pingWhenResetAtSlides
because its resetAt slides forward while the window is idle, so the per-tick
re-fetch is deliberate and is left alone.

Closes diegosouzapw#11904
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant