Skip to content

feat(providers): add expiry-first account fallback strategy - #14533

Merged
diegosouzapw merged 3 commits into
diegosouzapw:release/v3.8.51from
feci:feat/expiry-first-account-rotation
Sep 24, 2026
Merged

diegosouzapw merged 3 commits into
diegosouzapw:release/v3.8.51from
feci:feat/expiry-first-account-rotation

Conversation

@feci

@feci feci commented Sep 22, 2026

Copy link
Copy Markdown
Contributor

Problem

A provider with several accounts can rotate them with fill-first, round-robin, priority,
p2c, random, least-used, cost-optimized or strict-random. None of these considers when
an account's quota window resets
, so fill-first — the default — drains the top-priority account
while the others sit full until their windows roll over and the unspent quota is lost.

Observed on a four-account Codex pool over 14 days:

priority calls input tokens
1 4605 630.7M
2 2147 264.8M
3 1779 229.8M
4 203 25.3M

At the time of measurement priority 1 held 26% with ~145h to reset, while priority 4 held 88% with
~69h to reset. The account about to lose the most quota was the one barely being used.

What this adds

expiry-first, a per-provider account fallback strategy that spends the quota closest to being
lost. Each account is ranked by how much it must burn per hour to avoid wasting its leftover:

score = usable / hoursUntilNearestReset
  • usable is the tightest window's remaining fraction. Nested windows (a session cap inside a
    weekly one) decrement together, so an account can never spend more than its most constrained
    window allows.
  • The deadline is the nearest reset, because that is when the first tranche is lost.
  • An exhausted account scores zero, so it never wins on a near reset alone. That is why plain
    earliest-deadline-first is the wrong rule: on the pool above it would pick the 1%-remaining
    account purely because it resets in 8h.

On that pool: 0.26/145 = 0.0018, 0.88/145 = 0.0061, 1% excluded, 0.88/69.3 = 0.0127 → the last
one is selected, which is the account whose quota was about to expire unused.

Why not extend reset-aware

scoreResetAwareQuota answers a different question. It ranks mostly on leftover and adds
resetUrgency * (1 - remaining) — a recovery signal that favours a nearly empty account about
to refresh. Its urgency term is also relative to the nominal window length and clamps to zero
outside it, so for the two 88% accounts above it returns the same score (0.341725 each) and
cannot separate a 69h reset from a 145h one. Both strategies are useful; they are not the same
strategy, so this adds one rather than changing the meaning of the existing one.

Scope and safety

  • Ordering only. The strategy ranks accounts the existing quota, rate-limit and peak-hour filters
    have already allowed; it adds no upstream quota call (it reads the per-connection quota cache).
  • Session stickiness and prompt-cache affinity are untouched.
  • Accounts within expiryFirstTieBandPercent (default 5%, compared relatively since the score is a
    rate) rotate least-recently-used, so equivalent accounts still share load instead of pinning.
  • Falls back to the incoming priority order when no account reports quota telemetry.
  • New tunables, all defaulted: expiryFirstTieBandPercent, expiryFirstMinHours,
    expiryFirstExhaustedFloorPercent.

Tests

tests/unit/expiry-first-account-rotation.test.ts — 10 cases covering the ranking, the
equally-full-different-reset case, exhausted-account exclusion, tightest-window binding, tie
rotation, backoff skip, the no-telemetry fallback and the empty pool. The ordering is a pure
exported function with an injected quota lookup, so it needs no database or network.

Also ran the 25 existing quota/strategy/fallback unit test files: 213 passing, 0 failing.

A provider with several accounts could only rotate them with fill-first,
round-robin, p2c, random, least-used or cost-optimized. None of these looks at
when an account's quota window resets, so fill-first drains the top-priority
account while the rest sit full until their windows roll over and the unspent
quota is simply lost.

expiry-first ranks each account by how much quota it must spend per hour to
avoid wasting its leftover at the next reset:

    score = usable / hoursUntilNearestReset

usable is the tightest window's remaining fraction, because nested windows (a
session cap inside a weekly one) decrement together and an account can never
spend more than its most constrained window allows. The deadline is the nearest
reset, since that is when the first tranche is lost. An exhausted account scores
zero, so it never wins on a near reset alone — which is why plain
earliest-deadline-first is the wrong rule here.

Measured on a four-account Codex pool:

    26% left, resets in 145h  ->  0.0018 /h
    88% left, resets in 145h  ->  0.0061 /h
     1% left, resets in   8h  ->  excluded, exhausted
    88% left, resets in  69h  ->  0.0127 /h   <- selected

fill-first picks the first of those and lets the last one's 88% expire.

Deliberately separate from reset-aware rather than reusing it: that scorer ranks
mostly on leftover and adds resetUrgency * (1 - remaining), a recovery signal
favouring a nearly empty account about to refresh. Its urgency term is relative
to the nominal window length and saturates to zero outside it, so it scores the
two 88% accounts above identically (0.341725) and cannot separate a 69h reset
from a 145h one.

Session stickiness and prompt-cache affinity are untouched; the strategy only
orders accounts that the existing quota and rate-limit filters already allowed.
Accounts within expiryFirstTieBandPercent of the leader rotate
least-recently-used, so equivalent accounts still share load instead of pinning.

The ordering is a pure exported function with an injected quota lookup, so it is
covered without a database or any upstream quota call.
@feci
feci requested a review from diegosouzapw as a code owner September 22, 2026 16:04
@diegosouzapw

Copy link
Copy Markdown
Owner

Nice addition — the reset-pressure framing (score = usable / hoursUntilNearestReset) is a
clean generalization of the existing reset-aware/reset-window pair, and the 10-case test
suite is solid (ran it locally, 10/10 pass). One gap before merge: docs/routing/AUTO-COMBO.md
still says "19 routing strategies" and its strategy table doesn't list expiry-first yet —
could you add it there? Also flagging: the changelog fragment is named
14381-expiry-first-account-rotation.md — is #14381 the right issue number, or should that be
#14533?

Tomas Fecko and others added 2 commits September 24, 2026 02:39
Renamed changelog.d/features/14381-expiry-first-account-rotation.md to
14533-expiry-first-account-rotation.md — 14381 is an unrelated open issue
(config-security CLI relay slice); the correct PR number is 14533.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
…af module

Move buildConnectionQuotaWindowsView and selectExpiryFirstConnection out of
the frozen src/sse/services/auth.ts into src/sse/services/expiryFirstAccountSelection.ts,
and route the expiry-first strategy through the least-used branch so it shares
the lastUsedAt commit. auth.ts now carries only the import and a two-line
dispatch; selection behavior is unchanged.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
@diegosouzapw
diegosouzapw merged commit ab052ab into diegosouzapw:release/v3.8.51 Sep 24, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants