Skip to content

fix(resilience): lock permanently retired models instead of short backoff (Gemini ban prevention) - #11762

Merged
diegosouzapw merged 2 commits into
diegosouzapw:release/v3.8.51from
turbolego:fix/gemini-deprecated-model-lockout
Aug 29, 2026
Merged

diegosouzapw merged 2 commits into
diegosouzapw:release/v3.8.51from
turbolego:fix/gemini-deprecated-model-lockout

Conversation

@turbolego

Copy link
Copy Markdown
Contributor

Why this happened

We had a Gemini API key get banned after a burst of "spam-looking" traffic. Filtering our 24h request logs for provider: "gemini" showed three repeating patterns:

1. Genuine free-tier 429 quota errors (expected/handled correctly, not the bug):

[429]: You exceeded your current quota, please check your plan and billing details. ...
* Quota exceeded for metric: generativelanguage.googleapis.com/generate_content_free_tier_requests,
  limit: 20, model: gemini-3-flash
Please retry in 20.551466945s.

2. The actual bug — deprecated-model 404s, recurring every ~15–45 minutes, all day:

"model": "gemini-2.5-flash",
"status": 404,
"error": "[404]: This model models/gemini-2.5-flash is no longer available to new users.
          Please update your code to use models/gemini-3.6-flash for the latest features
          and improvements."
"model": "gemini-2.5-flash-lite",
"status": 404,
"error": "[404]: This model models/gemini-2.5-flash-lite is no longer available to new users.
          Please update your code to use models/gemini-3.5-flash-lite for the latest features
          and improvements."

3. Symptom of the above — the local rate-limit queue repeatedly getting congested:

"status": 502,
"error": "[502]: Request dropped: the local rate-limit queue for gemini/gemini-3.1-flash-lite
          was detected as wedged (stalled with nothing executing) and force-reset. This is
          OmniRoute's own queue recovering, not an upstream error."

Root cause

I traced checkFallbackError() in open-sse/services/accountFallback.ts end-to-end. Contrary to my first hypothesis, Gemini's "Please retry in Ns" 429 hint is already parsed and honored correctly (parseRetryFromErrorText + useUpstreamRetryHints, which defaults to true for apikey-category providers) — that part was not the bug.

The real gap: a 404 for a permanently retired model (Google explicitly says "no longer available to new users") had no dedicated classification. It fell through to the generic "all other errors" branch at the bottom of checkFallbackError, which only applies a short transient cooldown (a few seconds, escalating to a ~20 minute cap via BACKOFF_STEPS_MS). Because Gemini is a per-model-quota provider (hasPerModelQuota("gemini") === true), combo/auto-routing kept re-selecting the same dead model roughly every cooldown window — forever, since the model will never succeed. That produced a steady drumbeat of guaranteed-to-fail requests against the account, on top of the genuine 429 bursts on other Gemini model aliases. At volume, this combination looks like abusive/spam traffic to Google and is what got the account flagged and banned.

The fix

In open-sse/services/accountFallback.ts:

  • Added MODEL_PERMANENTLY_UNAVAILABLE_PATTERNS — matches phrasing like "no longer available", "no longer supported", "has reached its end of life", "model ... deprecated/retired/discontinued" (also fixes the same class of bug for Fireworks/etc. end-of-life 410s, e.g. minimaxai/minimax-m2.7 seen in the same logs).
  • Added isModelPermanentlyUnavailable() to test error text against those patterns.
  • checkFallbackError() now classifies a 404/410 matching those patterns with a 24-hour lockout (reason: "not_found"), instead of the short generic transient cooldown. The lockout is surfaced via quotaResetHintMs, which flows into combo's per-request model-lockout (open-sse/services/combo.ts) as an upstream-verified reset — so it is honored in full and is not clamped down to the normal ~20 minute model-lockout ceiling.
  • A generic/unrelated 404 (e.g. "Model not found, inaccessible, and/or not deployed") is untouched and still falls through to the existing short-backoff path — this change is narrowly scoped to the specific "permanently retired" phrasing.

How this prevents future bans

  • Once Google tells us a model is retired, OmniRoute now stops hammering it for 24h instead of retrying every few minutes forever — cutting the volume of guaranteed-to-fail requests against a Gemini account from "dozens per day, indefinitely" down to effectively zero after the first failure.
  • Combined with the free-tier 429 handling (already correct), this removes the main source of sustained abusive-looking traffic against a single Gemini key, which is the behavior Google's anti-abuse systems flag for bans.

Testing

Added tests/unit/gemini-deprecated-model-lockout.test.ts (6 tests, all passing):

  • Matches Gemini's deprecated-model phrasing and generic end-of-life 410 phrasing.
  • Does not match an unrelated/generic 404.
  • checkFallbackError returns a 24h lockout (reason: "not_found", quotaResetHintMs) for both the Gemini 404 and a Fireworks-style end-of-life 410.
  • A generic 404 still gets the short transient cooldown (no regression).

Ran locally:

node --import tsx/esm --test tests/unit/gemini-deprecated-model-lockout.test.ts
# 6/6 pass

node --import tsx/esm --test tests/integration/gemini-rate-limit-classification.test.ts \
  tests/unit/account-fallback-service.test.ts tests/unit/error-classification.test.ts
# 133/133 pass, no regressions

Files changed

  • open-sse/services/accountFallback.ts — new pattern set + classification branch.
  • tests/unit/gemini-deprecated-model-lockout.test.ts — new regression test.

…koff

Gemini's deprecated-model 404 ("This model models/gemini-2.5-flash is no
longer available to new users...") and similar provider end-of-life 410s
fell through checkFallbackError's generic transient-error branch, which
only applies a short cooldown (seconds, escalating to ~20min max). Combo
auto-routing kept re-selecting the dead model roughly every cooldown
window, all day, sending guaranteed-to-fail requests upstream. At volume
this looked like abusive traffic and was implicated in a Gemini free-tier
API key ban.

Add MODEL_PERMANENTLY_UNAVAILABLE_PATTERNS + isModelPermanentlyUnavailable()
and classify matching 404/410 responses with a 24h lockout instead, surfaced
via quotaResetHintMs so combo's per-request model-lockout honors it in full.
@diegosouzapw
diegosouzapw merged commit 87b3bdf into diegosouzapw:release/v3.8.51 Aug 29, 2026
16 checks passed
diegosouzapw pushed a commit that referenced this pull request Aug 29, 2026
…ng-suspended accounts (#11774)

Follow-up to #11762, same bug class hitting freeaiapikey (410 permanently-moved endpoint) and fireworks (412 billing-suspension) — both fell through checkFallbackError's generic transient-cooldown branch and got retried every ~1 minute for a full day.

Fix: `ENDPOINT_PERMANENTLY_MOVED_PATTERNS`/`isEndpointPermanentlyMoved()` → 24h lockout; `ACCOUNT_SUSPENDED_BILLING_PATTERNS`/`isAccountSuspendedForBilling()` → treated as credits-exhausted (1h cooldown), independent of status code so it also catches Fireworks' 412.

#11762 landed first and touched the same file — rebased/re-merged onto the updated tip (additive, no logic changes) and re-validated: 13/13 tests pass. Thanks for tracing this with real production logs again!
diegosouzapw pushed a commit that referenced this pull request Aug 29, 2026
… requests (#11781)

Follow-up to #11762/#11774, same bug class in combo's own model-lockout wiring: GitHub rejects several models (gpt-5.4, gpt-5.3-codex, etc.) with a 400 that's permanently unavailable for this account's Copilot integration, but nothing recorded a cross-request lockout — combo's #5249 in-request advance guard is correct but doesn't persist, so the same doomed model gets retried from scratch on every new request, indefinitely.

Fix: on a model-scoped 400 (`isModelScoped400`), call `lockModelIfPerModelQuota(provider, connectionId, rawModel, "model_capacity", 1h)`. GitHub already has per-model-quota enabled, so only the rejected model locks — siblings keep working. `isModelLocked()` is already checked pre-dispatch, so no other wiring needed.

Validated: 3/3 new tests + fixed a pre-existing test-isolation gap in combo-model-scoped-400-advance.test.ts (shared model name across sub-tests without clearing lockout state). Thanks!
diegosouzapw pushed a commit that referenced this pull request Aug 29, 2026
…EADME (#11772)

Finishes the Freepik → Magnific rebrand from #10594 across 40 locale files and 3 README feature-list bullets (README.md, docs/i18n/it, docs/i18n/tr) — legacy `freepik` alias intentionally left in code/tests/redirects for backward compatibility, and historical CHANGELOG entries left untouched as documented history.

The README bullet had base-drifted since the PR branched (release tip's "What's New" changelog snippet had already dropped two providers mentioned nowhere else in the codebase, unrelated to this PR's scope) — resolved by keeping the tip's current bullet shape and applying only the Freepik→Magnific rename on top, in both the combined-worktree validation and the pushed branch.

Validated: all 40 edited locale JSON files parse; re-verified after resync onto the updated tip (post #11762/#11774/#11781).
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…koff (Gemini ban prevention) (diegosouzapw#11762)

Root-caused via a real Gemini-ban incident log: deprecated-model 404/410s (e.g. gemini-2.5-flash "no longer available to new users") fell through checkFallbackError's generic transient-cooldown branch, so combo/auto-routing kept re-selecting a permanently dead model every cooldown window forever — the hammering that got the account flagged as abusive.

Fix: `MODEL_PERMANENTLY_UNAVAILABLE_PATTERNS` + `isModelPermanentlyUnavailable()` classify these as a 24h lockout instead, surfaced via `quotaResetHintMs` so combo's per-request model-lockout honors it in full.

Validated: 6/6 new tests + 133/133 existing accountFallback/error-classification tests, no regressions. Thanks for tracing this end-to-end with real production logs!
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…ng-suspended accounts (diegosouzapw#11774)

Follow-up to diegosouzapw#11762, same bug class hitting freeaiapikey (410 permanently-moved endpoint) and fireworks (412 billing-suspension) — both fell through checkFallbackError's generic transient-cooldown branch and got retried every ~1 minute for a full day.

Fix: `ENDPOINT_PERMANENTLY_MOVED_PATTERNS`/`isEndpointPermanentlyMoved()` → 24h lockout; `ACCOUNT_SUSPENDED_BILLING_PATTERNS`/`isAccountSuspendedForBilling()` → treated as credits-exhausted (1h cooldown), independent of status code so it also catches Fireworks' 412.

diegosouzapw#11762 landed first and touched the same file — rebased/re-merged onto the updated tip (additive, no logic changes) and re-validated: 13/13 tests pass. Thanks for tracing this with real production logs again!
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
… requests (diegosouzapw#11781)

Follow-up to diegosouzapw#11762/diegosouzapw#11774, same bug class in combo's own model-lockout wiring: GitHub rejects several models (gpt-5.4, gpt-5.3-codex, etc.) with a 400 that's permanently unavailable for this account's Copilot integration, but nothing recorded a cross-request lockout — combo's diegosouzapw#5249 in-request advance guard is correct but doesn't persist, so the same doomed model gets retried from scratch on every new request, indefinitely.

Fix: on a model-scoped 400 (`isModelScoped400`), call `lockModelIfPerModelQuota(provider, connectionId, rawModel, "model_capacity", 1h)`. GitHub already has per-model-quota enabled, so only the rejected model locks — siblings keep working. `isModelLocked()` is already checked pre-dispatch, so no other wiring needed.

Validated: 3/3 new tests + fixed a pre-existing test-isolation gap in combo-model-scoped-400-advance.test.ts (shared model name across sub-tests without clearing lockout state). Thanks!
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…EADME (diegosouzapw#11772)

Finishes the Freepik → Magnific rebrand from diegosouzapw#10594 across 40 locale files and 3 README feature-list bullets (README.md, docs/i18n/it, docs/i18n/tr) — legacy `freepik` alias intentionally left in code/tests/redirects for backward compatibility, and historical CHANGELOG entries left untouched as documented history.

The README bullet had base-drifted since the PR branched (release tip's "What's New" changelog snippet had already dropped two providers mentioned nowhere else in the codebase, unrelated to this PR's scope) — resolved by keeping the tip's current bullet shape and applying only the Freepik→Magnific rename on top, in both the combined-worktree validation and the pushed branch.

Validated: all 40 edited locale JSON files parse; re-verified after resync onto the updated tip (post diegosouzapw#11762/diegosouzapw#11774/diegosouzapw#11781).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants