fix(resilience): lock permanently retired models instead of short backoff (Gemini ban prevention) - #11762
Merged
diegosouzapw merged 2 commits intoAug 29, 2026
Conversation
…koff
Gemini's deprecated-model 404 ("This model models/gemini-2.5-flash is no
longer available to new users...") and similar provider end-of-life 410s
fell through checkFallbackError's generic transient-error branch, which
only applies a short cooldown (seconds, escalating to ~20min max). Combo
auto-routing kept re-selecting the dead model roughly every cooldown
window, all day, sending guaranteed-to-fail requests upstream. At volume
this looked like abusive traffic and was implicated in a Gemini free-tier
API key ban.
Add MODEL_PERMANENTLY_UNAVAILABLE_PATTERNS + isModelPermanentlyUnavailable()
and classify matching 404/410 responses with a 24h lockout instead, surfaced
via quotaResetHintMs so combo's per-request model-lockout honors it in full.
…yker tap.testFiles
This was referenced Aug 27, 2026
Merged
diegosouzapw
pushed a commit
that referenced
this pull request
Aug 29, 2026
…ng-suspended accounts (#11774) Follow-up to #11762, same bug class hitting freeaiapikey (410 permanently-moved endpoint) and fireworks (412 billing-suspension) — both fell through checkFallbackError's generic transient-cooldown branch and got retried every ~1 minute for a full day. Fix: `ENDPOINT_PERMANENTLY_MOVED_PATTERNS`/`isEndpointPermanentlyMoved()` → 24h lockout; `ACCOUNT_SUSPENDED_BILLING_PATTERNS`/`isAccountSuspendedForBilling()` → treated as credits-exhausted (1h cooldown), independent of status code so it also catches Fireworks' 412. #11762 landed first and touched the same file — rebased/re-merged onto the updated tip (additive, no logic changes) and re-validated: 13/13 tests pass. Thanks for tracing this with real production logs again!
diegosouzapw
pushed a commit
that referenced
this pull request
Aug 29, 2026
… requests (#11781) Follow-up to #11762/#11774, same bug class in combo's own model-lockout wiring: GitHub rejects several models (gpt-5.4, gpt-5.3-codex, etc.) with a 400 that's permanently unavailable for this account's Copilot integration, but nothing recorded a cross-request lockout — combo's #5249 in-request advance guard is correct but doesn't persist, so the same doomed model gets retried from scratch on every new request, indefinitely. Fix: on a model-scoped 400 (`isModelScoped400`), call `lockModelIfPerModelQuota(provider, connectionId, rawModel, "model_capacity", 1h)`. GitHub already has per-model-quota enabled, so only the rejected model locks — siblings keep working. `isModelLocked()` is already checked pre-dispatch, so no other wiring needed. Validated: 3/3 new tests + fixed a pre-existing test-isolation gap in combo-model-scoped-400-advance.test.ts (shared model name across sub-tests without clearing lockout state). Thanks!
diegosouzapw
pushed a commit
that referenced
this pull request
Aug 29, 2026
…EADME (#11772) Finishes the Freepik → Magnific rebrand from #10594 across 40 locale files and 3 README feature-list bullets (README.md, docs/i18n/it, docs/i18n/tr) — legacy `freepik` alias intentionally left in code/tests/redirects for backward compatibility, and historical CHANGELOG entries left untouched as documented history. The README bullet had base-drifted since the PR branched (release tip's "What's New" changelog snippet had already dropped two providers mentioned nowhere else in the codebase, unrelated to this PR's scope) — resolved by keeping the tip's current bullet shape and applying only the Freepik→Magnific rename on top, in both the combined-worktree validation and the pushed branch. Validated: all 40 edited locale JSON files parse; re-verified after resync onto the updated tip (post #11762/#11774/#11781).
muhamadgalihsaputra
pushed a commit
to niyatna/NiyatnaRoute
that referenced
this pull request
Sep 27, 2026
…koff (Gemini ban prevention) (diegosouzapw#11762) Root-caused via a real Gemini-ban incident log: deprecated-model 404/410s (e.g. gemini-2.5-flash "no longer available to new users") fell through checkFallbackError's generic transient-cooldown branch, so combo/auto-routing kept re-selecting a permanently dead model every cooldown window forever — the hammering that got the account flagged as abusive. Fix: `MODEL_PERMANENTLY_UNAVAILABLE_PATTERNS` + `isModelPermanentlyUnavailable()` classify these as a 24h lockout instead, surfaced via `quotaResetHintMs` so combo's per-request model-lockout honors it in full. Validated: 6/6 new tests + 133/133 existing accountFallback/error-classification tests, no regressions. Thanks for tracing this end-to-end with real production logs!
muhamadgalihsaputra
pushed a commit
to niyatna/NiyatnaRoute
that referenced
this pull request
Sep 27, 2026
…ng-suspended accounts (diegosouzapw#11774) Follow-up to diegosouzapw#11762, same bug class hitting freeaiapikey (410 permanently-moved endpoint) and fireworks (412 billing-suspension) — both fell through checkFallbackError's generic transient-cooldown branch and got retried every ~1 minute for a full day. Fix: `ENDPOINT_PERMANENTLY_MOVED_PATTERNS`/`isEndpointPermanentlyMoved()` → 24h lockout; `ACCOUNT_SUSPENDED_BILLING_PATTERNS`/`isAccountSuspendedForBilling()` → treated as credits-exhausted (1h cooldown), independent of status code so it also catches Fireworks' 412. diegosouzapw#11762 landed first and touched the same file — rebased/re-merged onto the updated tip (additive, no logic changes) and re-validated: 13/13 tests pass. Thanks for tracing this with real production logs again!
muhamadgalihsaputra
pushed a commit
to niyatna/NiyatnaRoute
that referenced
this pull request
Sep 27, 2026
… requests (diegosouzapw#11781) Follow-up to diegosouzapw#11762/diegosouzapw#11774, same bug class in combo's own model-lockout wiring: GitHub rejects several models (gpt-5.4, gpt-5.3-codex, etc.) with a 400 that's permanently unavailable for this account's Copilot integration, but nothing recorded a cross-request lockout — combo's diegosouzapw#5249 in-request advance guard is correct but doesn't persist, so the same doomed model gets retried from scratch on every new request, indefinitely. Fix: on a model-scoped 400 (`isModelScoped400`), call `lockModelIfPerModelQuota(provider, connectionId, rawModel, "model_capacity", 1h)`. GitHub already has per-model-quota enabled, so only the rejected model locks — siblings keep working. `isModelLocked()` is already checked pre-dispatch, so no other wiring needed. Validated: 3/3 new tests + fixed a pre-existing test-isolation gap in combo-model-scoped-400-advance.test.ts (shared model name across sub-tests without clearing lockout state). Thanks!
muhamadgalihsaputra
pushed a commit
to niyatna/NiyatnaRoute
that referenced
this pull request
Sep 27, 2026
…EADME (diegosouzapw#11772) Finishes the Freepik → Magnific rebrand from diegosouzapw#10594 across 40 locale files and 3 README feature-list bullets (README.md, docs/i18n/it, docs/i18n/tr) — legacy `freepik` alias intentionally left in code/tests/redirects for backward compatibility, and historical CHANGELOG entries left untouched as documented history. The README bullet had base-drifted since the PR branched (release tip's "What's New" changelog snippet had already dropped two providers mentioned nowhere else in the codebase, unrelated to this PR's scope) — resolved by keeping the tip's current bullet shape and applying only the Freepik→Magnific rename on top, in both the combined-worktree validation and the pushed branch. Validated: all 40 edited locale JSON files parse; re-verified after resync onto the updated tip (post diegosouzapw#11762/diegosouzapw#11774/diegosouzapw#11781).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why this happened
We had a Gemini API key get banned after a burst of "spam-looking" traffic. Filtering our 24h request logs for
provider: "gemini"showed three repeating patterns:1. Genuine free-tier 429 quota errors (expected/handled correctly, not the bug):
2. The actual bug — deprecated-model 404s, recurring every ~15–45 minutes, all day:
3. Symptom of the above — the local rate-limit queue repeatedly getting congested:
Root cause
I traced
checkFallbackError()inopen-sse/services/accountFallback.tsend-to-end. Contrary to my first hypothesis, Gemini's"Please retry in Ns"429 hint is already parsed and honored correctly (parseRetryFromErrorText+useUpstreamRetryHints, which defaults totruefor apikey-category providers) — that part was not the bug.The real gap: a 404 for a permanently retired model (Google explicitly says "no longer available to new users") had no dedicated classification. It fell through to the generic "all other errors" branch at the bottom of
checkFallbackError, which only applies a short transient cooldown (a few seconds, escalating to a ~20 minute cap viaBACKOFF_STEPS_MS). Because Gemini is a per-model-quota provider (hasPerModelQuota("gemini") === true), combo/auto-routing kept re-selecting the same dead model roughly every cooldown window — forever, since the model will never succeed. That produced a steady drumbeat of guaranteed-to-fail requests against the account, on top of the genuine 429 bursts on other Gemini model aliases. At volume, this combination looks like abusive/spam traffic to Google and is what got the account flagged and banned.The fix
In
open-sse/services/accountFallback.ts:MODEL_PERMANENTLY_UNAVAILABLE_PATTERNS— matches phrasing like "no longer available", "no longer supported", "has reached its end of life", "model ... deprecated/retired/discontinued" (also fixes the same class of bug for Fireworks/etc. end-of-life 410s, e.g.minimaxai/minimax-m2.7seen in the same logs).isModelPermanentlyUnavailable()to test error text against those patterns.checkFallbackError()now classifies a 404/410 matching those patterns with a 24-hour lockout (reason: "not_found"), instead of the short generic transient cooldown. The lockout is surfaced viaquotaResetHintMs, which flows into combo's per-request model-lockout (open-sse/services/combo.ts) as an upstream-verified reset — so it is honored in full and is not clamped down to the normal ~20 minute model-lockout ceiling.How this prevents future bans
Testing
Added
tests/unit/gemini-deprecated-model-lockout.test.ts(6 tests, all passing):checkFallbackErrorreturns a 24h lockout (reason: "not_found",quotaResetHintMs) for both the Gemini 404 and a Fireworks-style end-of-life 410.Ran locally:
Files changed
open-sse/services/accountFallback.ts— new pattern set + classification branch.tests/unit/gemini-deprecated-model-lockout.test.ts— new regression test.