fix(sse): bound Antigravity 429 retry loop and lock quota-exhausted accounts for full reset window - #3122
Conversation
A persistent 429 on the short-retry branch (retryAfterMs ≤ 60s) looped forever on the same endpoint because the branch did `urlIndex--; continue` without checking the shared retry counter. Production log showed 77 consecutive 429s on one daily endpoint/account with zero fallback. Gate the short-retry branch on `retryAttemptsByUrl[urlIndex] < MAX_AUTO_RETRIES` (mirroring the already-bounded sibling), so a persistent 429 retries at most 3× per endpoint across all 3 base URLs then returns the 429 to the account- fallback layer. Regression test: 'bounds a persistent short-retry 429' in tests/unit/executor-antigravity.test.ts — asserts 12 total attempts (3 endpoints × 4) and a returned 429 with zero hang.
After the retry-loop bound, OmniRoute fell over to the next account but re-selected the exhausted one first on every subsequent request (~60s wasted per request). Root cause: the 429 body 'Individual quota reached. Contact your administrator to enable overages. Resets in 164h27m24s.' was not recognized as quota exhaustion, so the model was locked for only ~5s instead of the real 6.8-day reset window. Two detector fixes (mirrors Antigravity-Manager rate_limit.rs set_lockout_until): 1. classify429.ts — add QUOTA_PATTERNS: /individual quota reached/i, /quota reached/i, /enable overages/i so looksLikeQuotaExhausted() fires. 2. accountFallback.ts — same patterns in classifyErrorText(); extend parseRetryFromErrorText() to parse 'Resets? in XhYmZs' (reusing the existing computeDurationMs helper) so the exact reset duration reaches recordModelLockoutFailure as exactCooldownMs (uncapped, per user choice). The lockout machinery already stores until = now + cooldownMs with no clamp, bypasses getScaledCooldown when exactCooldownMs > 0, and keeps the longer of existing/new, so the full 164h window flows through intact. New patterns stay specific — plain 'too many requests'/'rate limit exceeded' messages still classify as rate_limit. Verified end-to-end against the real message: - classify429 → quota_exhausted - parseRetryFromErrorText → 592044000 ms (164h27m24s exactly) - checkFallbackError → usedUpstreamRetryHint: true, cooldownMs: 592044000
There was a problem hiding this comment.
Code Review
This pull request introduces changes to handle Antigravity / Cloud Code quota exhaustion and prevent infinite retry loops. Specifically, it bounds short-retries in AntigravityExecutor using a per-URL attempt counter to avoid infinite loops on persistent 429 errors. It also updates error parsing and classification to recognize Antigravity's 'Resets in XhYmZs' phrasing and quota-related error messages, ensuring they are correctly classified as quota exhaustion with the appropriate long cooldown rather than transient rate limits. Comprehensive unit tests have been added to verify these behaviors. There are no review comments, so I have no additional feedback to provide.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
- Remove /quota reached/i pattern (subsumed by /individual quota reached/i) - Delegate classifyErrorText to looksLikeQuotaExhausted for quota patterns - Add 30-day cap to computeDurationMs (prevent adversarial lockouts) - Flatten parseRetryFromErrorText nested structure - Add test for CREDITS_EXHAUSTED_SIGNALS 'free tier exhausted' entry - Add test for 30-day cooldown cap Addressed MAJOR findings from code review: - Pattern duplication across classify429.ts and accountFallback.ts - Missing test coverage for CREDITS_EXHAUSTED_SIGNALS entry - Overly broad /quota reached/i pattern (false-positive risk) - Unbounded cooldown injection (30-day cap added) - Nested if chain flattened for readability
6c63884 to
2d95845
Compare
…afety - Implement a 30-day cap on parsed retry durations to prevent indefinite account lockouts. - Replace manual string matching with `looksLikeQuotaExhausted` for more robust quota detection. - Streamline regex logic in `parseRetryFromErrorText` for better readability. - Add unit tests for extreme cooldown values and free-tier exhaustion scenarios.
…vity-429-retry-loop
The bare /quota reached/ would also flag transient per-minute limits like 'request quota reached, retry in 60s' as quota_exhausted (multi-hour lock). The Antigravity message is still caught by /individual quota reached/. Added a regression assertion proving the transient case stays a rate_limit.
|
Thanks, @ahmet-cetinkaya! 🙏 Two solid fixes: bounding the Antigravity 429 short-retry branch with the per-URL |
Code Review SummaryStatus: No Issues Found | Recommendation: Merge Files Reviewed (4 files)
Reviewed by step-3.7-flash-20260528 · 708,913 tokens |
…ccounts for full reset window (diegosouzapw#3122) * fix(sse): bound Antigravity short-retry 429 loop per endpoint A persistent 429 on the short-retry branch (retryAfterMs ≤ 60s) looped forever on the same endpoint because the branch did `urlIndex--; continue` without checking the shared retry counter. Production log showed 77 consecutive 429s on one daily endpoint/account with zero fallback. Gate the short-retry branch on `retryAttemptsByUrl[urlIndex] < MAX_AUTO_RETRIES` (mirroring the already-bounded sibling), so a persistent 429 retries at most 3× per endpoint across all 3 base URLs then returns the 429 to the account- fallback layer. Regression test: 'bounds a persistent short-retry 429' in tests/unit/executor-antigravity.test.ts — asserts 12 total attempts (3 endpoints × 4) and a returned 429 with zero hang. * fix(sse): lock Antigravity quota-exhausted account for full reset window After the retry-loop bound, OmniRoute fell over to the next account but re-selected the exhausted one first on every subsequent request (~60s wasted per request). Root cause: the 429 body 'Individual quota reached. Contact your administrator to enable overages. Resets in 164h27m24s.' was not recognized as quota exhaustion, so the model was locked for only ~5s instead of the real 6.8-day reset window. Two detector fixes (mirrors Antigravity-Manager rate_limit.rs set_lockout_until): 1. classify429.ts — add QUOTA_PATTERNS: /individual quota reached/i, /quota reached/i, /enable overages/i so looksLikeQuotaExhausted() fires. 2. accountFallback.ts — same patterns in classifyErrorText(); extend parseRetryFromErrorText() to parse 'Resets? in XhYmZs' (reusing the existing computeDurationMs helper) so the exact reset duration reaches recordModelLockoutFailure as exactCooldownMs (uncapped, per user choice). The lockout machinery already stores until = now + cooldownMs with no clamp, bypasses getScaledCooldown when exactCooldownMs > 0, and keeps the longer of existing/new, so the full 164h window flows through intact. New patterns stay specific — plain 'too many requests'/'rate limit exceeded' messages still classify as rate_limit. Verified end-to-end against the real message: - classify429 → quota_exhausted - parseRetryFromErrorText → 592044000 ms (164h27m24s exactly) - checkFallbackError → usedUpstreamRetryHint: true, cooldownMs: 592044000 * refactor(account-fallback): simplify error parsing and add cooldown safety - Implement a 30-day cap on parsed retry durations to prevent indefinite account lockouts. - Replace manual string matching with `looksLikeQuotaExhausted` for more robust quota detection. - Streamline regex logic in `parseRetryFromErrorText` for better readability. - Add unit tests for extreme cooldown values and free-tier exhaustion scenarios. * fix(429): drop over-broad /quota reached/ pattern, keep specific matches The bare /quota reached/ would also flag transient per-minute limits like 'request quota reached, retry in 60s' as quota_exhausted (multi-hour lock). The Antigravity message is still caught by /individual quota reached/. Added a regression assertion proving the transient case stays a rate_limit. --------- Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>
…ors hall Adds entries for diegosouzapw#3097, diegosouzapw#3101 (deepseek-web diegosouzapw#2942/diegosouzapw#2820), diegosouzapw#3104, diegosouzapw#3105, diegosouzapw#3107, diegosouzapw#3109, diegosouzapw#3111, diegosouzapw#3113, diegosouzapw#3115, diegosouzapw#3122, diegosouzapw#3125, diegosouzapw#3127, diegosouzapw#3129, plus a Contributors section crediting all v3.8.9 contributors. Stamps the 3.8.9 release date.
…ccounts for full reset window (diegosouzapw#3122) * fix(sse): bound Antigravity short-retry 429 loop per endpoint A persistent 429 on the short-retry branch (retryAfterMs ≤ 60s) looped forever on the same endpoint because the branch did `urlIndex--; continue` without checking the shared retry counter. Production log showed 77 consecutive 429s on one daily endpoint/account with zero fallback. Gate the short-retry branch on `retryAttemptsByUrl[urlIndex] < MAX_AUTO_RETRIES` (mirroring the already-bounded sibling), so a persistent 429 retries at most 3× per endpoint across all 3 base URLs then returns the 429 to the account- fallback layer. Regression test: 'bounds a persistent short-retry 429' in tests/unit/executor-antigravity.test.ts — asserts 12 total attempts (3 endpoints × 4) and a returned 429 with zero hang. * fix(sse): lock Antigravity quota-exhausted account for full reset window After the retry-loop bound, OmniRoute fell over to the next account but re-selected the exhausted one first on every subsequent request (~60s wasted per request). Root cause: the 429 body 'Individual quota reached. Contact your administrator to enable overages. Resets in 164h27m24s.' was not recognized as quota exhaustion, so the model was locked for only ~5s instead of the real 6.8-day reset window. Two detector fixes (mirrors Antigravity-Manager rate_limit.rs set_lockout_until): 1. classify429.ts — add QUOTA_PATTERNS: /individual quota reached/i, /quota reached/i, /enable overages/i so looksLikeQuotaExhausted() fires. 2. accountFallback.ts — same patterns in classifyErrorText(); extend parseRetryFromErrorText() to parse 'Resets? in XhYmZs' (reusing the existing computeDurationMs helper) so the exact reset duration reaches recordModelLockoutFailure as exactCooldownMs (uncapped, per user choice). The lockout machinery already stores until = now + cooldownMs with no clamp, bypasses getScaledCooldown when exactCooldownMs > 0, and keeps the longer of existing/new, so the full 164h window flows through intact. New patterns stay specific — plain 'too many requests'/'rate limit exceeded' messages still classify as rate_limit. Verified end-to-end against the real message: - classify429 → quota_exhausted - parseRetryFromErrorText → 592044000 ms (164h27m24s exactly) - checkFallbackError → usedUpstreamRetryHint: true, cooldownMs: 592044000 * refactor(account-fallback): simplify error parsing and add cooldown safety - Implement a 30-day cap on parsed retry durations to prevent indefinite account lockouts. - Replace manual string matching with `looksLikeQuotaExhausted` for more robust quota detection. - Streamline regex logic in `parseRetryFromErrorText` for better readability. - Add unit tests for extreme cooldown values and free-tier exhaustion scenarios. * fix(429): drop over-broad /quota reached/ pattern, keep specific matches The bare /quota reached/ would also flag transient per-minute limits like 'request quota reached, retry in 60s' as quota_exhausted (multi-hour lock). The Antigravity message is still caught by /individual quota reached/. Added a regression assertion proving the transient case stays a rate_limit. --------- Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>
…ors hall Adds entries for diegosouzapw#3097, diegosouzapw#3101 (deepseek-web diegosouzapw#2942/diegosouzapw#2820), diegosouzapw#3104, diegosouzapw#3105, diegosouzapw#3107, diegosouzapw#3109, diegosouzapw#3111, diegosouzapw#3113, diegosouzapw#3115, diegosouzapw#3122, diegosouzapw#3125, diegosouzapw#3127, diegosouzapw#3129, plus a Contributors section crediting all v3.8.9 contributors. Stamps the 3.8.9 release date.
…ccounts for full reset window (diegosouzapw#3122) * fix(sse): bound Antigravity short-retry 429 loop per endpoint A persistent 429 on the short-retry branch (retryAfterMs ≤ 60s) looped forever on the same endpoint because the branch did `urlIndex--; continue` without checking the shared retry counter. Production log showed 77 consecutive 429s on one daily endpoint/account with zero fallback. Gate the short-retry branch on `retryAttemptsByUrl[urlIndex] < MAX_AUTO_RETRIES` (mirroring the already-bounded sibling), so a persistent 429 retries at most 3× per endpoint across all 3 base URLs then returns the 429 to the account- fallback layer. Regression test: 'bounds a persistent short-retry 429' in tests/unit/executor-antigravity.test.ts — asserts 12 total attempts (3 endpoints × 4) and a returned 429 with zero hang. * fix(sse): lock Antigravity quota-exhausted account for full reset window After the retry-loop bound, OmniRoute fell over to the next account but re-selected the exhausted one first on every subsequent request (~60s wasted per request). Root cause: the 429 body 'Individual quota reached. Contact your administrator to enable overages. Resets in 164h27m24s.' was not recognized as quota exhaustion, so the model was locked for only ~5s instead of the real 6.8-day reset window. Two detector fixes (mirrors Antigravity-Manager rate_limit.rs set_lockout_until): 1. classify429.ts — add QUOTA_PATTERNS: /individual quota reached/i, /quota reached/i, /enable overages/i so looksLikeQuotaExhausted() fires. 2. accountFallback.ts — same patterns in classifyErrorText(); extend parseRetryFromErrorText() to parse 'Resets? in XhYmZs' (reusing the existing computeDurationMs helper) so the exact reset duration reaches recordModelLockoutFailure as exactCooldownMs (uncapped, per user choice). The lockout machinery already stores until = now + cooldownMs with no clamp, bypasses getScaledCooldown when exactCooldownMs > 0, and keeps the longer of existing/new, so the full 164h window flows through intact. New patterns stay specific — plain 'too many requests'/'rate limit exceeded' messages still classify as rate_limit. Verified end-to-end against the real message: - classify429 → quota_exhausted - parseRetryFromErrorText → 592044000 ms (164h27m24s exactly) - checkFallbackError → usedUpstreamRetryHint: true, cooldownMs: 592044000 * refactor(account-fallback): simplify error parsing and add cooldown safety - Implement a 30-day cap on parsed retry durations to prevent indefinite account lockouts. - Replace manual string matching with `looksLikeQuotaExhausted` for more robust quota detection. - Streamline regex logic in `parseRetryFromErrorText` for better readability. - Add unit tests for extreme cooldown values and free-tier exhaustion scenarios. * fix(429): drop over-broad /quota reached/ pattern, keep specific matches The bare /quota reached/ would also flag transient per-minute limits like 'request quota reached, retry in 60s' as quota_exhausted (multi-hour lock). The Antigravity message is still caught by /individual quota reached/. Added a regression assertion proving the transient case stays a rate_limit. --------- Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>
…ors hall Adds entries for diegosouzapw#3097, diegosouzapw#3101 (deepseek-web diegosouzapw#2942/diegosouzapw#2820), diegosouzapw#3104, diegosouzapw#3105, diegosouzapw#3107, diegosouzapw#3109, diegosouzapw#3111, diegosouzapw#3113, diegosouzapw#3115, diegosouzapw#3122, diegosouzapw#3125, diegosouzapw#3127, diegosouzapw#3129, plus a Contributors section crediting all v3.8.9 contributors. Stamps the 3.8.9 release date.
🚀 Motivation and Context
Fixes two critical gaps in Antigravity 429 handling:
Executor retry loop — A persistent 429 on the short-retry branch (
retryAfterMs ≤ 60s) looped forever on the same endpoint/account because the branch didurlIndex--; continuewithout checking the shared retry counter. Production log showed 77 consecutive 429s on one daily endpoint with zero fallback.Quota-exhausted lockout — After the retry-loop bound, OmniRoute fell over to the next account but re-selected the exhausted one first on every subsequent request (~60s wasted per request). The 429 body
"Individual quota reached. Contact your administrator to enable overages. Resets in 164h27m24s."was not recognized as quota exhaustion, so the model was locked for only ~5s instead of the real 6.8-day reset window.⚙️ Implementation Details
Commit 1 —
fix(sse): bound Antigravity short-retry 429 loop per endpointretryAttemptsByUrl[urlIndex] < MAX_AUTO_RETRIES(mirroring the already-bounded sibling)tests/unit/executor-antigravity.test.ts— asserts 12 total attempts (3 × 4) and a returned 429Commit 2 —
fix(sse): lock Antigravity quota-exhausted account for full reset windowclassify429.ts— addQUOTA_PATTERNS:/individual quota reached/i,/enable overages/iaccountFallback.ts— delegate tolooksLikeQuotaExhausted()for quota patterns; extendparseRetryFromErrorText()to parse"Resets? in XhYmZs"recordModelLockoutFailureasexactCooldownMs(uncapped, per user choice)classify429→quota_exhausted;parseRetryFromErrorText→ 592044000 ms (164h27m24s exactly);checkFallbackError→usedUpstreamRetryHint: trueMirrors Antigravity-Manager
src-tauri/src/proxy/rate_limit.rsset_lockout_untilbehavior. Patterns are specific — plain rate-limit messages still classify asrate_limit.Commit 3 —
refactor(sse): address review findings from PR #3122/quota reached/ipattern (retained only/individual quota reached/i)classifyErrorText()tolooksLikeQuotaExhausted()for quota patterns (eliminates duplication)computeDurationMs(prevents adversarial/buggy upstream from locking indefinitely)parseRetryFromErrorTextnested structure for readabilityCREDITS_EXHAUSTED_SIGNALS"free tier of the model has been exhausted"entry📋 Checklist for Reviewer
🔗 Related