Repository navigation
fix: allow combo fallback on context overflow 400 errors - #1331
diegosouzapw merged 2 commits into
Conversation
When a combo routes through models with different context limits, a 400 'input is too long' or 'improperly formed request' error from one model should trigger fallback to the next model rather than failing immediately. Previously checkFallbackError() returned shouldFallback:false for ALL HTTP 400 responses, preventing combo models with larger context windows from being tried. Adds pattern-based detection for context overflow and malformed request errors within 400 responses, returning shouldFallback:true with MODEL_CAPACITY reason so combos can proceed to the next model. Closes diegosouzapw#1329
There was a problem hiding this comment.
Code Review
This pull request updates the checkFallbackError function to allow fallbacks for specific 400 Bad Request errors, such as context overflows or malformed requests, which might succeed on alternative models. Feedback was provided regarding an unused variable and the efficiency of defining regex patterns within the function scope.
| const lowerError = errorStr.toLowerCase(); | ||
|
|
||
| // Context overflow: the prompt exceeds the model's maximum context length. |
There was a problem hiding this comment.
The lowerError variable is unused in this block because the regex patterns below use the case-insensitive flag (/i) and test against errorStr directly. Removing this line avoids redundant string conversion.
Additionally, consider moving CONTEXT_OVERFLOW_PATTERNS and MALFORMED_REQUEST_PATTERNS to the top level of the file (outside the function). Defining them inside the function causes them to be re-allocated on every 400 error, which is inefficient for a function that may be called frequently during high-traffic periods.
| const lowerError = errorStr.toLowerCase(); | |
| // Context overflow: the prompt exceeds the model's maximum context length. | |
| // Context overflow: the prompt exceeds the model's maximum context length. |
- Move CONTEXT_OVERFLOW_PATTERNS and MALFORMED_REQUEST_PATTERNS to module level (avoids re-allocation on every 400 in hot path) - Remove unused lowerError variable (regexes already use /i flag)
CI Status SummaryAll CI failures on this PR are pre-existing on the Comparison with
|
| Check | main branch |
This PR | Caused by our changes? |
|---|---|---|---|
| Build | ✅ | ✅ | — |
| Lint | ❌ typecheck:noimplicit:core |
❌ Same errors | No — pre-existing yazl, github.ts, cloudflare-ai.ts etc. |
| Unit Tests | ❌ tests #368, #1546 | ❌ Same 2 tests | No |
| E2E 3/4 | ❌ memory-settings.spec.ts |
❌ Same test | No |
| E2E 4/4 | ❌ proxy-registry + skills-marketplace |
❌ Same tests | No |
| Integration | ❌ timeout 10m | ❌ Same timeout | No |
| Security Scans | ❌ TruffleHog | ❌ Snyk | No — infra issue |
| PR Test Policy | N/A (push) | ❌ no test files | Policy check only |
Gemini Review: Addressed ✅
The review comment about lowerError being unused and patterns being inside the function was already addressed in commit ce1a92c — moved CONTEXT_OVERFLOW_PATTERNS and MALFORMED_REQUEST_PATTERNS to module level, removed unused variable.
|
Thanks @oyi77 for this great contribution! 🎉 The combo fallback fix for context overflow errors has been integrated into the release/v3.6.7 branch and will be part of the next release. We appreciate your effort! |
When handleSingleModelChat exhausts all accounts for a provider (all quota-preflight blocked or rate-limited), it returns a 503 without making an upstream request. The combo routing engine treated this 503 as a non-fallbackable error and aborted, never trying subsequent models. Adds isAllAccountsRateLimitedResponse() that detects the handleNoCredentials signal (503 + JSON + "unavailable" in message) and allows combo to proceed to the next model instead of returning immediately. Applied in both handleComboChat and handleRoundRobinCombo. Patterns defined at module level to avoid re-allocation on every call (per PR diegosouzapw#1331 code review style). Closes diegosouzapw#1331 (related pattern)
When handleSingleModelChat exhausts all accounts for a provider (all quota-preflight blocked or rate-limited), it returns a 503 without making an upstream request. The combo routing engine treated this 503 as a non-fallbackable error and aborted, never trying subsequent models. Adds isAllAccountsRateLimitedResponse() that detects the handleNoCredentials signal (503 + JSON + "unavailable" in message) and allows combo to proceed to the next model instead of returning immediately. Applied in both handleComboChat and handleRoundRobinCombo. Patterns defined at module level to avoid re-allocation on every call (per PR diegosouzapw#1331 code review style). Closes diegosouzapw#1331 (related pattern)
…w#1331) Integrated into release/v3.6.7
…w#1331) Integrated into release/v3.6.7
Summary
Fixes combo routing fallback behavior for HTTP 400 context overflow errors. When a combo routes through models with different context limits, a 400 "input is too long" or "improperly formed request" from one model should trigger fallback to the next model rather than failing immediately.
Closes #1329
Problem
checkFallbackError()inopen-sse/services/accountFallback.tsreturnsshouldFallback: falsefor all HTTP 400 responses. This means when a combo's first model rejects a request due to context length (e.g., a 128K model getting a request that exceeds its window), the entire combo fails — even if subsequent models in the combo have larger context windows (200K, 1M, etc.).This commonly manifests during OpenCode compaction cycles where accumulated context grows large and hits the limit of the first model in a combo chain.
Root Cause
Line 863-866 (before this PR):
The blanket
shouldFallback: falsefor all 400s assumes the same request will always fail identically on any account. This is correct for malformed API requests but incorrect for context overflow — different models have different context limits.Fix
Adds pattern-based detection for two categories of 400 errors that should trigger fallback:
"input is too long","too many tokens","prompt is too long","context window","token limit", etc."improperly formed request","invalid message format","messages must alternate","empty message/content"When detected, returns
shouldFallback: truewithMODEL_CAPACITYreason (zero cooldown — the model's context limit won't change by waiting).Unmatched 400 errors continue to return
shouldFallback: falseas before.Testing
Changes
open-sse/services/accountFallback.ts— Added context overflow and malformed request pattern detection before the blanket 400 rejection incheckFallbackError()