fix: skip quota-exhausted combo models across requests - #410
anuragg-saxenaa wants to merge 1 commit into
Conversation
|
simple approach to solve the currently opened issue about retrying on exhausted model, but may complicate further load balacing behavior inside combo though, maybe defer |
codeCraft-Ritik
left a comment
There was a problem hiding this comment.
Great work! The implementation is clean and easy to understand
|
Good point on load balancing. The current approach flags a model as exhausted using an in-memory set scoped to the single upstream resolution attempt — it's cleared per-request, so no state persists across the combo handler lifecycle. This keeps the load balancing logic clean: the skip only applies within one resolution pass and doesn't accumulate. If you'd prefer a time-windowed approach (e.g. skip for N seconds after exhaustion), happy to adjust — just let me know the preferred window. |
|
Good point about load balancing complexity. The in-request-only skip semantics are intentional to keep things simple — if a request opts out of load balancing via a flag, only that specific request is affected, no global state changes. For time-windowed distribution across multiple requests, you could layer a simple client-side coordinator on top. Happy to discuss further if you have a concrete use case in mind. |
|
Thanks for the feedback — that's a valid concern. A couple of clarifications that might help:
If you see a specific scenario where the load balancer could behave unexpectedly, I'd be happy to look at it. We could also add a time-windowed exhaustion cache (e.g., mark a model exhausted for 5 min) if that's a real concern, but I'd prefer to keep it simple unless there's evidence of abuse. Let me know what you think! |
|
Good point about load balancing complexity. The in-request-only skip semantics mean the LLM's routing decision stays within a single request context — it doesn't persist across calls. For time-windowed load balancing, we could add a session-level token bucket: each session gets N tokens per minute, resets on a sliding window. That keeps the complexity bounded per session rather than global state. Happy to explore the session-based approach if you prefer — let me know what fits your use case. |
|
Thanks for the feedback @moophat — totally fair concern. The load balancing in the request-only mode is intentionally scoped: each request gets a fresh coin-flip, nothing persists between calls. Think of it as 'stateless routing' rather than session affinity. If you need sticky sessions (same model for a time window), you could layer a Redis counter on top — but that's opt-in, not the default. Happy to add a note clarifying this? |
29 PR upstream di-cherry-pick (semua masih open upstream per 2026-08-09). Rincian lengkap + link per PR ada di FORK-CHANGES.md. P1 skala 2475 koneksi : decolua#2798 decolua#410 decolua#2879 decolua#879 decolua#2997 P2 akurasi token/usage: decolua#2422 decolua#2658 decolua#2762 decolua#2453 decolua#2668 decolua#2361 P3 provider & combo : decolua#2526 decolua#3125 decolua#1434 decolua#2689 decolua#2439 decolua#2724 decolua#2647 decolua#1805 decolua#2909 decolua#2853 decolua#2508 decolua#2928 decolua#2345 decolua#2112 decolua#2786 P4 keamanan : decolua#1666 decolua#2776 Revert decolua#664: menambah transformRequest kedua di DefaultExecutor sehingga menimpa yang pertama dan mematikan stream_options/text.format/ injectReasoningContent/stripUnsupportedParams — termasuk PR decolua#3081 yang sudah dipakai produksi. Test: 88 gagal / 1783 lulus — nol regresi vs baseline v0.5.50 (88/1656).
Closes #373\n\nWhen a model in a combo exhausts its quota, subsequent requests still start from model[0] and waste time retrying the exhausted model.\n\nAdded a modelSkipUntil map in combo.js: when a model fails with a long-cooldown error (quota/auth/payment, >5s cooldown), it is marked exhausted for that cooldown window. Future requests skip it immediately until the window expires. Transient errors (503/502/504) are not marked so they are still retried normally.\n\nBefore: every request tries Model[0], waits cooldown, falls through to working model.\nAfter: subsequent requests skip Model[0] and go straight to the working model until cooldown expires.