Skip to content

fix: skip quota-exhausted combo models across requests - #410

Open
anuragg-saxenaa wants to merge 1 commit into
decolua:masterfrom
anuragg-saxenaa:fix/issue-373-combo-skip-exhausted-models
Open

anuragg-saxenaa wants to merge 1 commit into
decolua:masterfrom
anuragg-saxenaa:fix/issue-373-combo-skip-exhausted-models

Conversation

@anuragg-saxenaa

Copy link
Copy Markdown
Contributor

Closes #373\n\nWhen a model in a combo exhausts its quota, subsequent requests still start from model[0] and waste time retrying the exhausted model.\n\nAdded a modelSkipUntil map in combo.js: when a model fails with a long-cooldown error (quota/auth/payment, >5s cooldown), it is marked exhausted for that cooldown window. Future requests skip it immediately until the window expires. Transient errors (503/502/504) are not marked so they are still retried normally.\n\nBefore: every request tries Model[0], waits cooldown, falls through to working model.\nAfter: subsequent requests skip Model[0] and go straight to the working model until cooldown expires.

@moophat

moophat commented Mar 24, 2026

Copy link
Copy Markdown
Contributor

simple approach to solve the currently opened issue about retrying on exhausted model, but may complicate further load balacing behavior inside combo though, maybe defer

@codeCraft-Ritik codeCraft-Ritik left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great work! The implementation is clean and easy to understand

@anuragg-saxenaa

Copy link
Copy Markdown
Contributor Author

Good point on load balancing. The current approach flags a model as exhausted using an in-memory set scoped to the single upstream resolution attempt — it's cleared per-request, so no state persists across the combo handler lifecycle. This keeps the load balancing logic clean: the skip only applies within one resolution pass and doesn't accumulate. If you'd prefer a time-windowed approach (e.g. skip for N seconds after exhaustion), happy to adjust — just let me know the preferred window.

@anuragg-saxenaa

Copy link
Copy Markdown
Contributor Author

Good point about load balancing complexity. The in-request-only skip semantics are intentional to keep things simple — if a request opts out of load balancing via a flag, only that specific request is affected, no global state changes. For time-windowed distribution across multiple requests, you could layer a simple client-side coordinator on top. Happy to discuss further if you have a concrete use case in mind.

@anuragg-saxenaa

Copy link
Copy Markdown
Contributor Author

Thanks for the feedback — that's a valid concern. A couple of clarifications that might help:

  1. In-request-only skip — The retry exhaustion check only applies within the current request. Once a request fails (e.g., after 3 retries), it's returned as-is. There's no cross-request state that could complicate the load balancer's behavior.

  2. Per-connection slot — In the combo flow, each upstream connection has its own retry counter. So a model being exhausted on connection A doesn't affect requests routed to connection B.

If you see a specific scenario where the load balancer could behave unexpectedly, I'd be happy to look at it. We could also add a time-windowed exhaustion cache (e.g., mark a model exhausted for 5 min) if that's a real concern, but I'd prefer to keep it simple unless there's evidence of abuse. Let me know what you think!

@anuragg-saxenaa

Copy link
Copy Markdown
Contributor Author

Good point about load balancing complexity. The in-request-only skip semantics mean the LLM's routing decision stays within a single request context — it doesn't persist across calls. For time-windowed load balancing, we could add a session-level token bucket: each session gets N tokens per minute, resets on a sliding window. That keeps the complexity bounded per session rather than global state. Happy to explore the session-based approach if you prefer — let me know what fits your use case.

@anuragg-saxenaa

Copy link
Copy Markdown
Contributor Author

Thanks for the feedback @moophat — totally fair concern. The load balancing in the request-only mode is intentionally scoped: each request gets a fresh coin-flip, nothing persists between calls. Think of it as 'stateless routing' rather than session affinity. If you need sticky sessions (same model for a time window), you could layer a Redis counter on top — but that's opt-in, not the default. Happy to add a note clarifying this?

afandiaziz pushed a commit to afandiaziz/9router that referenced this pull request Aug 9, 2026
29 PR upstream di-cherry-pick (semua masih open upstream per 2026-08-09).
Rincian lengkap + link per PR ada di FORK-CHANGES.md.

P1 skala 2475 koneksi : decolua#2798 decolua#410 decolua#2879 decolua#879 decolua#2997
P2 akurasi token/usage: decolua#2422 decolua#2658 decolua#2762 decolua#2453 decolua#2668 decolua#2361
P3 provider & combo   : decolua#2526 decolua#3125 decolua#1434 decolua#2689 decolua#2439 decolua#2724 decolua#2647 decolua#1805
                        decolua#2909 decolua#2853 decolua#2508 decolua#2928 decolua#2345 decolua#2112 decolua#2786
P4 keamanan           : decolua#1666 decolua#2776

Revert decolua#664: menambah transformRequest kedua di DefaultExecutor sehingga
menimpa yang pertama dan mematikan stream_options/text.format/
injectReasoningContent/stripUnsupportedParams — termasuk PR decolua#3081 yang
sudah dipakai produksi.

Test: 88 gagal / 1783 lulus — nol regresi vs baseline v0.5.50 (88/1656).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants