Skip to content

feat(auth): model-level rate limit locking for multi-bucket providers - #120

Merged
decolua merged 1 commit into
decolua:masterfrom
rothnic:feat/granular-model-rate-limiting
Feb 15, 2026
Merged

decolua merged 1 commit into
decolua:masterfrom
rothnic:feat/granular-model-rate-limiting

Conversation

@rothnic

@rothnic rothnic commented Feb 13, 2026

Copy link
Copy Markdown
Contributor

Summary

Adds model-level rate limit locking for providers with separate per-model quota buckets (e.g. Antigravity).

Closes #110

Problem

Antigravity (and Vertex AI) maintain independent quota buckets per model family — Claude models and Gemini models have separate rate limits under the same account. When 9router receives a 429 Too Many Requests for any model, it currently locks the entire account in the database, preventing access to all models including those with full quota remaining.

Example: A 429 on claude-opus (Account A) locks Account A entirely. Subsequent gemini-pro requests skip Account A and burn Account B's quota, even though Account A's Gemini bucket is 100% available.

Solution

In-memory model-aware locking that tracks rate limits per connectionId:model pair rather than per account.

How it works

  1. On 429 (multi-bucket providers only): Instead of locking the account in the database, we set a temporary in-memory lock for that specific model (default 5 minutes or the calculated backoff duration).
  2. On request: The connection filter checks both the standard account-level DB lock AND the in-memory model lock. If connectionId:model is locked, we skip that account for that model only — other models remain accessible.
  3. Fallback behavior: When all accounts are model-locked for a specific model, returns a retryAfter response with a short 1-minute retry window (vs the standard account-level backoff).

Design decisions

  • In-memory only — No database schema changes. Locks clear on restart. Low risk, zero migration.
  • Provider-gated — Only enabled for providers in MULTI_BUCKET_PROVIDERS set (currently antigravity). Other providers use the existing account-level locking unchanged.
  • Backward compatible — The model parameter defaults to null. When not provided (e.g. from combo handlers), behavior is identical to before.

Changes

File Change
src/sse/services/auth.js Add modelLocks Map, isModelLocked(), lockModel(), isMultiBucketProvider(). Update getProviderCredentials() and markAccountUnavailable() to accept model param and apply model-level filtering.
src/sse/handlers/chat.js Pass model string to getProviderCredentials() and markAccountUnavailable() calls.

Testing

Validated in production with 2 Antigravity accounts serving mixed Claude/Gemini workloads. After applying this patch:

  • Claude 429 on Account A only blocks Claude on Account A
  • Gemini requests continue routing to Account A normally
  • Standard account-level locking remains unchanged for non-Antigravity providers

Future work (from #110)

This PR implements the core locking mechanism. Potential follow-ups:

  • Configuration: Add rateLimitScope: "model" | "account" to Provider settings in the dashboard
  • Regex buckets: Allow defining model groups (e.g., ^claude-* = Bucket A) for more granular control
  • Persisted state: Optionally store model locks in the database for persistence across restarts

…ders

Providers like Antigravity maintain separate quota buckets per model family
(e.g. Claude vs Gemini). A 429 on claude-opus previously locked the entire
account, preventing gemini-pro requests even though its quota was full.

This adds in-memory per-model locking so that only the specific model is
skipped during account selection while other models remain accessible.

Changes:
- Add model-aware lock tracking in auth.js (Map<connectionId:model, expiry>)
- Pass model context from chat handler to auth service
- Multi-bucket behavior gated to known providers (MULTI_BUCKET_PROVIDERS set)
- No database schema changes — locks are in-memory and clear on restart

Closes decolua#110
@rothnic

rothnic commented Feb 13, 2026

Copy link
Copy Markdown
Contributor Author

Ideally this would be stored in the database, but i wanted to avoid making too extreme of changes since I wasn't sure about any other preferred alternative approaches. This at least is a pretty simple and focused fix.

I did notice that some previous discussion seemed to think that antigravity is token limited, but from everything I've seen it is more request or prompt limited. Maybe could use some investigation whether that is true and if the available usage api could be leveraged in some way to have more intelligent fallbacks or routing. For example, i wouldn't mind being able to fallback from claude models to gemini model at some percentage of usage so i can retain some opus 4.6 usage for opencode's antigravity auth use.

@Wladefant

Copy link
Copy Markdown

GH Copilot is request based for sure. I was told that Antigravity work by analyzing work done, and seems like you have only a limited work done requests. But then some told that it is token based.

And I agree with you completely. A much smarter routing system is needed. that would always try to use the models up for 50 percent first and after that they can go to opus once again. And also of course the routing really needs to account the rollover time window, so that these models are used first that rollover earliest

@decolua
decolua merged commit 202fee7 into decolua:master Feb 15, 2026
LinearSakana pushed a commit to LinearSakana/9router that referenced this pull request Feb 15, 2026
…ders (decolua#120)

Providers like Antigravity maintain separate quota buckets per model family
(e.g. Claude vs Gemini). A 429 on claude-opus previously locked the entire
account, preventing gemini-pro requests even though its quota was full.

This adds in-memory per-model locking so that only the specific model is
skipped during account selection while other models remain accessible.

Changes:
- Add model-aware lock tracking in auth.js (Map<connectionId:model, expiry>)
- Pass model context from chat handler to auth service
- Multi-bucket behavior gated to known providers (MULTI_BUCKET_PROVIDERS set)
- No database schema changes — locks are in-memory and clear on restart

Closes decolua#110

(cherry picked from commit 202fee7)
@rothnic

rothnic commented Feb 17, 2026

Copy link
Copy Markdown
Contributor Author

GH Copilot is request based for sure. I was told that Antigravity work by analyzing work done, and seems like you have only a limited work done requests. But then some told that it is token based.

And I agree with you completely. A much smarter routing system is needed. that would always try to use the models up for 50 percent first and after that they can go to opus once again. And also of course the routing really needs to account the rollover time window, so that these models are used first that rollover earliest

Sorry for the slow follow-up, ended up getting banned by antigravity, so i've been figuring out what the next best options are. From my understanding antigravity was request-based, but i could be misunderstanding. I thought it was N requests per 6 hours or something like that. Copilot it like 1200 requests per month, no hour-based timeout, but is only $40 and has free gpt-5-mini access.

At one point when i was having issues, i almost put litellm in front if 9router until i realized it wasn't going to help. It did kind of make me wonder if using something like litellm internally would be a good idea since you could kind of provide its configuration as an advanced mode.

I ended up grabbing a month of kimi for $40 for now until i have more time to figure out the next best option.

@Wladefant

Copy link
Copy Markdown

no way, all my accoutns are banned now from day to the other. That is crazy. Is there a way to fix this?

@rothnic

rothnic commented Feb 19, 2026

Copy link
Copy Markdown
Contributor Author

no way, all my accoutns are banned now from day to the other. That is crazy. Is there a way to fix this?

Don't think so... there are a bunch of people that got banned over the past week or so. It was nice while it lasted.

kwanLeeFrmVi pushed a commit to kwanLeeFrmVi/9router that referenced this pull request Feb 26, 2026
…ders (decolua#120)

Providers like Antigravity maintain separate quota buckets per model family
(e.g. Claude vs Gemini). A 429 on claude-opus previously locked the entire
account, preventing gemini-pro requests even though its quota was full.

This adds in-memory per-model locking so that only the specific model is
skipped during account selection while other models remain accessible.

Changes:
- Add model-aware lock tracking in auth.js (Map<connectionId:model, expiry>)
- Pass model context from chat handler to auth service
- Multi-bucket behavior gated to known providers (MULTI_BUCKET_PROVIDERS set)
- No database schema changes — locks are in-memory and clear on restart

Closes decolua#110
involvex added a commit to involvex/involvex-claude-router that referenced this pull request Feb 28, 2026
### Bug Fixes

* remove duplicated opencode prefix from model IDs and handle it in executor ([180f884](180f884))
* update opencode endpoint and model IDs ([6be8a03](6be8a03))

### Features

* add opencode provider ([03c846e](03c846e))

## [0.2.96](5645d0a...v0.2.96) (2026-02-24)

### Bug Fixes

*  GitHub Copilot model ([95fd950](95fd950))
* **auth:** allow HTTP for local network ([0a394d0](0a394d0))
* **auth:** prevent auto-login after logout ([49df3dc](49df3dc))
* **codex:** use user-agent detection for Droid CLI compatibility ([8c6e3b8](8c6e3b8))
* Correct indentation for clarity in chatCore and claude-to-openai response handlers ([fa06226](fa06226))
* correct token extraction for Claude non-streaming responses ([decolua#131](https://github.com/involvex/involvex-claude-router/issues/131)) ([9fbd6e6](9fbd6e6))
* **dashboard:** resolve 'Provider not found' for free providers ([45a4d3b](45a4d3b))
* **db:** improve error handling and null checks ([e6ef852](e6ef852))
* **gemini:** improve base64 image data parsing ([5645d0a](5645d0a))
* **github:** Implement dynamic fallback for Codex models requiring /responses endpoint ([decolua#127](https://github.com/involvex/involvex-claude-router/issues/127)) ([6913129](6913129))
* improve code formatting and reduce auto-refresh interval ([7f71916](7f71916))
* improve cursor auto-import reliability on macOS ([decolua#161](https://github.com/involvex/involvex-claude-router/issues/161)) ([d7e06c3](d7e06c3))
* **login:** avoid infinite loading on settings fetch failure ([01c9410](01c9410))
* **open-sse:** emit [DONE] in passthrough SSE mode ([decolua#142](https://github.com/involvex/involvex-claude-router/issues/142)) ([b9a6979](b9a6979))
* prevent race conditions in sticky round-robin ([3ad2f8d](3ad2f8d))
* resolve SonarQube findings and Next.js Image warnings ([7058b06](7058b06))
* update Codex executor for gpt-5.3-codex support ([d7d5dc9](d7d5dc9))

### Features

* add /v1/embeddings endpoint (OpenAI-compatible) ([decolua#146](https://github.com/involvex/involvex-claude-router/issues/146)) ([e1b8361](e1b8361)), closes [decolua#117](https://github.com/involvex/involvex-claude-router/issues/117)
* Add Anthropic Compatible provider support ([da5bdef](da5bdef))
* add API endpoint dimension to usage statistics dashboard ([decolua#152](https://github.com/involvex/involvex-claude-router/issues/152)) ([806bd4a](806bd4a))
* add Claude Opus 4.6 to GitHub Copilot provider ([decolua#97](https://github.com/involvex/involvex-claude-router/issues/97)) ([3d60597](3d60597))
* Add Claude Sonnet 4.6 to GitHub Copilot ([decolua#149](https://github.com/involvex/involvex-claude-router/issues/149)) ([4e2a3f8](4e2a3f8))
* add CLI entry points for claude-router commands ([62c632d](62c632d))
* add enable/disable toggle for provider connections ([ed796d2](ed796d2))
* add Gemini 3.1 Pro models to provider ([f2025cc](f2025cc))
* add Gemini embeddings support + Letta compatibility fixes ([a57a8ce](a57a8ce)), closes [decolua#148](decolua#148)
* add GLM 5 and MiniMax M2.5 models to providerModels.js; add Claude Sonnet 4.6 to CLI tools ([e1e5a81](e1e5a81))
* add GLM Coding (China) provider and Usage by API Keys statistics ([1ae4e31](1ae4e31))
* add GPT 4o to GitHub Copilot provider ([decolua#98](https://github.com/involvex/involvex-claude-router/issues/98)) ([c090bb0](c090bb0))
* add GPT 5.3 Codex Spark model to pricing and provider models ([decolua#133](https://github.com/involvex/involvex-claude-router/issues/133)) ([c7d4410](c7d4410))
* Add GPT 5.3 Codex to GitHub Copilot ([decolua#150](https://github.com/involvex/involvex-claude-router/issues/150)) ([c4aa424](c4aa424))
* add GPT-3.5 Turbo to GitHub Copilot provider ([e3dbd44](e3dbd44))
* add GPT-4 to GitHub Copilot provider ([6ade8ef](6ade8ef))
* add GPT-4o mini to GitHub Copilot provider ([053e490](053e490))
* add models management and router process control commands ([c65fd42](c65fd42))
* Add OpenAI-compatible provider nodes ([0a28f9f](0a28f9f))
* add password change functionality and dependencies ([23cfb19](23cfb19))
* add pause/resume functionality for API keys ([decolua#158](https://github.com/involvex/involvex-claude-router/issues/158)) ([73388a0](73388a0))
* add Qwen3.5 Coder Model configuration ([decolua#156](https://github.com/involvex/involvex-claude-router/issues/156)) ([f933dd9](f933dd9))
* add request logging functionality and usage metrics display ([e476907](e476907))
* add round-robin routing strategy ([9ebd7d3](9ebd7d3))
* add sticky round-robin routing strategy ([4f292aa](4f292aa))
* add support for GLM 5 (if) ([decolua#123](https://github.com/involvex/involvex-claude-router/issues/123)) ([03ab554](03ab554))
* add URL-based tab state persistence in usage page ([decolua#129](https://github.com/involvex/involvex-claude-router/issues/129)) ([6caef7f](6caef7f))
* allow custom user data directory via DATA_DIR environment variable ([d83bd86](d83bd86))
* **antigravity:** initial steps for Antigravity anti-ban alignment ([a229d79](a229d79)), closes [decolua#141](decolua#141)
* **antigravity:** integrate Antigravity tool with MITM support and update CLI tools ([2e854bd](2e854bd))
* **auth:** add model-level rate limit locking for multi-bucket providers ([decolua#120](https://github.com/involvex/involvex-claude-router/issues/120)) ([202fee7](202fee7)), closes [decolua#110](https://github.com/involvex/involvex-claude-router/issues/110)
* **auth:** Enhance authentication flow and settings management ([249fc28](249fc28))
* **cli-tools:** update CLI tools and add new models ([a2122e3](a2122e3))
* **cli-tools:** update default local endpoint port to 20128 ([6c41573](6c41573))
* **cli:** add new CLI package with basic scaffolding ([9cf4628](9cf4628))
* **cloud:** harden sync/auth flow, SSE fallback, and update changelog ([3d43983](3d43983))
* **codex:** add GPT 5.3, fix API translation, add thinking levels ([127475d](127475d))
* **codex:** Cursor compatibility + Next.js 16 proxy migration ([1c6dd6d](1c6dd6d))
* **codex:** Cursor compatibility + Next.js 16 proxy migration ([7b864a9](7b864a9))
* **codex:** Cursor compatibility + Next.js 16 proxy migration ([e9b0a73](e9b0a73))
* **config:** add Cloudflare MCP server and account ID configuration ([67297ea](67297ea))
* **cursor:** Add cursor Provider ([0a026c7](0a026c7))
* **cursor:** Integrate Cursor IDE support with OAuth import token flow ([137f315](137f315))
* **docker:** add Docker setup, environment examples, and architecture docs ([5e4a15b](5e4a15b))
* enhance disconnect handling and request tracking in chatCore.js ([decolua#126](https://github.com/involvex/involvex-claude-router/issues/126)) ([3d29b86](3d29b86))
* enhance request handling and error management in chatCore and streamToJsonConverter ([e2db638](e2db638))
* enhance usage stats with sortable columns and improved data handling ([bf6e09b](bf6e09b))
* Enhance usage tracking across response handlers ([a33924b](a33924b))
* **executors:**  Improved UI components for displaying provider limits and usage statistics in the dashboard. ([32aefe5](32aefe5))
* **iflow:** add IFlowExecutor with HMAC-SHA256 signature and enable models ([bd23ab4](bd23ab4))
* **iflow:** add kimi-k2.5 model support ([9e357a7](9e357a7))
* implement API key requirement toggle ([4cf25dc](4cf25dc))
* Implement buffer addition to usage tracking for improved context handling ([7881db8](7881db8))
* implement lazy loading for UsagePage with suspense fallback ([decolua#136](https://github.com/involvex/involvex-claude-router/issues/136)) ([05b09e6](05b09e6))
* implement provider connection reordering on create, update, and delete ([f2abcc6](f2abcc6))
* implement real project ID fetching for Antigravity ([decolua#170](https://github.com/involvex/involvex-claude-router/issues/170)) ([ea67742](ea67742))
* implement request tracking and enhance usage stats display ([e4f92cd](e4f92cd))
* implement usage tracking for AI requests ([9c3d6f4](9c3d6f4))
* Improve Antigravity quota monitoring and fix Droid CLI compatibility ([3c65e0c](3c65e0c))
* **open-sse:** add Claude Sonnet 4.6 ([b057c43](b057c43))
* OpenAI compatibility improvements & build fixes ([d9b8e48](d9b8e48)), closes [decolua#18](https://github.com/involvex/involvex-claude-router/issues/18)
* **provider:** add free providers and enhance error handling ([bdbe816](bdbe816))
* **providers:** add Minimax Coding (China) provider ([7c609d7](7c609d7))
* **providers:** add provider icons to dashboard ([60bd686](60bd686))
* **providers:** auto-validate API keys on save ([b275dfd](b275dfd))
* rename app to involvex-claude-router and add lenient JSON parsing ([bec206d](bec206d))
* **responses:** respect client streaming preference + string input support ([decolua#121](https://github.com/involvex/involvex-claude-router/issues/121)) ([ac7cedd](ac7cedd))
* **translator:** add thinking parameter support in OpenAI → Claude ([54e01d6](54e01d6))
* **ui:** add cost tracking to usage dashboard and pricing settings ([f302c88](f302c88))
* **ui:** add model support for custom providers and improve UX ([a7a52be](a7a52be))
* Update response handling and logging for improved usage tracking ([df0e1d6](df0e1d6))
* **usage:** implement cost tracking backend and pricing configuration ([a36afaa](a36afaa))
idbara pushed a commit to idbara/9router that referenced this pull request Apr 23, 2026
…ders (decolua#120)

Providers like Antigravity maintain separate quota buckets per model family
(e.g. Claude vs Gemini). A 429 on claude-opus previously locked the entire
account, preventing gemini-pro requests even though its quota was full.

This adds in-memory per-model locking so that only the specific model is
skipped during account selection while other models remain accessible.

Changes:
- Add model-aware lock tracking in auth.js (Map<connectionId:model, expiry>)
- Pass model context from chat handler to auth service
- Multi-bucket behavior gated to known providers (MULTI_BUCKET_PROVIDERS set)
- No database schema changes — locks are in-memory and clear on restart

Closes decolua#110
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Feature Request: Granular Rate Limit Locking for Multi-Bucket Providers (Antigravity/Vertex)

3 participants