Repository navigation
feat(sse): opt-in per-model concurrency caps per connection - #13700
Notaloop763 wants to merge 2 commits into
Conversation
|
Thanks for this — clean extension of the existing hierarchical admission gate Two small non-blocking notes for a possible follow-up: (1)
This looks merge-ready to us pending the GitHub Actions checks finishing green (only Mergify |
|
Re-homed to |
Adds optional per-model concurrency caps on a provider connection, enforced by the chatCore admission gate alongside the existing per-account cap. Squashed from the original series (incl. merge-time helper extraction): - feat(sse): add opt-in per-model concurrency caps per connection - feat(dashboard): use i18n keys for per-model concurrency editor - docs(changelog): add fragment for per-model concurrency caps - extract per-model concurrency UI/gate helpers under file-size ceilings - test(sse): align admission-gate source assertions with resolveModelSemaphore helper - fix(i18n): mirror per-model concurrency keys into all locales - chore(quality): rebaseline per-model concurrency file-size ceilings - fix(shared): move model concurrency bounds to a server-free leaf
…axWaitMs - parseModelConcurrencyInput now reports the offending entry instead of a hard-coded English sentence; the modal renders it through the new providers.rateLimitOverridesModelConcurrencyInvalid i18n key. - sanitizeRateLimitOverrides accepts executionMaxWaitMs, matching the Zod schema; previously a PATCH carrying it passed validation and then failed with "Refusing to persist rateLimitOverrides with rejected keys". - buildRateLimitOverridesFromForm carries over an API-set executionMaxWaitMs (the form has no field for it) so a dashboard save no longer wipes it. - Tests: form-helper coverage in model-concurrency-input.test.ts and the allowlist regression in columns-validation.test.ts.
56bed30 to
c1d61a9
Compare
Summary
Adds an optional, generic per-connection, per-model upstream concurrency cap via
rateLimitOverrides.modelConcurrency. The motivating case is Z.AI PAYG, which has per-model limits, but nothing here is provider-specific: there are no hard-coded provider limits and no limit discovery. OmniRoute just queues excess requests locally against exact ceilings the operator configures.{ "rateLimitOverrides": { "maxConcurrent": 4, "modelConcurrency": { "glm-5": 1, "glm-4.7": 3 } } }No DB migration: it reuses the existing
rate_limit_overrides_jsoncolumn.Behavior
glm-5). A client-sidezai/glm-5alias does not match. Caps are positive integers 1–10,000; keys are ≤128 chars.maxConcurrent. Both gates join the same atomic composite acquisition (global → provider → account → model), so the stricter limit wins.SEMAPHORE_TIMEOUT/SEMAPHORE_QUEUE_FULLadmission errors, and no new error codes. A saturated model gate never disables the provider or creates a model lockout; 429/cooldown/fallback stay the upstream backstop.modelConcurrency.<key>) happens only on write, so operator intent is never dropped silently.model=capper line). It loads the stored map and writes it back, so maps set through the API survive unrelated saves. A malformed entry blocks the save with a localized error (providers.rateLimitOverridesModelConcurrencyInvalid).docs/architecture/RESILIENCE_GUIDE.md.Incidental fix:
executionMaxWaitMscould not be persistedupdateProviderConnectionSchemaacceptsrateLimitOverrides.executionMaxWaitMs, but the DB-side allowlist insanitizeRateLimitOverridesdid not. A PATCH carrying it passed Zod and then threwRefusing to persist rateLimitOverrides with rejected keys: executionMaxWaitMs. This is the same class of bug as the #11251maxWaitMsfollow-up. The allowlist now matches the schema, and the dashboard form carries an API-setexecutionMaxWaitMsthrough saves, since the form has no field for it. Regression tests are intests/unit/columns-validation.test.tsand the form/modal tests below.Out of scope (deliberately excluded)
Hard-coded provider limits, limit discovery/scraping, new executors, global behavior changes, retries, and billing/quota/cooldown changes.
Validation
Rebased onto
release/v3.8.52@23a1148486. Run from a worktree on Linux, Node 24.21.0:npm run test:scopedresolves to the full suite because hub files (the locale catalogs) are touched, so the full unit suite was run as well. Its one failure,tests/unit/verified-connection-activation-11446.test.ts("testSingleConnection activates a connection once a test actually passes", which times out after about 16 s withupstream_error), reproduces identically on a cleanrelease/v3.8.52@23a1148486checkout. It is unrelated to this PR: the test stubsglobalThis.fetch, which the validation probe appears to bypass. It is not yet listed in #15306.Changed/added test files
tests/unit/model-concurrency-gate.test.ts(new)tests/unit/model-concurrency-input.test.ts(new: parser and form-helper coverage)tests/unit/provider-patch-model-concurrency.test.ts(new)tests/unit/dashboard/edit-connection-modal-model-concurrency.test.tsx(new, Vitest)tests/unit/account-concurrency-cap.test.ts(extended)tests/unit/chatcore-hierarchical-admission.test.ts(extended)tests/unit/columns-validation.test.ts(extended:executionMaxWaitMsregression)File-size baseline:
open-sse/handlers/chatCore.ts(+14) andsrc/sse/services/auth.ts(+3) keep only call-site wiring; the logic lives in non-frozen helpers. The justification is inconfig/quality/file-size-baseline.jsonunder_rebaseline_2026_09_25_13700_per_model_concurrency, re-measured after the rebase.Commits
The branch was rebuilt on
release/v3.8.52. Its earlierrelease/v3.8.51sync merges contained real conflict-resolution content (the helper extraction), which a plain rebase would have dropped. So the original series is squashed into one feature commit, with the follow-up on top:feat(sse): opt-in per-model concurrency caps per connectionfix(dashboard): localize per-model cap save error and keep executionMaxWaitMs