Skip to content

feat(sse): opt-in per-model concurrency caps per connection - #13700

Open
Notaloop763 wants to merge 2 commits into
diegosouzapw:release/v3.8.52from
Notaloop763:feat/per-model-concurrency
Open

Notaloop763 wants to merge 2 commits into
diegosouzapw:release/v3.8.52from
Notaloop763:feat/per-model-concurrency

Conversation

@Notaloop763

@Notaloop763 Notaloop763 commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Adds an optional, generic per-connection, per-model upstream concurrency cap via rateLimitOverrides.modelConcurrency. The motivating case is Z.AI PAYG, which has per-model limits, but nothing here is provider-specific: there are no hard-coded provider limits and no limit discovery. OmniRoute just queues excess requests locally against exact ceilings the operator configures.

{ "rateLimitOverrides": { "maxConcurrent": 4, "modelConcurrency": { "glm-5": 1, "glm-4.7": 3 } } }

No DB migration: it reuses the existing rate_limit_overrides_json column.

⚠️ base-red inherited: #15306. The failures tracked there are pre-existing on release/v3.8.52 and don't touch this PR's paths.

Behavior

  • Opt-in; absent config = unchanged. With no map (or a blank dashboard field), no model gate is added.
  • Exact model-key match on the model string the executor receives after routing resolution, which is the bare upstream id (e.g. glm-5). A client-side zai/glm-5 alias does not match. Caps are positive integers 1–10,000; keys are ≤128 chars.
  • Composes with maxConcurrent. Both gates join the same atomic composite acquisition (global → provider → account → model), so the stricter limit wins.
  • Same queue/timeout semantics as the existing gates: typed SEMAPHORE_TIMEOUT / SEMAPHORE_QUEUE_FULL admission errors, and no new error codes. A saturated model gate never disables the provider or creates a model lockout; 429/cooldown/fallback stay the upstream backstop.
  • Fail-open propagation. A malformed or missing map means no model cap at credential selection. Strict rejection (reported as modelConcurrency.<key>) happens only on write, so operator intent is never dropped silently.
  • Per-connection, per-process scope. Connections that share an upstream key do not coordinate.
  • Dashboard: a new Per-model concurrency caps textarea (one model=cap per line). It loads the stored map and writes it back, so maps set through the API survive unrelated saves. A malformed entry blocks the save with a localized error (providers.rateLimitOverridesModelConcurrencyInvalid).
  • Docs: a new Per-model concurrency caps subsection in docs/architecture/RESILIENCE_GUIDE.md.

Incidental fix: executionMaxWaitMs could not be persisted

updateProviderConnectionSchema accepts rateLimitOverrides.executionMaxWaitMs, but the DB-side allowlist in sanitizeRateLimitOverrides did not. A PATCH carrying it passed Zod and then threw Refusing to persist rateLimitOverrides with rejected keys: executionMaxWaitMs. This is the same class of bug as the #11251 maxWaitMs follow-up. The allowlist now matches the schema, and the dashboard form carries an API-set executionMaxWaitMs through saves, since the form has no field for it. Regression tests are in tests/unit/columns-validation.test.ts and the form/modal tests below.

Out of scope (deliberately excluded)

Hard-coded provider limits, limit discovery/scraping, new executors, global behavior changes, retries, and billing/quota/cooldown changes.

Validation

Rebased onto release/v3.8.52 @ 23a1148486. Run from a worktree on Linux, Node 24.21.0:

node --import tsx/esm --test tests/unit/account-concurrency-cap.test.ts tests/unit/chatcore-hierarchical-admission.test.ts \
  tests/unit/model-concurrency-gate.test.ts tests/unit/model-concurrency-input.test.ts tests/unit/provider-patch-model-concurrency.test.ts \
  tests/unit/columns-validation.test.ts tests/unit/rate-limit-execution-per-conn.test.ts tests/unit/provider-connection-rpd-override.test.ts \
  tests/unit/db-providers-split.test.ts
# 94 tests, 94 pass, 0 fail
npx vitest run --config vitest.config.ts tests/unit/dashboard/edit-connection-modal-*.test.tsx tests/unit/ui/edit-connection-modal-free-models.test.tsx \
  tests/unit/dashboard/providers/components/providerCardWarningIndicators.test.tsx
# 31 tests, all pass (incl. the new edit-connection-modal-model-concurrency.test.tsx)
npm run test:unit                 # 45,457 tests: 45,426 pass, 30 skipped, 1 fail (pre-existing, see below)
npm run lint                      # clean
npm run typecheck:core            # clean
npm run check:dashboard-typecheck # OK, within frozen baseline
npm run check:file-size           # OK
npm run check:db-rules            # OK
npm run check:docs-all            # PASS (incl. fabricated-docs)
npm run i18n:check-ui-coverage && npm run i18n:check-glossary   # PASS
BASE_REF=upstream/release/v3.8.52 npm run i18n:check-value-drift # PASS

npm run test:scoped resolves to the full suite because hub files (the locale catalogs) are touched, so the full unit suite was run as well. Its one failure, tests/unit/verified-connection-activation-11446.test.ts ("testSingleConnection activates a connection once a test actually passes", which times out after about 16 s with upstream_error), reproduces identically on a clean release/v3.8.52 @ 23a1148486 checkout. It is unrelated to this PR: the test stubs globalThis.fetch, which the validation probe appears to bypass. It is not yet listed in #15306.

Changed/added test files

  • tests/unit/model-concurrency-gate.test.ts (new)
  • tests/unit/model-concurrency-input.test.ts (new: parser and form-helper coverage)
  • tests/unit/provider-patch-model-concurrency.test.ts (new)
  • tests/unit/dashboard/edit-connection-modal-model-concurrency.test.tsx (new, Vitest)
  • tests/unit/account-concurrency-cap.test.ts (extended)
  • tests/unit/chatcore-hierarchical-admission.test.ts (extended)
  • tests/unit/columns-validation.test.ts (extended: executionMaxWaitMs regression)

File-size baseline: open-sse/handlers/chatCore.ts (+14) and src/sse/services/auth.ts (+3) keep only call-site wiring; the logic lives in non-frozen helpers. The justification is in config/quality/file-size-baseline.json under _rebaseline_2026_09_25_13700_per_model_concurrency, re-measured after the rebase.

Commits

The branch was rebuilt on release/v3.8.52. Its earlier release/v3.8.51 sync merges contained real conflict-resolution content (the helper extraction), which a plain rebase would have dropped. So the original series is squashed into one feature commit, with the follow-up on top:

  • feat(sse): opt-in per-model concurrency caps per connection
  • fix(dashboard): localize per-model cap save error and keep executionMaxWaitMs

@diegosouzapw

Copy link
Copy Markdown
Owner

Thanks for this — clean extension of the existing hierarchical admission gate
(global → provider → account) to a fourth, opt-in per-model layer, and it reuses
acquireMany/acquireConcurrencyGates instead of building a parallel mechanism, which is
exactly the right call here. Review checked out the branch, ran the full focused test suite
you listed (40/40 pass), eslint on every touched file under the repo's suppressions
baseline (clean), and tsc for both open-sse and the dashboard (clean) — no conflicts
against the current release/v3.8.51 tip either.

Two small non-blocking notes for a possible follow-up: (1) rateLimitOverrides.modelConcurrency
isn't yet reflected in docs/openapi.yaml/docs/reference/API_REFERENCE.md, worth adding so
generated API clients pick it up; (2) ProviderCredentials.modelConcurrency is declared in
both open-sse/types.d.ts and the local type in open-sse/executors/base.ts — pre-existing
duplication (same as maxConcurrent) that this PR just followed, not something to fix here.

⚠️ base-red inherited: #12732 — the hard failures tracked there are pre-existing on
release/v3.8.51 and don't touch any file this PR changes.

This looks merge-ready to us pending the GitHub Actions checks finishing green (only Mergify
and the semgrep scan had reported by the time of this review).

@diegosouzapw diegosouzapw added the deferred-v3.8.52 Grande demais / suspeito para o lote atual; precisa de sessão dedicada no ciclo v3.8.52 label Sep 25, 2026
@diegosouzapw diegosouzapw changed the title feat(sse): opt-in per-model concurrency caps per connection [defer] feat(sse): opt-in per-model concurrency caps per connection Sep 25, 2026
@diegosouzapw
diegosouzapw changed the base branch from release/v3.8.51 to release/v3.8.52 September 29, 2026 11:25
@diegosouzapw

Copy link
Copy Markdown
Owner

Re-homed to release/v3.8.52: v3.8.51 entered its release freeze, so the branch now belongs to the release captain and development continues on the next cycle. Nothing is wrong with this PR — it just needed a live base. No action needed from you; CI will re-run against the new base.

@diegosouzapw diegosouzapw removed the deferred-v3.8.52 Grande demais / suspeito para o lote atual; precisa de sessão dedicada no ciclo v3.8.52 label Oct 1, 2026
@diegosouzapw diegosouzapw changed the title [defer] feat(sse): opt-in per-model concurrency caps per connection feat(sse): opt-in per-model concurrency caps per connection Oct 1, 2026
Adds optional per-model concurrency caps on a provider connection, enforced
by the chatCore admission gate alongside the existing per-account cap.

Squashed from the original series (incl. merge-time helper extraction):
- feat(sse): add opt-in per-model concurrency caps per connection
- feat(dashboard): use i18n keys for per-model concurrency editor
- docs(changelog): add fragment for per-model concurrency caps
- extract per-model concurrency UI/gate helpers under file-size ceilings
- test(sse): align admission-gate source assertions with resolveModelSemaphore helper
- fix(i18n): mirror per-model concurrency keys into all locales
- chore(quality): rebaseline per-model concurrency file-size ceilings
- fix(shared): move model concurrency bounds to a server-free leaf
…axWaitMs

- parseModelConcurrencyInput now reports the offending entry instead of a
  hard-coded English sentence; the modal renders it through the new
  providers.rateLimitOverridesModelConcurrencyInvalid i18n key.
- sanitizeRateLimitOverrides accepts executionMaxWaitMs, matching the Zod
  schema; previously a PATCH carrying it passed validation and then failed
  with "Refusing to persist rateLimitOverrides with rejected keys".
- buildRateLimitOverridesFromForm carries over an API-set executionMaxWaitMs
  (the form has no field for it) so a dashboard save no longer wipes it.
- Tests: form-helper coverage in model-concurrency-input.test.ts and the
  allowlist regression in columns-validation.test.ts.
@Notaloop763
Notaloop763 force-pushed the feat/per-model-concurrency branch from 56bed30 to c1d61a9 Compare October 3, 2026 19:38

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants