feat(gateway): per-org concurrency + overload caps - #2668
Conversation
|
Note Reviews pausedIt looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the Use the following commands to manage reviews:
Use the checkboxes below for quick actions:
WalkthroughAdds gateway backpressure middleware that caps per-pod in-flight requests and returns HTTP 529 with Retry-After when overloaded. It also adds 529-aware error rendering, wires the middleware into the app, updates server startup behavior, and documents the overload flow. ChangesBackpressure Load-Shedding Flow
Estimated code review effort: 4 (Complex) | ~45 minutes Possibly related PRs
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: cc2fe4cc84
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
|
||
| // Shed excess load early (after logging, before auth/routing) so each pod | ||
| // fast-fails with a retryable 529 instead of piling up unbounded connections. | ||
| app.use("*", backpressureMiddleware); |
There was a problem hiding this comment.
When the in-flight cap is reached, this middleware returns the 529 without calling next(). Because it is registered before the CORS middleware, cross-origin browser clients hitting non-exempt /v1/... routes during overload receive no Access-Control-Allow-Origin, so fetch reports a CORS/network failure instead of exposing the retryable 529 response or Retry-After guidance. Move CORS ahead of this early-return middleware or add the CORS headers on the shed response.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Actionable comments posted: 1
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
apps/gateway/src/app.ts (1)
86-103:⚠️ Potential issue | 🟠 Major | ⚡ Quick winPlace backpressure after CORS so overloaded responses keep CORS headers.
With Line 86 before
cors(...), shed requests return 529 withoutAccess-Control-Allow-Originand related headers. Browser clients can fail to surface overload responses/retry hints correctly.Suggested reorder
app.use("*", tracingMiddleware); app.use("*", requestLifecycleMiddleware); app.use("*", honoRequestLogger); - -// Shed excess load early (after logging, before auth/routing) so each pod -// fast-fails with a retryable 529 instead of piling up unbounded connections. -app.use("*", backpressureMiddleware); app.use( "*", cors({ origin: "*", ... }), ); + +// Shed excess load early (after logging/CORS, before auth/routing). +app.use("*", backpressureMiddleware);🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@apps/gateway/src/app.ts` around lines 86 - 103, The backpressure middleware is registered before the CORS middleware so overload (529) responses miss CORS headers; move the backpressureMiddleware registration to after the cors(...) registration so both use app.use("*", ...) in that order, i.e., keep the existing cors({ ... }) call as-is and re-register backpressureMiddleware immediately after it (still on the "*" route) so shed responses include the Access-Control-Allow-* headers.
🧹 Nitpick comments (2)
apps/gateway/src/lib/backpressure.ts (1)
13-13: ⚡ Quick winAvoid import-time caching of
GATEWAY_MAX_INFLIGHT_REQUESTS(Line 13).This value is frozen at module load, which is why tests currently need module resets/re-imports to vary behavior. Parse on use (with positive-integer validation) instead of caching the parsed env value.
Suggested refactor
-const MAX_INFLIGHT = Number(process.env.GATEWAY_MAX_INFLIGHT_REQUESTS) || 1000; +function getMaxInflight(): number { + const parsed = Number(process.env.GATEWAY_MAX_INFLIGHT_REQUESTS); + return Number.isFinite(parsed) && parsed > 0 ? parsed : 1000; +} ... - if (inFlight >= MAX_INFLIGHT) { + const maxInflight = getMaxInflight(); + if (inFlight >= maxInflight) { c.header("Retry-After", "1"); return renderGatewayError(c, 529, "Gateway overloaded, please retry"); }As per coding guidelines,
**/*.{ts,tsx,js,jsx}: “Do not add explicit caching or memoization around process.env reads or parsed env-var values unless there is a measured hot-path need.”🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@apps/gateway/src/lib/backpressure.ts` at line 13, The module currently caches GATEWAY_MAX_INFLIGHT at import time via the MAX_INFLIGHT constant; change this to read and validate the env var at call time by replacing MAX_INFLIGHT with a small accessor function (e.g., getMaxInflight()) that reads process.env.GATEWAY_MAX_INFLIGHT, parses it with Number/parseInt, validates it is a positive integer, and returns the parsed value or the default 1000 on invalid/missing input; update every use of MAX_INFLIGHT to call getMaxInflight() so tests and runtime variations take effect and no import-time caching occurs.Source: Coding guidelines
apps/gateway/src/serve.ts (1)
28-29: ⚡ Quick winAvoid module-scope caching for the new
LISTEN_BACKLOGparse.Line 28 introduces a cached parsed env value at module scope. This code path is startup-only, so this can be read at use-site (or in
startServer) without long-lived caching.♻️ Guideline-aligned refactor
-const listenBacklog = Number(process.env.LISTEN_BACKLOG) || 1024; +const getListenBacklog = () => Number(process.env.LISTEN_BACKLOG) || 1024;async function startServer() { + const listenBacklog = getListenBacklog(); logger.info("Server starting", { port, backlog: listenBacklog });As per coding guidelines, "Do not add explicit caching or memoization around
process.envreads or parsed env-var values unless there is a measured hot-path need."🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@apps/gateway/src/serve.ts` around lines 28 - 29, Remove the module-scope cached parse of LISTEN_BACKLOG (the listenBacklog const) and instead read and parse process.env.LISTEN_BACKLOG at use-site (for example inside startServer or wherever the server is started); replace references to the module-level listenBacklog with a local Number(process.env.LISTEN_BACKLOG) || 1024 computed when starting the server so there is no long-lived env-value caching.Source: Coding guidelines
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@apps/gateway/src/serve.ts`:
- Around line 46-50: startServer currently calls
createAdaptorServer(...).listen(...) and returns immediately, so bind failures
become uncaught; change startServer to await the listen completion by wrapping
server.listen(...) in a Promise that resolves on the server's 'listening' event
and rejects on the server's 'error' event (attach once listeners and clean them
up after resolution), and keep the existing logger.info call in the 'listening'
handler; ensure an 'error' listener is present (or the Promise rejection will
propagate) so bind failures are surfaced to the caller instead of as an uncaught
exception.
---
Outside diff comments:
In `@apps/gateway/src/app.ts`:
- Around line 86-103: The backpressure middleware is registered before the CORS
middleware so overload (529) responses miss CORS headers; move the
backpressureMiddleware registration to after the cors(...) registration so both
use app.use("*", ...) in that order, i.e., keep the existing cors({ ... }) call
as-is and re-register backpressureMiddleware immediately after it (still on the
"*" route) so shed responses include the Access-Control-Allow-* headers.
---
Nitpick comments:
In `@apps/gateway/src/lib/backpressure.ts`:
- Line 13: The module currently caches GATEWAY_MAX_INFLIGHT at import time via
the MAX_INFLIGHT constant; change this to read and validate the env var at call
time by replacing MAX_INFLIGHT with a small accessor function (e.g.,
getMaxInflight()) that reads process.env.GATEWAY_MAX_INFLIGHT, parses it with
Number/parseInt, validates it is a positive integer, and returns the parsed
value or the default 1000 on invalid/missing input; update every use of
MAX_INFLIGHT to call getMaxInflight() so tests and runtime variations take
effect and no import-time caching occurs.
In `@apps/gateway/src/serve.ts`:
- Around line 28-29: Remove the module-scope cached parse of LISTEN_BACKLOG (the
listenBacklog const) and instead read and parse process.env.LISTEN_BACKLOG at
use-site (for example inside startServer or wherever the server is started);
replace references to the module-level listenBacklog with a local
Number(process.env.LISTEN_BACKLOG) || 1024 computed when starting the server so
there is no long-lived env-value caching.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository UI
Review profile: CHILL
Plan: Pro
Run ID: aab7ab30-cd66-4a57-a072-0510cc040bf5
📒 Files selected for processing (8)
apps/gateway/src/app.tsapps/gateway/src/lib/backpressure.spec.tsapps/gateway/src/lib/backpressure.tsapps/gateway/src/lib/error-response.spec.tsapps/gateway/src/lib/error-response.tsapps/gateway/src/serve.tspackages/instrumentation/src/index.tspackages/instrumentation/src/metrics.ts
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 242479ed82
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| path.startsWith("/mcp") || | ||
| path.startsWith("/oauth") |
There was a problem hiding this comment.
Exempt the OAuth discovery routes too
This predicate exempts /oauth*, but registerMcpOAuthRoutes also exposes OAuth metadata at /.well-known/oauth-authorization-server and /.well-known/oauth-authorization-server/mcp. When the in-flight cap is reached, MCP/OAuth clients that start by fetching RFC 8414 discovery will get a 529 before they can reach the exempt /oauth/* routes, so the advertised OAuth flow no longer stays available while the pod is shedding load.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: db6cb45d32
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| if (inFlight >= getMaxInflight()) { | ||
| c.header("Retry-After", "1"); | ||
| return renderGatewayError(c, 529, "Gateway overloaded, please retry"); |
There was a problem hiding this comment.
Expose Retry-After to browser callers
For cross-origin browser clients hitting non-exempt /v1/... routes during overload, this newly set Retry-After header is still hidden because the CORS config only exposes Content-Length and mcp-session-id; the client can see the 529 after moving backpressure behind CORS, but cannot honor the documented retry delay and may retry too aggressively. Add Retry-After to the CORS exposeHeaders list when returning this overload response.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: e74cecbcf3
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| return ( | ||
| path === "/" || | ||
| path === "/metrics" || | ||
| path.startsWith("/mcp") || |
There was a problem hiding this comment.
Count long-lived MCP streams in the cap
When clients open /mcp with Accept: text/event-stream, mcpHandler returns a keep-alive ReadableStream and stores the session until disconnect; this exemption means those long-lived MCP connections never increment the new in-flight counter and can still pile up unbounded, bypassing the overload protection this middleware is meant to provide. Consider exempting only cheap MCP preflight/discovery-style requests or applying a separate cap for MCP streams.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Pull request overview
This PR improves gateway overload resilience by increasing the Node listen backlog, changing process-level error handling to avoid restart-cycling under load, and adding a backpressure middleware that sheds excess concurrent requests with a retryable HTTP 529 (plus metrics, tests, and docs) so pods degrade gracefully instead of timing out behind the LB.
Changes:
- Increased accept backlog via
createAdaptorServer()+server.listen({ backlog })and added startup bind validation/logging. - Added per-pod in-flight request shedding middleware (HTTP 529 +
Retry-After) with Prometheus metrics for tuning/alerting. - Refactored shared error rendering for provider-compatible 529 envelopes and documented the new overload behavior.
Reviewed changes
Copilot reviewed 10 out of 10 changed files in this pull request and generated 2 comments.
Show a summary per file
| File | Description |
|---|---|
| packages/instrumentation/src/metrics.ts | Adds gateway_inflight_requests gauge and gateway_requests_shed_total counter for overload tuning/visibility. |
| packages/instrumentation/src/index.ts | Exposes the new gateway-level overload metrics from the instrumentation package. |
| apps/gateway/src/serve.ts | Raises listen backlog and changes process-level error handling to log-and-continue; improves bind failure behavior. |
| apps/gateway/src/lib/error-response.ts | Adds 529 mapping and extracts renderGatewayError for reuse by middleware/handlers. |
| apps/gateway/src/lib/error-response.spec.ts | Adds unit coverage for 529 mapping across OpenAI + Anthropic error envelopes. |
| apps/gateway/src/lib/backpressure.ts | Introduces in-flight concurrency cap middleware that sheds with 529 and updates metrics. |
| apps/gateway/src/lib/backpressure.spec.ts | Adds unit tests for shedding behavior, slot release, and exempt paths. |
| apps/gateway/src/app.ts | Wires backpressure middleware early in the pipeline (after CORS, before auth/routing). |
| apps/docs/content/resources/rate-limits.mdx | Documents overload 529 behavior and retry guidance alongside rate-limit guidance. |
| apps/docs/content/resources/error-handling.mdx | Adds 529 to the public status→OpenAI error-type mapping table. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| if (outgoing) { | ||
| outgoing.once("close", decrement); | ||
| return await next(); | ||
| } |
| // the GKE L7 LB surfaces as "connection timeout". Raise it so bursts queue | ||
| // instead of being dropped. The kernel caps the effective value at | ||
| // net.core.somaxconn (4096 on modern COS nodes), so 1024 is honored. | ||
| const listenBacklog = Number(process.env.LISTEN_BACKLOG) || 1024; |
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
1 similar comment
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
Three independent, env-tunable resilience changes so each pod sheds
load instead of falling over when a slow upstream piles up connections:
1. Listen backlog — raise Node's default accept backlog (511) to 1024
via createAdaptorServer so connection bursts queue instead of being
dropped (surfaced by the LB as "connection timeout"). Knob:
LISTEN_BACKLOG.
2. Process-error handling — log + continue on uncaughtException /
unhandledRejection instead of self-terminating, which caused
restart-cycling under load. SIGTERM/SIGINT still shut down gracefully.
3. In-flight backpressure — new middleware caps concurrent requests per
pod and fast-fails excess with a retryable HTTP 529 + Retry-After.
Readiness/metrics/MCP/OAuth paths stay exempt so the pod keeps
passing readiness while shedding. Tracked via a new
gateway_inflight_requests gauge. Knob: GATEWAY_MAX_INFLIGHT_REQUESTS.
renderGatewayError moves to error-response.ts so the middleware can
reuse it, and 529 maps to OpenAI {type/code: overloaded} and Anthropic
{type: overloaded_error}.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Register backpressure after CORS so shed 529s carry Access-Control-* headers; browser clients can now see the response and Retry-After hint instead of a CORS/network failure. - Await the listen bind in startServer (resolve on 'listening', reject on 'error') so bind failures (e.g. EADDRINUSE) exit(1) instead of being swallowed by the new log-and-continue uncaughtException handler. - Read GATEWAY_MAX_INFLIGHT_REQUESTS and LISTEN_BACKLOG at use-site instead of caching at module load, per the repo env-read guideline. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add a "Gateway Overload (529)" section explaining the transient, retryable load-shedding response (Retry-After, OpenAI/Anthropic envelopes) and how it differs from a 429. Add the 529 -> overloaded row to the error-handling status-code reference. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add a gateway_requests_shed_total counter incremented when the backpressure middleware rejects a request with HTTP 529. Enables a direct sheds/sec signal for dashboards and alerting instead of inferring it from the in-flight gauge riding the cap. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Scope the pod-wide 529 backpressure cap to inference POSTs only (cheap routes are never counted or shed) and skip internal app.request() hops, which the previous revision double-counted. Add a fleet-wide per-org in-flight budget in the org rate-limit middleware: slots live in a Redis zset for the response's full lifetime, stale slots from crashed pods are reaped, and over-limit requests get a retryable 429. Enterprise orgs skip RPM limits but get an elevated concurrency ceiling instead of an exemption. Slot ops use explicit pipelines because bare auto-pipelined commands issued from the response-close context wedge ioredis. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 4a77de6b2d
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| const checkResults = await redisClient | ||
| .pipeline() | ||
| .zremrangebyscore(key, "-inf", staleBefore) | ||
| .zcard(key) | ||
| .exec(); |
There was a problem hiding this comment.
Make the concurrency check-and-add atomic
During a same-organization burst across gateway pods, multiple acquisitions can all complete this pipeline, observe currentCount < limit, and then independently execute the separate ZADD at line 377. The number of admitted requests can therefore exceed the configured ceiling by the size of the concurrent burst, defeating the fleet-wide protection precisely under overload; perform the reap, count check, and insertion in one atomic Redis script or transaction.
Useful? React with 👍 / 👎.
| "spend_cap_daily", | ||
| "spend_cap_monthly", | ||
| "topup_velocity", | ||
| "concurrency", |
There was a problem hiding this comment.
Surface concurrency hits in the admin dashboard
Once concurrency rejections are flushed, this API accepts and returns the new type, but the summary has no concurrencyHits bucket while totalHits includes it, and the unchanged ee/admin/src/app/limit-hits/page.tsx filter and columns only cover RPM, spend caps, and top-ups. As a result, totals no longer reconcile and admins cannot select concurrency hits from the overview; add the corresponding summary field and update the admin type lists, labels, and column.
Useful? React with 👍 / 👎.
Chat plan concurrency drops to 10 and dev plan to 50. Regular (PAYG) org ceilings now scale with the trust tier (100/200/400/1000/2000 for tiers 0-4, env-overridable per tier), resolved lazily only once an org reaches the tier-0 base — mirroring the RPM multiplier. Concurrency denials get full admin limit-hit parity: a concurrencyHits column in the admin summary API, a filter option and column on the Limit Hits page, and a labeled badge on the per-org rate-limits breakdown. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 624cb761b4
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| if (outgoing) { | ||
| outgoing.once("close", runCleanup); | ||
| return await next(); |
There was a problem hiding this comment.
Release slots when the response already closed
If a client disconnects while the org/token/RPM/Redis lookups are still awaiting, the Node adapter can emit close before acquireOrgInflightSlot returns. This helper then registers a listener for an event that has already occurred, so the newly added Redis slot is retained until the 30-minute stale timeout; a burst of timed-out requests can therefore exhaust a chat org's 10-slot budget and falsely reject subsequent traffic. Check outgoing.destroyed/completion before registering the listener and run cleanup immediately when it is already closed.
Useful? React with 👍 / 👎.
| if (c.req.method === "POST" && INFLIGHT_LIMITED_KEYS.has(config.key)) { | ||
| const baseLimit = getOrgInflightLimit(planClass, isEnterprise); | ||
| const acquisition = await acquireOrgInflightSlot( |
There was a problem hiding this comment.
Expose the enterprise concurrency ceiling in the Limits UI
Enterprise requests now receive the new per-org concurrency check, but the customer Limits page still says enterprise organizations have no per-organization rate limits and that throughput is bounded only by credits/upstream providers (apps/ui/src/app/dashboard/[orgId]/org/limits/_components/limits-client.tsx:103-107), while the limits API does not return inflightLimit. Enterprise users who encounter the new 429 therefore see a dashboard that explicitly denies the enforced 2,000-request ceiling; return and render the concurrency limit, including in the enterprise branch.
Useful? React with 👍 / 👎.
Problem
Under a traffic spike or a slow upstream, gateway pods accumulated unbounded in-flight connections until they fell over. The first revision of this PR added a global per-pod cap applied to every route pre-auth — but only inference requests are long-lived (streams can hold connections for minutes), so cheap routes were being punished for inference pile-ups, and a single organization's runaway concurrency could starve every other tenant on the pod.
Approach
Two complementary layers, both scoped to inference
POSTs (chat completions, messages, responses, embeddings, moderations, rerank, OCR, images, speech, transcriptions, videos, AI SDK) and both skipping trusted internalapp.request()re-dispatch hops so forwarded requests are not double-counted:lib/backpressure.ts, in-memory per-pod cap (GATEWAY_MAX_INFLIGHT_REQUESTS, default 1000). Over the cap, inference requests are shed instantly with529+Retry-After: 1; health probes,/v1/models, MCP/OAuth and docs are never counted, so the pod keeps passing readiness while shedding.ServerResponsecloseevent) and stale slots from crashed pods are reaped afterGATEWAY_ORG_INFLIGHT_STALE_SECONDS(default 1800). Over-limit requests get a retryable429 rate_limit_error+Retry-After: 1. Fails open on Redis errors.Limits
Regular (PAYG) org ceilings scale with the same trust tier that drives the RPM multiplier and spend caps, resolved lazily only once an org reaches the tier-0 base (mirroring the RPM multiplier's lazy spend lookup). Dev/chat plans are flat, and enterprise orgs are not exempt — they get an elevated ceiling instead, since unbounded single-tenant concurrency exhausts shared capacity regardless of plan.
GATEWAY_SPEND_TIER_<N>_INFLIGHT_LIMIT)All env-overridable;
0disables a class.Admin insights
Concurrency denials get the same limit-hit tracking as the RPM/spend/top-up limits: recorded per endpoint via
recordLimitHit(newconcurrencylimit type), flushed by the existing worker job, aggregated asconcurrencyHitsin the admin summary API, filterable and shown as a column on the admin Limit Hits page, and labeled in the per-org rate-limits breakdown.Per-org rate-limits breakdown (light + dark)
Supporting changes
serve.ts: listen backlog raised toLISTEN_BACKLOG(default 1024), bind awaited soEADDRINUSEexits instead of being swallowed, anduncaughtException/unhandledRejectionnow log-and-continue instead of restart-cycling under load.gateway_inflight_requestsnow counts inference requests only (purer HPA signal, name unchanged);gateway_requests_shed_totalgains ascopelabel (pod= 529,org= 429).Non-obvious constraint
The per-org slot ops use explicit
redisClient.pipeline()calls, never bare commands: a bare auto-pipelined command issued from the response-close/finallycontext can wedge ioredis's auto-pipelining (the queued command's flush is never scheduled), silently hanging every later command on the shared client. This surfaced as deterministic gateway integration-test hangs and is why the release path is pipeline-based.Verification
api.spec.ts, onboarding) re-run green.429 rate_limit_errorwithRetry-After: 1;/v1/modelsstays 200 during saturation; the slot zset is empty after the stream closes;gateway_requests_shed_total{scope="org"}increments and the inflight gauge returns to 0.pnpm format+pnpm buildpass; thelimitTypetext-enum addition produces no migration (verified with drizzle-kit).🤖 Generated with Claude Code