Skip to content

feat(gateway): per-org concurrency + overload caps - #2668

Merged
steebchen merged 6 commits into
mainfrom
colombo-v2
Aug 21, 2026
Merged

steebchen merged 6 commits into
mainfrom
colombo-v2

Conversation

@steebchen

@steebchen steebchen commented Jun 13, 2026 •

Copy link
Copy Markdown
Member

Problem

Under a traffic spike or a slow upstream, gateway pods accumulated unbounded in-flight connections until they fell over. The first revision of this PR added a global per-pod cap applied to every route pre-auth — but only inference requests are long-lived (streams can hold connections for minutes), so cheap routes were being punished for inference pile-ups, and a single organization's runaway concurrency could starve every other tenant on the pod.

Approach

Two complementary layers, both scoped to inference POSTs (chat completions, messages, responses, embeddings, moderations, rerank, OCR, images, speech, transcriptions, videos, AI SDK) and both skipping trusted internal app.request() re-dispatch hops so forwarded requests are not double-counted:

  1. Pod-wide backpressure (529) — lib/backpressure.ts, in-memory per-pod cap (GATEWAY_MAX_INFLIGHT_REQUESTS, default 1000). Over the cap, inference requests are shed instantly with 529 + Retry-After: 1; health probes, /v1/models, MCP/OAuth and docs are never counted, so the pod keeps passing readiness while shedding.
  2. Per-org fleet-wide concurrency limit (429) — added to the existing org rate-limit middleware (feat: per-org rate limits, spend & top-up caps #2646), reusing its token → org resolution. One Redis zset budget per org across all inference endpoints; a slot is held for the response's full lifetime (streaming included, released on the raw ServerResponse close event) and stale slots from crashed pods are reaped after GATEWAY_ORG_INFLIGHT_STALE_SECONDS (default 1800). Over-limit requests get a retryable 429 rate_limit_error + Retry-After: 1. Fails open on Redis errors.

Limits

Regular (PAYG) org ceilings scale with the same trust tier that drives the RPM multiplier and spend caps, resolved lazily only once an org reaches the tier-0 base (mirroring the RPM multiplier's lazy spend lookup). Dev/chat plans are flat, and enterprise orgs are not exempt — they get an elevated ceiling instead, since unbounded single-tenant concurrency exhausts shared capacity regardless of plan.

Plan Concurrent requests
Regular tier 0–4 100 / 200 / 400 / 1,000 / 2,000 (GATEWAY_SPEND_TIER_<N>_INFLIGHT_LIMIT)
Dev plan 50
Chat plan 10
Enterprise 2,000

All env-overridable; 0 disables a class.

Admin insights

Concurrency denials get the same limit-hit tracking as the RPM/spend/top-up limits: recorded per endpoint via recordLimitHit (new concurrency limit type), flushed by the existing worker job, aggregated as concurrencyHits in the admin summary API, filterable and shown as a column on the admin Limit Hits page, and labeled in the per-org rate-limits breakdown.

Admin limit-hits page (light) with the new Concurrency filter and column

Admin limit-hits page (dark)

Per-org rate-limits breakdown (light + dark) Per-org anti-abuse limit hits with Concurrency rows (light) Per-org anti-abuse limit hits with Concurrency rows (dark)

Supporting changes

  • serve.ts: listen backlog raised to LISTEN_BACKLOG (default 1024), bind awaited so EADDRINUSE exits instead of being swallowed, and uncaughtException/unhandledRejection now log-and-continue instead of restart-cycling under load.
  • Metrics: gateway_inflight_requests now counts inference requests only (purer HPA signal, name unchanged); gateway_requests_shed_total gains a scope label (pod = 529, org = 429).
  • Docs: rate-limits page gains a "Concurrent Request Limits" section (tier ladder included) and the 529 overload section; the trust-tier table now shows the concurrency column.

Non-obvious constraint

The per-org slot ops use explicit redisClient.pipeline() calls, never bare commands: a bare auto-pipelined command issued from the response-close/finally context can wedge ioredis's auto-pipelining (the queued command's flush is never scheduled), silently hanging every later command on the shared client. This surfaced as deterministic gateway integration-test hangs and is why the release path is pipeline-based.

Verification

  • Full unit suite green against an isolated stack (321 files, 5232 tests); after the tier rework, all scoped specs (lib, middleware, backpressure, spend-limit, admin-limit-hits) plus the gateway integration suites (api.spec.ts, onboarding) re-run green.
  • Live check on a local gateway with a concurrency limit of 1: concurrent second request gets the 429 rate_limit_error with Retry-After: 1; /v1/models stays 200 during saturation; the slot zset is empty after the stream closes; gateway_requests_shed_total{scope="org"} increments and the inflight gauge returns to 0.
  • Admin screenshots above are from a locally seeded stack.
  • pnpm format + pnpm build pass; the limitType text-enum addition produces no migration (verified with drizzle-kit).

🤖 Generated with Claude Code

@coderabbitai

coderabbitai Bot commented Jun 13, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

Walkthrough

Adds gateway backpressure middleware that caps per-pod in-flight requests and returns HTTP 529 with Retry-After when overloaded. It also adds 529-aware error rendering, wires the middleware into the app, updates server startup behavior, and documents the overload flow.

Changes

Backpressure Load-Shedding Flow

Layer / File(s) Summary
Backpressure Middleware & Instrumentation
apps/gateway/src/lib/backpressure.ts, apps/gateway/src/lib/backpressure.spec.ts, packages/instrumentation/src/metrics.ts, packages/instrumentation/src/index.ts
Backpressure middleware caps concurrent in-flight requests using GATEWAY_MAX_INFLIGHT_REQUESTS, exempts health/metrics/MCP/OAuth routes, returns HTTP 529 with Retry-After: 1 when overloaded, and decrements via response close or a finally fallback. A new gatewayInflightRequests gauge and gatewayRequestsShedTotal counter are added and re-exported, with tests covering shedding, metrics, and exemptions.
Overloaded Error Response Support
apps/gateway/src/lib/error-response.ts, apps/gateway/src/lib/error-response.spec.ts
Maps HTTP 529 to OpenAI "overloaded" meta and adds renderGatewayError that emits Anthropic or OpenAI JSON envelopes based on request path; tests verify 529 metadata for both providers.
Gateway App Integration
apps/gateway/src/app.ts
Wires backpressureMiddleware early in the middleware chain and replaces the in-file renderGatewayError helper with the shared implementation; app.onError uses the imported renderer.
Server Startup & Error Handling
apps/gateway/src/serve.ts
Uses createAdaptorServer, adds configurable listenBacklog (default 1024) and waiting for the listening event during startup, updates startup logs, and changes uncaught/unhandled rejection handlers to log and continue running.
Gateway Overload Docs
apps/docs/content/resources/error-handling.mdx, apps/docs/content/resources/rate-limits.mdx
Adds HTTP 529 to the error-handling status table and expands the rate-limits docs with a Gateway Overload section, example payloads, Anthropic notes, and updated retry guidance.

Estimated code review effort: 4 (Complex) | ~45 minutes

Possibly related PRs

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 57.14% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly identifies the gateway concurrency and overload-cap changes, which are central aspects of the pull request.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch colombo-v2

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: cc2fe4cc84

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread apps/gateway/src/app.ts

// Shed excess load early (after logging, before auth/routing) so each pod
// fast-fails with a retryable 529 instead of piling up unbounded connections.
app.use("*", backpressureMiddleware);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Run backpressure after CORS

When the in-flight cap is reached, this middleware returns the 529 without calling next(). Because it is registered before the CORS middleware, cross-origin browser clients hitting non-exempt /v1/... routes during overload receive no Access-Control-Allow-Origin, so fetch reports a CORS/network failure instead of exposing the retryable 529 response or Retry-After guidance. Move CORS ahead of this early-return middleware or add the CORS headers on the shed response.

Useful? React with 👍 / 👎.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
apps/gateway/src/app.ts (1)

86-103: ⚠️ Potential issue | 🟠 Major | ⚡ Quick win

Place backpressure after CORS so overloaded responses keep CORS headers.

With Line 86 before cors(...), shed requests return 529 without Access-Control-Allow-Origin and related headers. Browser clients can fail to surface overload responses/retry hints correctly.

Suggested reorder
 app.use("*", tracingMiddleware);
 app.use("*", requestLifecycleMiddleware);
 app.use("*", honoRequestLogger);
-
-// Shed excess load early (after logging, before auth/routing) so each pod
-// fast-fails with a retryable 529 instead of piling up unbounded connections.
-app.use("*", backpressureMiddleware);
 
 app.use(
 	"*",
 	cors({
 		origin: "*",
 		...
 	}),
 );
+
+// Shed excess load early (after logging/CORS, before auth/routing).
+app.use("*", backpressureMiddleware);
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@apps/gateway/src/app.ts` around lines 86 - 103, The backpressure middleware
is registered before the CORS middleware so overload (529) responses miss CORS
headers; move the backpressureMiddleware registration to after the cors(...)
registration so both use app.use("*", ...) in that order, i.e., keep the
existing cors({ ... }) call as-is and re-register backpressureMiddleware
immediately after it (still on the "*" route) so shed responses include the
Access-Control-Allow-* headers.
🧹 Nitpick comments (2)
apps/gateway/src/lib/backpressure.ts (1)

13-13: ⚡ Quick win

Avoid import-time caching of GATEWAY_MAX_INFLIGHT_REQUESTS (Line 13).

This value is frozen at module load, which is why tests currently need module resets/re-imports to vary behavior. Parse on use (with positive-integer validation) instead of caching the parsed env value.

Suggested refactor
-const MAX_INFLIGHT = Number(process.env.GATEWAY_MAX_INFLIGHT_REQUESTS) || 1000;
+function getMaxInflight(): number {
+	const parsed = Number(process.env.GATEWAY_MAX_INFLIGHT_REQUESTS);
+	return Number.isFinite(parsed) && parsed > 0 ? parsed : 1000;
+}
...
-	if (inFlight >= MAX_INFLIGHT) {
+	const maxInflight = getMaxInflight();
+	if (inFlight >= maxInflight) {
 		c.header("Retry-After", "1");
 		return renderGatewayError(c, 529, "Gateway overloaded, please retry");
 	}

As per coding guidelines, **/*.{ts,tsx,js,jsx}: “Do not add explicit caching or memoization around process.env reads or parsed env-var values unless there is a measured hot-path need.”

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@apps/gateway/src/lib/backpressure.ts` at line 13, The module currently caches
GATEWAY_MAX_INFLIGHT at import time via the MAX_INFLIGHT constant; change this
to read and validate the env var at call time by replacing MAX_INFLIGHT with a
small accessor function (e.g., getMaxInflight()) that reads
process.env.GATEWAY_MAX_INFLIGHT, parses it with Number/parseInt, validates it
is a positive integer, and returns the parsed value or the default 1000 on
invalid/missing input; update every use of MAX_INFLIGHT to call getMaxInflight()
so tests and runtime variations take effect and no import-time caching occurs.

Source: Coding guidelines

apps/gateway/src/serve.ts (1)

28-29: ⚡ Quick win

Avoid module-scope caching for the new LISTEN_BACKLOG parse.

Line 28 introduces a cached parsed env value at module scope. This code path is startup-only, so this can be read at use-site (or in startServer) without long-lived caching.

♻️ Guideline-aligned refactor
-const listenBacklog = Number(process.env.LISTEN_BACKLOG) || 1024;
+const getListenBacklog = () => Number(process.env.LISTEN_BACKLOG) || 1024;
 async function startServer() {
+	const listenBacklog = getListenBacklog();
 	logger.info("Server starting", { port, backlog: listenBacklog });

As per coding guidelines, "Do not add explicit caching or memoization around process.env reads or parsed env-var values unless there is a measured hot-path need."

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@apps/gateway/src/serve.ts` around lines 28 - 29, Remove the module-scope
cached parse of LISTEN_BACKLOG (the listenBacklog const) and instead read and
parse process.env.LISTEN_BACKLOG at use-site (for example inside startServer or
wherever the server is started); replace references to the module-level
listenBacklog with a local Number(process.env.LISTEN_BACKLOG) || 1024 computed
when starting the server so there is no long-lived env-value caching.

Source: Coding guidelines

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@apps/gateway/src/serve.ts`:
- Around line 46-50: startServer currently calls
createAdaptorServer(...).listen(...) and returns immediately, so bind failures
become uncaught; change startServer to await the listen completion by wrapping
server.listen(...) in a Promise that resolves on the server's 'listening' event
and rejects on the server's 'error' event (attach once listeners and clean them
up after resolution), and keep the existing logger.info call in the 'listening'
handler; ensure an 'error' listener is present (or the Promise rejection will
propagate) so bind failures are surfaced to the caller instead of as an uncaught
exception.

---

Outside diff comments:
In `@apps/gateway/src/app.ts`:
- Around line 86-103: The backpressure middleware is registered before the CORS
middleware so overload (529) responses miss CORS headers; move the
backpressureMiddleware registration to after the cors(...) registration so both
use app.use("*", ...) in that order, i.e., keep the existing cors({ ... }) call
as-is and re-register backpressureMiddleware immediately after it (still on the
"*" route) so shed responses include the Access-Control-Allow-* headers.

---

Nitpick comments:
In `@apps/gateway/src/lib/backpressure.ts`:
- Line 13: The module currently caches GATEWAY_MAX_INFLIGHT at import time via
the MAX_INFLIGHT constant; change this to read and validate the env var at call
time by replacing MAX_INFLIGHT with a small accessor function (e.g.,
getMaxInflight()) that reads process.env.GATEWAY_MAX_INFLIGHT, parses it with
Number/parseInt, validates it is a positive integer, and returns the parsed
value or the default 1000 on invalid/missing input; update every use of
MAX_INFLIGHT to call getMaxInflight() so tests and runtime variations take
effect and no import-time caching occurs.

In `@apps/gateway/src/serve.ts`:
- Around line 28-29: Remove the module-scope cached parse of LISTEN_BACKLOG (the
listenBacklog const) and instead read and parse process.env.LISTEN_BACKLOG at
use-site (for example inside startServer or wherever the server is started);
replace references to the module-level listenBacklog with a local
Number(process.env.LISTEN_BACKLOG) || 1024 computed when starting the server so
there is no long-lived env-value caching.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro

Run ID: aab7ab30-cd66-4a57-a072-0510cc040bf5

📥 Commits

Reviewing files that changed from the base of the PR and between 0bd7853 and cc2fe4c.

📒 Files selected for processing (8)
  • apps/gateway/src/app.ts
  • apps/gateway/src/lib/backpressure.spec.ts
  • apps/gateway/src/lib/backpressure.ts
  • apps/gateway/src/lib/error-response.spec.ts
  • apps/gateway/src/lib/error-response.ts
  • apps/gateway/src/serve.ts
  • packages/instrumentation/src/index.ts
  • packages/instrumentation/src/metrics.ts

Comment thread apps/gateway/src/serve.ts

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 242479ed82

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread apps/gateway/src/lib/backpressure.ts Outdated
Comment on lines +28 to +29
path.startsWith("/mcp") ||
path.startsWith("/oauth")

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Exempt the OAuth discovery routes too

This predicate exempts /oauth*, but registerMcpOAuthRoutes also exposes OAuth metadata at /.well-known/oauth-authorization-server and /.well-known/oauth-authorization-server/mcp. When the in-flight cap is reached, MCP/OAuth clients that start by fetching RFC 8414 discovery will get a 529 before they can reach the exempt /oauth/* routes, so the advertised OAuth flow no longer stays available while the pod is shedding load.

Useful? React with 👍 / 👎.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: db6cb45d32

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +41 to +43
if (inFlight >= getMaxInflight()) {
c.header("Retry-After", "1");
return renderGatewayError(c, 529, "Gateway overloaded, please retry");

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Expose Retry-After to browser callers

For cross-origin browser clients hitting non-exempt /v1/... routes during overload, this newly set Retry-After header is still hidden because the CORS config only exposes Content-Length and mcp-session-id; the client can see the 529 after moving backpressure behind CORS, but cannot honor the documented retry delay and may retry too aggressively. Add Retry-After to the CORS exposeHeaders list when returning this overload response.

Useful? React with 👍 / 👎.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: e74cecbcf3

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread apps/gateway/src/lib/backpressure.ts Outdated
return (
path === "/" ||
path === "/metrics" ||
path.startsWith("/mcp") ||

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Count long-lived MCP streams in the cap

When clients open /mcp with Accept: text/event-stream, mcpHandler returns a keep-alive ReadableStream and stores the session until disconnect; this exemption means those long-lived MCP connections never increment the new in-flight counter and can still pile up unbounded, bypassing the overload protection this middleware is meant to provide. Consider exempting only cheap MCP preflight/discovery-style requests or applying a separate cap for MCP streams.

Useful? React with 👍 / 👎.

Copilot AI review requested due to automatic review settings July 3, 2026 15:16

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR improves gateway overload resilience by increasing the Node listen backlog, changing process-level error handling to avoid restart-cycling under load, and adding a backpressure middleware that sheds excess concurrent requests with a retryable HTTP 529 (plus metrics, tests, and docs) so pods degrade gracefully instead of timing out behind the LB.

Changes:

  • Increased accept backlog via createAdaptorServer() + server.listen({ backlog }) and added startup bind validation/logging.
  • Added per-pod in-flight request shedding middleware (HTTP 529 + Retry-After) with Prometheus metrics for tuning/alerting.
  • Refactored shared error rendering for provider-compatible 529 envelopes and documented the new overload behavior.

Reviewed changes

Copilot reviewed 10 out of 10 changed files in this pull request and generated 2 comments.

Show a summary per file
File Description
packages/instrumentation/src/metrics.ts Adds gateway_inflight_requests gauge and gateway_requests_shed_total counter for overload tuning/visibility.
packages/instrumentation/src/index.ts Exposes the new gateway-level overload metrics from the instrumentation package.
apps/gateway/src/serve.ts Raises listen backlog and changes process-level error handling to log-and-continue; improves bind failure behavior.
apps/gateway/src/lib/error-response.ts Adds 529 mapping and extracts renderGatewayError for reuse by middleware/handlers.
apps/gateway/src/lib/error-response.spec.ts Adds unit coverage for 529 mapping across OpenAI + Anthropic error envelopes.
apps/gateway/src/lib/backpressure.ts Introduces in-flight concurrency cap middleware that sheds with 529 and updates metrics.
apps/gateway/src/lib/backpressure.spec.ts Adds unit tests for shedding behavior, slot release, and exempt paths.
apps/gateway/src/app.ts Wires backpressure middleware early in the pipeline (after CORS, before auth/routing).
apps/docs/content/resources/rate-limits.mdx Documents overload 529 behavior and retry guidance alongside rate-limit guidance.
apps/docs/content/resources/error-handling.mdx Adds 529 to the public status→OpenAI error-type mapping table.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread apps/gateway/src/lib/backpressure.ts Outdated
Comment on lines +70 to +73
if (outgoing) {
outgoing.once("close", decrement);
return await next();
}
Comment thread apps/gateway/src/serve.ts
// the GKE L7 LB surfaces as "connection timeout". Raise it so bursts queue
// instead of being dropped. The kernel caps the effective value at
// net.core.somaxconn (4096 on modern COS nodes), so 1024 is honored.
const listenBacklog = Number(process.env.LISTEN_BACKLOG) || 1024;
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

1 similar comment
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

steebchen and others added 5 commits August 20, 2026 00:20
Three independent, env-tunable resilience changes so each pod sheds
load instead of falling over when a slow upstream piles up connections:

1. Listen backlog — raise Node's default accept backlog (511) to 1024
   via createAdaptorServer so connection bursts queue instead of being
   dropped (surfaced by the LB as "connection timeout"). Knob:
   LISTEN_BACKLOG.
2. Process-error handling — log + continue on uncaughtException /
   unhandledRejection instead of self-terminating, which caused
   restart-cycling under load. SIGTERM/SIGINT still shut down gracefully.
3. In-flight backpressure — new middleware caps concurrent requests per
   pod and fast-fails excess with a retryable HTTP 529 + Retry-After.
   Readiness/metrics/MCP/OAuth paths stay exempt so the pod keeps
   passing readiness while shedding. Tracked via a new
   gateway_inflight_requests gauge. Knob: GATEWAY_MAX_INFLIGHT_REQUESTS.

renderGatewayError moves to error-response.ts so the middleware can
reuse it, and 529 maps to OpenAI {type/code: overloaded} and Anthropic
{type: overloaded_error}.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Register backpressure after CORS so shed 529s carry Access-Control-*
  headers; browser clients can now see the response and Retry-After hint
  instead of a CORS/network failure.
- Await the listen bind in startServer (resolve on 'listening', reject on
  'error') so bind failures (e.g. EADDRINUSE) exit(1) instead of being
  swallowed by the new log-and-continue uncaughtException handler.
- Read GATEWAY_MAX_INFLIGHT_REQUESTS and LISTEN_BACKLOG at use-site
  instead of caching at module load, per the repo env-read guideline.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add a "Gateway Overload (529)" section explaining the transient,
retryable load-shedding response (Retry-After, OpenAI/Anthropic
envelopes) and how it differs from a 429. Add the 529 -> overloaded
row to the error-handling status-code reference.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add a gateway_requests_shed_total counter incremented when the
backpressure middleware rejects a request with HTTP 529. Enables a
direct sheds/sec signal for dashboards and alerting instead of
inferring it from the in-flight gauge riding the cap.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Scope the pod-wide 529 backpressure cap to inference POSTs only (cheap
routes are never counted or shed) and skip internal app.request() hops,
which the previous revision double-counted. Add a fleet-wide per-org
in-flight budget in the org rate-limit middleware: slots live in a Redis
zset for the response's full lifetime, stale slots from crashed pods are
reaped, and over-limit requests get a retryable 429. Enterprise orgs
skip RPM limits but get an elevated concurrency ceiling instead of an
exemption. Slot ops use explicit pipelines because bare auto-pipelined
commands issued from the response-close context wedge ioredis.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@steebchen steebchen changed the title feat(gateway): degrade gracefully under overload feat(gateway): per-org concurrency + overload caps Aug 20, 2026

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 4a77de6b2d

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +343 to +347
const checkResults = await redisClient
.pipeline()
.zremrangebyscore(key, "-inf", staleBefore)
.zcard(key)
.exec();

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Make the concurrency check-and-add atomic

During a same-organization burst across gateway pods, multiple acquisitions can all complete this pipeline, observe currentCount < limit, and then independently execute the separate ZADD at line 377. The number of admitted requests can therefore exceed the configured ceiling by the size of the concurrent burst, defeating the fleet-wide protection precisely under overload; perform the reap, count check, and insertion in one atomic Redis script or transaction.

Useful? React with 👍 / 👎.

"spend_cap_daily",
"spend_cap_monthly",
"topup_velocity",
"concurrency",

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Surface concurrency hits in the admin dashboard

Once concurrency rejections are flushed, this API accepts and returns the new type, but the summary has no concurrencyHits bucket while totalHits includes it, and the unchanged ee/admin/src/app/limit-hits/page.tsx filter and columns only cover RPM, spend caps, and top-ups. As a result, totals no longer reconcile and admins cannot select concurrency hits from the overview; add the corresponding summary field and update the admin type lists, labels, and column.

Useful? React with 👍 / 👎.

Chat plan concurrency drops to 10 and dev plan to 50. Regular (PAYG)
org ceilings now scale with the trust tier (100/200/400/1000/2000 for
tiers 0-4, env-overridable per tier), resolved lazily only once an org
reaches the tier-0 base — mirroring the RPM multiplier. Concurrency
denials get full admin limit-hit parity: a concurrencyHits column in
the admin summary API, a filter option and column on the Limit Hits
page, and a labeled badge on the per-org rate-limits breakdown.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 624cb761b4

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +33 to +35
if (outgoing) {
outgoing.once("close", runCleanup);
return await next();

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Release slots when the response already closed

If a client disconnects while the org/token/RPM/Redis lookups are still awaiting, the Node adapter can emit close before acquireOrgInflightSlot returns. This helper then registers a listener for an event that has already occurred, so the newly added Redis slot is retained until the 30-minute stale timeout; a burst of timed-out requests can therefore exhaust a chat org's 10-slot budget and falsely reject subsequent traffic. Check outgoing.destroyed/completion before registering the listener and run cleanup immediately when it is already closed.

Useful? React with 👍 / 👎.

Comment on lines +148 to +150
if (c.req.method === "POST" && INFLIGHT_LIMITED_KEYS.has(config.key)) {
const baseLimit = getOrgInflightLimit(planClass, isEnterprise);
const acquisition = await acquireOrgInflightSlot(

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Expose the enterprise concurrency ceiling in the Limits UI

Enterprise requests now receive the new per-org concurrency check, but the customer Limits page still says enterprise organizations have no per-organization rate limits and that throughput is bounded only by credits/upstream providers (apps/ui/src/app/dashboard/[orgId]/org/limits/_components/limits-client.tsx:103-107), while the limits API does not return inflightLimit. Enterprise users who encounter the new 429 therefore see a dashboard that explicitly denies the enforced 2,000-request ceiling; return and render the concurrency limit, including in the enterprise branch.

Useful? React with 👍 / 👎.

@steebchen
steebchen merged commit bf88026 into main Aug 21, 2026
17 checks passed
@steebchen
steebchen deleted the colombo-v2 branch August 21, 2026 19:16
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants