Skip to content

fix(gateway): respect RPM limits in low-uptime fallback - #2042

Merged
steebchen merged 3 commits into
mainfrom
fix-uptime-fallback-rpm
May 7, 2026
Merged

steebchen merged 3 commits into
mainfrom
fix-uptime-fallback-rpm

Conversation

@steebchen

@steebchen steebchen commented Apr 20, 2026 •

Copy link
Copy Markdown
Member

Summary

  • The low-uptime fallback path in apps/gateway/src/chat/chat.ts selected an alternative provider based only on uptime and price, ignoring per-provider RPM/RPD caps. A request that was rerouted away from a low-uptime provider could land on another provider that was already at its rate limit, causing an avoidable 429 at checkProviderRateLimit right after.
  • Filter availableModelProviders through filterRateLimitedProviders before scoring, with fail-open if every alternative is capped. This mirrors the existing rate-limit fallback path (same file, lines ~2056–2071), so the two fallback paths now share the same RPM-aware candidate selection.

Test plan

  • pnpm build — passed locally
  • pnpm lint — passed locally
  • Manually exercise: requested provider below 90% uptime, best-priced alternative at RPM cap → request routes to the next-cheapest non-capped alternative, not the capped one
  • Manually exercise: all alternatives are capped → fails open and still attempts fallback instead of getting stuck on the low-uptime provider

🤖 Generated with Claude Code

Summary by CodeRabbit

  • Bug Fixes
    • Improved fallback provider selection to prefer non-rate-limited candidates when available, ensuring uptime comparisons, scoring, and routing use healthier alternatives while still allowing fail-open behavior when all candidates are capped.
  • Tests
    • Added coverage for selecting non-rate-limited candidates, deduplication of rate-limit checks, empty-input handling, and fail-open scenarios.

Copilot AI review requested due to automatic review settings April 20, 2026 17:27
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@coderabbitai

coderabbitai Bot commented Apr 20, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

Caution

Review failed

The pull request is closed.

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro

Run ID: 4ba90610-45e0-445f-9a97-635735ff3925

📥 Commits

Reviewing files that changed from the base of the PR and between 8002624 and 7b09a51.

📒 Files selected for processing (3)
  • apps/gateway/src/chat/chat.ts
  • apps/gateway/src/lib/provider-rate-limit.spec.ts
  • apps/gateway/src/lib/provider-rate-limit.ts

Walkthrough

Adds exported helper pickNonRateLimitedCandidates; uses it in chat routing (requested-provider reroute and low-uptime fallback) to filter out rate-limited provider/model candidates (with fail-open), and adds tests covering filtering, fail-open, deduping, and empty-input behavior.

Changes

Rate-Limit Candidate Selection

Layer / File(s) Summary
Helper implementation
apps/gateway/src/lib/provider-rate-limit.ts
Adds exported pickNonRateLimitedCandidates that deduplicates provider+model peeks, calls filterRateLimitedProviders, filters candidates, and returns original candidates if filtering yields none (fail-open).
Chat import & requested-provider reroute
apps/gateway/src/chat/chat.ts
Imports pickNonRateLimitedCandidates and replaces inline filterRateLimitedProviders + manual fail-open logic in the "requested provider is rate-limited" reroute with a call to the helper.
Low-uptime fallback scoring
apps/gateway/src/chat/chat.ts
Compute uptimeFallbackCandidates via the helper, gate fallback scoring on its length, and use it for metrics combination lookups and collapseProvidersToBestRegionPerProvider.
Tests
apps/gateway/src/lib/provider-rate-limit.spec.ts
Adds tests for pickNonRateLimitedCandidates: dropping rate-limited candidates, failing open when all are limited, deduping region-expanded variants (single DB/Redis peek), and early return on empty candidate lists; updates imports.

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly related PRs

  • theopenco/llmgateway#2033: Modifies the same low-uptime fallback routing logic in apps/gateway/src/chat/chat.ts to update usedRegion handling in the fallback path.
  • theopenco/llmgateway#2042: Also modifies low-uptime fallback candidate selection in apps/gateway/src/chat/chat.ts; similar intent to prefer non–rate-limited alternatives before uptime/price scoring.
🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and specifically describes the main change: applying RPM rate-limit filtering to the low-uptime fallback routing logic. It directly matches the PR's core objective of respecting rate limits during fallback selection.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix-uptime-fallback-rpm

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Aligns the gateway’s low-uptime fallback routing with the existing rate-limit fallback routing by avoiding (when possible) providers already at their configured RPM/RPD caps, reducing avoidable immediate 429s after reroute.

Changes:

  • Filters low-uptime fallback alternatives via filterRateLimitedProviders, with fail-open behavior when all alternatives are capped.
  • Routes scoring/metrics collection using the resulting uptimeFallbackCandidates instead of the full availableModelProviders.
  • Keeps behavior consistent with the pre-existing rate-limit fallback candidate filtering logic in the same file.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread apps/gateway/src/chat/chat.ts Outdated
Comment on lines +2215 to +2221
const rateLimitedAlternatives = await filterRateLimitedProviders(
project.organizationId,
availableModelProviders.map((p) => ({
providerId: p.providerId,
model: baseModelId,
providerModelName: p.modelName,
})),

Copilot AI Apr 20, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

filterRateLimitedProviders is called once per entry in availableModelProviders, which likely includes multiple regions per provider (it’s later collapsed via collapseProvidersToBestRegionPerProvider). This can trigger redundant rate-limit peeks (DB + Redis) for the same provider/model, increasing latency in the low-uptime fallback path. Consider de-duplicating the candidates (e.g., by providerId + providerModelName) before calling filterRateLimitedProviders, while still filtering the full availableModelProviders list using the resulting Set.

Suggested change
const rateLimitedAlternatives = await filterRateLimitedProviders(
project.organizationId,
availableModelProviders.map((p) => ({
providerId: p.providerId,
model: baseModelId,
providerModelName: p.modelName,
})),
const uniqueRateLimitCandidates = Array.from(
new Map(
availableModelProviders.map((p) => [
`${p.providerId}:${p.modelName}`,
{
providerId: p.providerId,
model: baseModelId,
providerModelName: p.modelName,
},
]),
).values(),
);
const rateLimitedAlternatives = await filterRateLimitedProviders(
project.organizationId,
uniqueRateLimitCandidates,

Copilot uses AI. Check for mistakes.
Comment thread apps/gateway/src/chat/chat.ts Outdated
Comment on lines +2212 to +2229
// Exclude alternatives that are already at their RPM/RPD cap so the
// low-uptime fallback doesn't route into a rate-limited provider.
// Fail-open: if all alternatives are rate-limited, keep them all.
const rateLimitedAlternatives = await filterRateLimitedProviders(
project.organizationId,
availableModelProviders.map((p) => ({
providerId: p.providerId,
model: baseModelId,
providerModelName: p.modelName,
})),
);
const nonRateLimitedAlternatives = availableModelProviders.filter(
(p) => !rateLimitedAlternatives.has(p.providerId),
);
const uptimeFallbackCandidates =
nonRateLimitedAlternatives.length > 0
? nonRateLimitedAlternatives
: availableModelProviders;

Copilot AI Apr 20, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This change introduces new routing behavior (skipping rate-limited alternatives during low-uptime fallback) but there doesn’t appear to be an integration test covering the scenario where the cheapest/healthiest alternative is at its provider RPM/RPD cap and the router should pick the next-best candidate (and also the fail-open case when all alternatives are capped). Please add a regression test (likely alongside existing low-uptime fallback tests in apps/gateway/src/fallback.spec.ts) to prevent this from regressing.

Copilot uses AI. Check for mistakes.
Comment thread apps/gateway/src/chat/chat.ts Outdated
Comment on lines +2212 to +2229
// Exclude alternatives that are already at their RPM/RPD cap so the
// low-uptime fallback doesn't route into a rate-limited provider.
// Fail-open: if all alternatives are rate-limited, keep them all.
const rateLimitedAlternatives = await filterRateLimitedProviders(
project.organizationId,
availableModelProviders.map((p) => ({
providerId: p.providerId,
model: baseModelId,
providerModelName: p.modelName,
})),
);
const nonRateLimitedAlternatives = availableModelProviders.filter(
(p) => !rateLimitedAlternatives.has(p.providerId),
);
const uptimeFallbackCandidates =
nonRateLimitedAlternatives.length > 0
? nonRateLimitedAlternatives
: availableModelProviders;

Copilot AI Apr 20, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This rate-limit-aware candidate filtering duplicates the earlier rate-limit fallback block (same file around the rate-limit-fallback routing). Consider extracting a small helper for “pick fallback candidates with fail-open rate-limit filtering” to avoid the two paths drifting over time (and to centralize any future tweaks to the filtering rules).

Copilot uses AI. Check for mistakes.
steebchen and others added 2 commits May 7, 2026 20:55
The low-uptime fallback path selected an alternative provider purely
on uptime and price, ignoring per-provider RPM/RPD caps. Filter out
rate-limited alternatives (fail-open if all are capped), matching the
rate-limit fallback behavior.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
availableModelProviders comes from the region-expanded list, so
multiple variants of the same provider+model triggered redundant
peekProviderRateLimit calls (Redis hit per region). Dedupe by
providerId+modelName before calling filterRateLimitedProviders.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
@steebchen
steebchen force-pushed the fix-uptime-fallback-rpm branch from b4e6f99 to 8002624 Compare May 7, 2026 13:59
Both rate-limit and low-uptime fallback paths reused the same
dedupe → peek → filter → fail-open dance. Pull it into a single
helper next to filterRateLimitedProviders so future tweaks to the
filtering rules apply to both paths at once. Add unit tests covering
the next-best, fail-open, and region-dedupe cases.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants