Skip to content

docs(routing): update routing internals & enterprise config - #2462

Merged
steebchen merged 1 commit into
mainfrom
docs/routing-update
May 31, 2026
Merged

steebchen merged 1 commit into
mainfrom
docs/routing-update

Conversation

@steebchen

@steebchen steebchen commented May 31, 2026 •

Copy link
Copy Markdown
Member

Summary

The routing docs (apps/docs/content/features/routing.mdx) had drifted from the actual implementation. This brings them back in line and documents the Enterprise per-project routing configuration.

What changed

  • Corrected the scoring weights. The docs claimed fixed percentages (uptime 50% / throughput 20% / price 20% / latency 10%). The real defaults are relative weights normalized by the sum of active weights: price 0.6, uptime 0.5, throughput 0.05, latency 0.025, cache 0.2, and imagePrice 1.0 (replaces price for image models). Added a table and an explanation of the ratio-based normalized scoring. Source: packages/shared/src/routing-config.ts, packages/actions/src/get-cheapest-from-available-providers.ts.
  • Fixed the metrics window. The docs repeatedly said "last 5 minutes". Routing actually aggregates a time-decayed 60-minute window (most recent 1m weighted 10×, 5m weighted 3×, remainder 1×). Source: packages/db/src/provider-metrics-history.ts.
  • Documented provider priority (default 1, 0 disables a provider) — previously undocumented.
  • Noted the configurable thresholds (uptime-penalty, exploration rate, low-uptime fallback).
  • New "Per-Project Routing Configuration (Enterprise)" section documenting that weights, thresholds, retry, timeouts, history window, sticky preference, and provider priorities can all be customized per project on the Enterprise plan via Project Settings → Routing. Source: apps/gateway/src/lib/routing-config-loader.ts (gated on orgPlan === "enterprise"), UI at apps/ui/.../settings/routing.

Verification

  • pnpm format ✅
  • pnpm build ✅ (17/17, docs MDX compiles)

🤖 Generated with Claude Code

Summary by CodeRabbit

  • Documentation
    • Updated Smart Routing Algorithm documentation with a new weighted scoring model, including metrics windows and cache support details
    • Added clarity on streaming vs non-streaming latency handling and uptime penalty behavior
    • Documented new enterprise-only Per-Project Routing Configuration feature for customizing routing parameters via dashboard

Bring the routing docs in line with the current implementation:

- Replace the stale weight percentages (uptime 50/throughput 20/price
  20/latency 10) with the actual default relative weights (price 0.6,
  uptime 0.5, throughput 0.05, latency 0.025, cache 0.2, imagePrice 1.0)
  and explain the ratio-based normalized scoring.
- Document the time-decayed 60-minute metrics window (1m ×10, 5m ×3,
  rest ×1) instead of the inaccurate "last 5 minutes" snapshot.
- Add the previously undocumented provider-priority mechanism.
- Note that the uptime-penalty, exploration-rate, and low-uptime
  fallback thresholds are configurable.
- Add a "Per-Project Routing Configuration (Enterprise)" section
  documenting that all routing values can be customized per project on
  the Enterprise plan via Project Settings -> Routing.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings May 31, 2026 10:03
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@coderabbitai

coderabbitai Bot commented May 31, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

Walkthrough

This PR updates the Smart Routing documentation to describe a new weighted scoring algorithm replacing the older "last 5 minutes" model, adds explicit default weights and normalization mechanics, documents 60-minute time-decayed metrics aggregation, and introduces cache weighting for large prompts. It also documents provider priorities, epsilon-greedy exploration, and adds a new Enterprise section for per-project routing configuration.

Changes

Routing Algorithm and Enterprise Configuration

Layer / File(s) Summary
Smart Routing Algorithm Specification
apps/docs/content/features/routing.mdx
Replaced prior algorithm explanation with new weighted scoring system covering factor weights, normalization method, streaming-only latency handling, 60-minute time-decayed metrics window, and conditional cache weighting for large prompts (≥ 5,000 tokens). Extended narrative clarifies exponential uptime penalty thresholding, provider priority effects including disabling, and epsilon-greedy exploration defaults.
Routing Metadata and Uptime Protection
apps/docs/content/features/routing.mdx
Updated routing metadata example to include priority and cache support in per-provider score components. Reworded "Low-Uptime Protection" to base provider checks on time-decayed metrics window instead of fixed "last 5 minutes".
Enterprise Per-Project Routing Configuration
apps/docs/content/features/routing.mdx
Added new Enterprise-only section describing dashboard-configurable routing parameters including weights, thresholds, retry behavior, timeouts, metrics windows, sticky routing, and provider priorities. Updated self-hosted configuration callout to note Enterprise plan requirement for per-project customization.

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~10 minutes

Possibly related PRs

  • theopenco/llmgateway#2095: Implements cache-support weighting for large prompts (≥5,000 tokens) and adds cacheSupported routing metadata field that this documentation now describes.
🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The pull request title accurately summarizes the main changes: updating routing documentation with internal algorithm details and adding enterprise configuration section.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch docs/routing-update

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Updates the Routing documentation to match the current routing implementation and adds documentation for Enterprise per-project routing overrides.

Changes:

  • Replaces outdated fixed-percentage scoring descriptions with the current normalized relative-weight scoring model (including cache and image pricing behavior).
  • Corrects the described metrics window to the current time-decayed 60-minute aggregation.
  • Documents provider priority and adds an Enterprise “Per-Project Routing Configuration” section outlining configurable groups and defaults.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

- **Throughput (20%)** - Favors providers with higher tokens per second generation speed
- **Price (20%)** - Considers cost efficiency while maintaining quality
- **Latency (10%)** - Considers time to first token (only applied for streaming requests)
Each factor has a **relative weight**. The factors are scored as ratios against the best provider in the candidate set (e.g. a provider that is twice as expensive as the cheapest scores `1.0` on price), and each ratio is multiplied by its weight divided by the sum of all active weights. The provider with the lowest (best) total score wins.
Comment on lines +64 to +66
- The most recent **1 minute** is weighted **10×**
- The most recent **5 minutes** are weighted **3×**
- The remainder of the 60-minute window is weighted **1×**

**Provider Priority**:

Each provider has a **priority** value (default `1`) that nudges routing toward or away from it independently of live metrics:
| **Timeouts** | Per-request time limits (end-to-end, streaming, non-streaming). Capped at the infrastructure defaults — an override can only lower them | `gatewayMs 1,500,000`, `streamingMs 1,200,000`, `plainMs 600,000` |
| **History** | The metrics window and the time-decay tier boundaries and weights | `windowMinutes 60` (max 120), `tier1Minutes 1`, `tier2Minutes 5`, `tier1Weight 10`, `tier2Weight 3`, `tier3Weight 1` |
| **Sticky** | Stable-provider preference: on/off, TTL, hard-switch uptime floor, soft-switch score margin | `enabled true`, `ttlSeconds 3600`, `uptimeThreshold 85`, `scoreMargin 0.15` |
| **Provider priorities** | Per-provider priority multipliers; set a provider to `0` to disable it for that project | `1` for every provider |
| **Weights** | Relative importance of each scoring factor | `price 0.6`, `imagePrice 1.0`, `uptime 0.5`, `throughput 0.05`, `latency 0.025`, `cache 0.2` |
| **Thresholds** | Cache prompt-size threshold, uptime-penalty threshold, exploration rate, and the assumed defaults used when no metrics exist | `cachePromptTokens 5000`, `uptimePenalty 95`, `defaultUptime 100`, `defaultLatency 1000`, `defaultThroughput 50`, `explorationRate 0.01` |
| **Retry** | Max cross-provider fallback attempts and the low-uptime reroute threshold | `maxRetries 2`, `lowUptimeFallbackThreshold 90` |
| **Timeouts** | Per-request time limits (end-to-end, streaming, non-streaming). Capped at the infrastructure defaults — an override can only lower them | `gatewayMs 1,500,000`, `streamingMs 1,200,000`, `plainMs 600,000` |
@steebchen
steebchen enabled auto-merge May 31, 2026 10:07

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (1)
apps/docs/content/features/routing.mdx (1)

64-66: 💤 Low value

Consider clarifying the time-decay tier boundaries.

The phrasing "most recent 1 minute" and "most recent 5 minutes" could be interpreted as overlapping ranges. Consider making it explicit whether these are exclusive tiers or cumulative:

If the tiers are exclusive (most likely):

  • Minutes 0–1: weighted 10×
  • Minutes 1–5: weighted 3×
  • Minutes 5–60: weighted 1×
♻️ Suggested clarification
-Provider metrics (uptime, throughput, latency) are not a flat "last N minutes" snapshot. They are aggregated over a rolling **60-minute window** with a time-decay weighting so very recent behavior dominates while older data still contributes:
-
-- The most recent **1 minute** is weighted **10×**
-- The most recent **5 minutes** are weighted **3×**
-- The remainder of the 60-minute window is weighted **1×**
+Provider metrics (uptime, throughput, latency) are not a flat "last N minutes" snapshot. They are aggregated over a rolling **60-minute window** with a time-decay weighting so very recent behavior dominates while older data still contributes:
+
+- **0–1 minutes ago**: weighted **10×**
+- **1–5 minutes ago**: weighted **3×**
+- **5–60 minutes ago**: weighted **1×**
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@apps/docs/content/features/routing.mdx` around lines 64 - 66, Update the
phrasing around the time-decay tiers to explicitly show exclusive boundaries so
readers don’t interpret them as overlapping; replace the ambiguous lines "most
recent 1 minute", "most recent 5 minutes", "remainder of the 60-minute window"
with explicit ranges such as "Minutes 0–1: weighted 10×", "Minutes 1–5: weighted
3×", and "Minutes 5–60: weighted 1×" (referencing the existing phrases "most
recent 1 minute", "most recent 5 minutes", and "60-minute window" in the current
text).
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@apps/docs/content/features/routing.mdx`:
- Line 41: Update the routing docs text to reflect how priceScore is actually
computed: state that priceScore = (providerPrice / minPrice) - 1 so the cheapest
provider scores 0 and a provider twice as expensive scores 1.0, and explicitly
mention the “-1 offset” rather than describing it as a raw ratio; reference the
terms priceScore and minPrice used in the implementation for clarity.

---

Nitpick comments:
In `@apps/docs/content/features/routing.mdx`:
- Around line 64-66: Update the phrasing around the time-decay tiers to
explicitly show exclusive boundaries so readers don’t interpret them as
overlapping; replace the ambiguous lines "most recent 1 minute", "most recent 5
minutes", "remainder of the 60-minute window" with explicit ranges such as
"Minutes 0–1: weighted 10×", "Minutes 1–5: weighted 3×", and "Minutes 5–60:
weighted 1×" (referencing the existing phrases "most recent 1 minute", "most
recent 5 minutes", and "60-minute window" in the current text).
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro

Run ID: ddeab870-cffa-40d5-86a3-76f04c0acb25

📥 Commits

Reviewing files that changed from the base of the PR and between 3b14c4a and ab6ebf8.

📒 Files selected for processing (1)
  • apps/docs/content/features/routing.mdx

- **Throughput (20%)** - Favors providers with higher tokens per second generation speed
- **Price (20%)** - Considers cost efficiency while maintaining quality
- **Latency (10%)** - Considers time to first token (only applied for streaming requests)
Each factor has a **relative weight**. The factors are scored as ratios against the best provider in the candidate set (e.g. a provider that is twice as expensive as the cheapest scores `1.0` on price), and each ratio is multiplied by its weight divided by the sum of all active weights. The provider with the lowest (best) total score wins.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor | ⚡ Quick win

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# Search for the price ratio calculation in the routing logic
rg -nP -C5 --type=ts 'price.*ratio|score.*price' --glob '*get-cheapest-from-available-providers.ts' --glob '*routing*.ts'

Repository: theopenco/llmgateway

Length of output: 2465


🏁 Script executed:

#!/bin/bash
set -euo pipefail

FILE="packages/actions/src/get-cheapest-from-available-providers.ts"

# 1) Find likely scoring/ratio variables
rg -n "score|ratio|normalize|normalized|bestPrice|minPrice|priceScore|priceWeight|effectivePriceWeight|uptime|timeDecay|decay" "$FILE"

# 2) Pull the sections around providerScores construction (where scoring is likely computed)
#    Use a broader context window around "providerScores" occurrences
rg -n "providerScores" "$FILE" -C20

# 3) Find explicit price normalization/comparison
rg -n "best.*price|min.*price|price.*(best|min)|cheapest.*price|price.*ratio" "$FILE" -C10

# 4) If weights are used to compute a total score, locate "totalScore" / "weighted" math
rg -n "totalScore|weighted|weights\.|sum.*weights|active weights|dominates|dominant" "$FILE" -C20

Repository: theopenco/llmgateway

Length of output: 23995


Correct the wording for the “price ratio” example in routing docs: the implementation computes priceScore as (providerPrice / minPrice) - 1 (where minPrice is the cheapest provider), so a provider that is 2x as expensive yields priceScore = 1.0—the example at apps/docs/content/features/routing.mdx line 41 is numerically consistent. Update the text to clarify the -1 offset (cheapest scores 0, 2x scores 1.0) rather than implying a raw ratio.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@apps/docs/content/features/routing.mdx` at line 41, Update the routing docs
text to reflect how priceScore is actually computed: state that priceScore =
(providerPrice / minPrice) - 1 so the cheapest provider scores 0 and a provider
twice as expensive scores 1.0, and explicitly mention the “-1 offset” rather
than describing it as a raw ratio; reference the terms priceScore and minPrice
used in the implementation for clarity.

@steebchen
steebchen added this pull request to the merge queue May 31, 2026
Merged via the queue into main with commit 611acdf May 31, 2026
13 checks passed
@steebchen
steebchen deleted the docs/routing-update branch May 31, 2026 10:19
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants