Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions apps/api/src/routes/logs.ts
Original file line number Diff line number Diff line change
Expand Up @@ -125,6 +125,7 @@ const logSchema = z.object({
throughput: z.number().optional(),
price: z.number().optional(),
priority: z.number().optional(),
cacheSupported: z.boolean().optional(),
failed: z.boolean().optional(),
status_code: z.number().optional(),
error_type: z.string().optional(),
Expand Down
2 changes: 2 additions & 0 deletions apps/code/src/lib/api/v1.d.ts
Original file line number Diff line number Diff line change
Expand Up @@ -992,6 +992,7 @@ export interface paths {
throughput?: number;
price?: number;
priority?: number;
cacheSupported?: boolean;
failed?: boolean;
status_code?: number;
error_type?: string;
Expand Down Expand Up @@ -1236,6 +1237,7 @@ export interface paths {
throughput?: number;
price?: number;
priority?: number;
cacheSupported?: boolean;
failed?: boolean;
status_code?: number;
error_type?: string;
Expand Down
4 changes: 4 additions & 0 deletions apps/docs/content/features/routing.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -49,6 +49,10 @@ The algorithm calculates a weighted score for each available provider and select

For non-streaming requests, the latency weight (10%) is redistributed proportionally to the other factors since time-to-first-token is less relevant when waiting for the complete response.

**Cache Support for Large Prompts**:

When the estimated prompt is at least 5,000 tokens, an additional 20% weight is added to the score for whether each provider supports prompt caching (advertised via a cached input price). Providers that support caching score better than ones that do not, since caching can substantially reduce the cost of large or repeated prompts. Below the 5k threshold, this weight is dropped entirely — caching has little impact on small prompts, so cache support is ignored. The selected provider's cache support is exposed as `cacheSupported` on the routing metadata.

Copilot AI Apr 27, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The docs say cache support is exposed as cacheSupported "on the routing metadata", but in code it’s added on each entry of metadata.providerScores (and not as a top-level RoutingMetadata.cacheSupported). Please clarify the exact field path (e.g. routingMetadata.providerScores[].cacheSupported) to avoid consumers looking for a non-existent top-level property.

Suggested change
When the estimated prompt is at least 5,000 tokens, an additional 20% weight is added to the score for whether each provider supports prompt caching (advertised via a cached input price). Providers that support caching score better than ones that do not, since caching can substantially reduce the cost of large or repeated prompts. Below the 5k threshold, this weight is dropped entirely — caching has little impact on small prompts, so cache support is ignored. The selected provider's cache support is exposed as `cacheSupported` on the routing metadata.
When the estimated prompt is at least 5,000 tokens, an additional 20% weight is added to the score for whether each provider supports prompt caching (advertised via a cached input price). Providers that support caching score better than ones that do not, since caching can substantially reduce the cost of large or repeated prompts. Below the 5k threshold, this weight is dropped entirely — caching has little impact on small prompts, so cache support is ignored. Cache support is exposed on each provider score entry in the routing metadata as `routingMetadata.providerScores[].cacheSupported`.

Copilot uses AI. Check for mistakes.

**Exponential Uptime Penalty**:

Providers with uptime below 95% receive an additional exponential penalty that increases rapidly as uptime drops:
Expand Down
1 change: 0 additions & 1 deletion apps/gateway/package.json
Original file line number Diff line number Diff line change
Expand Up @@ -40,7 +40,6 @@
"@opentelemetry/sdk-node": "0.205.0",
"decimal.js": "10.5.0",
"dotenv": "16.5.0",
"gpt-tokenizer": "3.0.1",
"hono": "^4.12.7",
"hono-openapi": "^1.1.1",
"ioredis": "5.7.0",
Expand Down
126 changes: 50 additions & 76 deletions apps/gateway/src/chat/chat.ts
Original file line number Diff line number Diff line change
@@ -1,5 +1,4 @@
import { createRoute, OpenAPIHono, z } from "@hono/zod-openapi";
import { encode } from "gpt-tokenizer";
import { HTTPException } from "hono/http-exception";
import { streamSSE } from "hono/streaming";

Expand Down Expand Up @@ -245,6 +244,7 @@ function collapseProvidersToBestRegionPerProvider(
options: {
metricsMap: Map<string, ProviderMetrics>;
isStreaming: boolean;
promptTokens?: number;
},
): ProviderModelMapping[] {
const providersById = new Map<string, ProviderModelMapping[]>();
Expand Down Expand Up @@ -1087,6 +1087,18 @@ chat.openapi(completions, async (c) => {
}
}

// Estimate prompt tokens once so all routing decisions can reuse the
// same value (e.g. cache-support weighting kicks in for large prompts).
// Uses a cheap chars/4 heuristic — accuracy is intentionally traded
// for throughput on the gateway hot path.
let routingPromptTokens = 0;
if (messages && messages.length > 0) {
routingPromptTokens = encodeChatMessages(messages);
}
if (tools && tools.length > 0) {
routingPromptTokens += Math.round(JSON.stringify(tools).length / 4);
}
Comment thread
steebchen marked this conversation as resolved.

// Extract and validate source from x-source header with HTTP-Referer fallback
let source = validateSource(
c.req.header("x-source"),
Expand Down Expand Up @@ -1498,34 +1510,9 @@ chat.openapi(completions, async (c) => {
(usedProvider === "llmgateway" && usedModel === "auto") ||
usedModel === "auto"
) {
// Estimate prompt/input tokens first so auto-routing can react to large prompts
let estimatedInputTokens = 0;

// Estimate prompt tokens from messages
if (messages && messages.length > 0) {
try {
estimatedInputTokens = encodeChatMessages(messages);
} catch {
// Fallback to simple estimation if encoding fails
const messageTokens = messages.reduce(
(acc, m) => acc + (m.content?.length ?? 0),
0,
);
estimatedInputTokens = Math.max(1, Math.round(messageTokens / 4));
}
}

// Add tool definitions to context estimation
if (tools && tools.length > 0) {
try {
const toolsString = JSON.stringify(tools);
const toolTokens = Math.round(toolsString.length / 4);
estimatedInputTokens += toolTokens;
} catch {
// Fallback estimation for tools
estimatedInputTokens += tools.length * 100; // Rough estimate per tool
}
}
// Reuse the prompt-token estimate computed earlier so auto-routing can
// react to large prompts when picking a model.
const estimatedInputTokens = routingPromptTokens;

// Estimate the full context needed based on the request
let requiredContextSize = estimatedInputTokens;
Expand Down Expand Up @@ -1744,13 +1731,18 @@ chat.openapi(completions, async (c) => {
{
metricsMap,
isStreaming: stream,
promptTokens: routingPromptTokens,
},
);

const cheapestResult = getCheapestFromAvailableProviders(
providerAgnosticSelectedProviders,
selectedModel,
{ metricsMap, isStreaming: stream },
{
metricsMap,
isStreaming: stream,
promptTokens: routingPromptTokens,
},
);

if (cheapestResult) {
Expand Down Expand Up @@ -1910,6 +1902,7 @@ chat.openapi(completions, async (c) => {
{
metricsMap,
isStreaming: stream,
promptTokens: routingPromptTokens,
},
);

Expand Down Expand Up @@ -2094,6 +2087,7 @@ chat.openapi(completions, async (c) => {
{
metricsMap: allMetricsMap,
isStreaming: stream,
promptTokens: routingPromptTokens,
},
);

Expand Down Expand Up @@ -2231,7 +2225,11 @@ chat.openapi(completions, async (c) => {
collapseProvidersToBestRegionPerProvider(
availableModelProviders,
modelWithPricing,
{ metricsMap: allMetricsMap, isStreaming: stream },
{
metricsMap: allMetricsMap,
isStreaming: stream,
promptTokens: routingPromptTokens,
},
);

// Filter to only providers with better uptime than the original
Expand All @@ -2256,7 +2254,11 @@ chat.openapi(completions, async (c) => {
const cheapestResult = getCheapestFromAvailableProviders(
betterUptimeProviders,
modelWithPricing,
{ metricsMap: allMetricsMap, isStreaming: stream },
{
metricsMap: allMetricsMap,
isStreaming: stream,
promptTokens: routingPromptTokens,
},
);

// Get price info for the original requested provider to include in scores
Expand Down Expand Up @@ -2433,13 +2435,21 @@ chat.openapi(completions, async (c) => {
collapseProvidersToBestRegionPerProvider(
routingCandidates,
modelWithPricing,
{ metricsMap, isStreaming: stream },
{
metricsMap,
isStreaming: stream,
promptTokens: routingPromptTokens,
},
);

const cheapestResult = getCheapestFromAvailableProviders(
providerAgnosticCandidates,
modelWithPricing,
{ metricsMap, isStreaming: stream },
{
metricsMap,
isStreaming: stream,
promptTokens: routingPromptTokens,
},
);

if (cheapestResult) {
Expand Down Expand Up @@ -2606,6 +2616,7 @@ chat.openapi(completions, async (c) => {
{
metricsMap,
isStreaming: stream,
promptTokens: routingPromptTokens,
},
);

Expand Down Expand Up @@ -6866,27 +6877,8 @@ chat.openapi(completions, async (c) => {
imageTokens = 258 + Math.ceil(imageByteSize / 750);
}

// Skip expensive token encoding for image responses - use simple estimation
// Token encoding on large base64 content causes CPU spikes
if (imageByteSize > 0) {
const textTokens = estimateTokensFromContent(fullContent);
calculatedCompletionTokens = textTokens + imageTokens;
} else {
try {
const textTokens = fullContent
? encode(JSON.stringify(fullContent)).length
: 0;
calculatedCompletionTokens = textTokens + imageTokens;
} catch (error) {
// Fallback to simple estimation if encoding fails
logger.error(
"Failed to encode completion text in streaming",
error instanceof Error ? error : new Error(String(error)),
);
const textTokens = estimateTokensFromContent(fullContent);
calculatedCompletionTokens = textTokens + imageTokens;
}
}
const textTokens = estimateTokensFromContent(fullContent);
calculatedCompletionTokens = textTokens + imageTokens;
}

calculatedTotalTokens =
Expand All @@ -6896,17 +6888,8 @@ chat.openapi(completions, async (c) => {
// Estimate reasoning tokens if not provided but reasoning content exists
let calculatedReasoningTokens = reasoningTokens;
if (!reasoningTokens && fullReasoningContent) {
try {
calculatedReasoningTokens = encode(fullReasoningContent).length;
} catch (error) {
// Fallback to simple estimation if encoding fails
logger.error(
"Failed to encode reasoning text in streaming",
error instanceof Error ? error : new Error(String(error)),
);
calculatedReasoningTokens =
estimateTokensFromContent(fullReasoningContent);
}
calculatedReasoningTokens =
estimateTokensFromContent(fullReasoningContent);
Comment on lines 6889 to +6892

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

Thread the estimated reasoning tokens through the downstream accounting.

calculatedReasoningTokens is populated here, but the later calculateCosts(...), total-token math, usage chunks, and transformResponseToOpenai(...) still use raw reasoningTokens. When upstream omits reasoning usage, the new fallback only affects some log fields while response metadata and billing-related totals still undercount reasoning.

Also applies to: 8940-8942

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@apps/gateway/src/chat/chat.ts` around lines 6887 - 6890,
calculatedReasoningTokens is computed as a fallback but never propagated; update
downstream logic to use calculatedReasoningTokens wherever reasoningTokens is
currently referenced so the fallback is honoured: pass calculatedReasoningTokens
into calculateCosts(...), include it in total-token math and any usage chunk
computations (e.g. usageChunks/usage totals), and ensure
transformResponseToOpenai(...) receives the calculatedReasoningTokens for
response metadata. Locate uses of reasoningTokens in the same scope and replace
or augment them to prefer calculatedReasoningTokens when reasoningTokens is
falsy so billing and metadata reflect the estimated reasoning usage (also apply
the same change near the other occurrence around the 8940–8942 region).

}

if (
Expand Down Expand Up @@ -8962,16 +8945,7 @@ chat.openapi(completions, async (c) => {
// Estimate reasoning tokens if not provided but reasoning content exists
let calculatedReasoningTokens = reasoningTokens;
if (!reasoningTokens && reasoningContent) {
try {
calculatedReasoningTokens = encode(reasoningContent).length;
} catch (error) {
// Fallback to simple estimation if encoding fails
logger.error(
"Failed to encode reasoning text",
error instanceof Error ? error : new Error(String(error)),
);
calculatedReasoningTokens = estimateTokensFromContent(reasoningContent);
}
calculatedReasoningTokens = estimateTokensFromContent(reasoningContent);
}
const costs = await calculateCosts(
usedModel,
Expand Down
10 changes: 5 additions & 5 deletions apps/gateway/src/chat/tools/estimate-tokens-from-content.ts
Original file line number Diff line number Diff line change
@@ -1,9 +1,9 @@
import { estimateTokensFromText } from "@llmgateway/shared";

/**
* Estimates tokens from content length using simple division
* Estimates tokens from content length using a chars/4 heuristic. Backed by
* the shared text-only estimator.
*/
export function estimateTokensFromContent(content: string): number {
if (!content) {
return 0;
}
return Math.max(1, Math.round(content.length / 4));
return estimateTokensFromText(content);
}
23 changes: 5 additions & 18 deletions apps/gateway/src/chat/tools/estimate-tokens.ts
Original file line number Diff line number Diff line change
@@ -1,13 +1,12 @@
import { encode } from "gpt-tokenizer";

import { logger } from "@llmgateway/logger";

import { estimateTokensFromContent } from "./estimate-tokens-from-content.js";
import { encodeChatMessages } from "./tokenizer.js";

import type { Provider } from "@llmgateway/models";

/**
* Estimates token counts when not provided by the API using gpt-tokenizer
* Estimates token counts when not provided by the API. Uses a cheap
* length-based heuristic rather than running a tokenizer on the
* gateway hot path.
*/
export function estimateTokens(
usedProvider: Provider,
Expand All @@ -19,25 +18,13 @@ export function estimateTokens(
let calculatedPromptTokens = promptTokens;
let calculatedCompletionTokens = completionTokens;

// Always estimate missing tokens for any provider
if (!promptTokens || !completionTokens) {
// Estimate prompt tokens using encodeChat for better accuracy
if (!promptTokens && messages && messages.length > 0) {
calculatedPromptTokens = encodeChatMessages(messages);
}

// Estimate completion tokens using encode for better accuracy
if (!completionTokens && content) {
try {
calculatedCompletionTokens = encode(JSON.stringify(content)).length;
} catch (error) {
// Fallback to simple estimation if encoding fails
logger.error(
"Failed to encode completion text",
error instanceof Error ? error : new Error(String(error)),
);
calculatedCompletionTokens = content.length / 4;
}
calculatedCompletionTokens = estimateTokensFromContent(content);
}
Comment thread
coderabbitai[bot] marked this conversation as resolved.
}

Expand Down
46 changes: 10 additions & 36 deletions apps/gateway/src/chat/tools/tokenizer.ts
Original file line number Diff line number Diff line change
@@ -1,12 +1,10 @@
import { encodeChat } from "gpt-tokenizer";

import { logger } from "@llmgateway/logger";

import { DEFAULT_TOKENIZER_MODEL } from "./types.js";
import { estimateChatMessageTokens } from "@llmgateway/shared";

/**
* Converts a message content value (string, array of content parts, null, or
* undefined) to a plain string suitable for the gpt-tokenizer library.
* undefined) to a plain string. Used by call sites that need a flat string
* (e.g. for cost estimation) — not for token counting; see
* `encodeChatMessages` for that.
*/
export function messageContentToString(
content: string | unknown[] | null | undefined,
Expand All @@ -21,36 +19,12 @@ export function messageContentToString(
}

/**
* Encodes an array of chat messages and returns the token count. Handles
* messages whose content may be a string, an array of content parts, null, or
* undefined – all of which are valid shapes in the OpenAI chat format but
* would otherwise crash gpt-tokenizer.
* Rough length-based prompt-token estimate for a chat message array.
*
* Backed by the shared `estimateChatMessageTokens` helper, which only counts
* text and ignores multimodal parts (image_url, file, etc.). Image input
* billing is handled separately in costs.ts.
*/
export function encodeChatMessages(messages: any[]): number {
try {
const chatMessages = messages.map((m) => ({
role: m.role as "user" | "assistant" | "system" | undefined,
content: messageContentToString(m.content),
...(m.name !== null && m.name !== undefined && { name: m.name }),
}));
return encodeChat(chatMessages, DEFAULT_TOKENIZER_MODEL).length;
} catch (error) {
logger.error("Failed to encode chat messages", {
error: error instanceof Error ? error.message : String(error),
messageCount: messages.length,
messageRoles: messages.map((m) => m.role),
messageContentTypes: messages.map((m) => typeof m.content),
});
// Fallback: rough 4-chars-per-token estimate
return Math.max(
1,
Math.round(
messages.reduce(
(acc: number, m: any) =>
acc + messageContentToString(m.content).length,
0,
) / 4,
),
);
}
return estimateChatMessageTokens(messages);
}
Comment on lines 28 to 30

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# Verify current `any[]` usage and inspect encodeChatMessages call sites before tightening types.
rg -n --type=ts -C2 'export function encodeChatMessages\(messages:\s*any\[\]\)'
rg -n --type=ts -C2 '\bencodeChatMessages\('

Repository: theopenco/llmgateway

Length of output: 2749


🏁 Script executed:

#!/bin/bash
# Check the actual types of messages at call sites
# 1. Check costs.ts to see type of fullOutput.messages
head -160 apps/gateway/src/lib/costs.ts | tail -20

# 2. Check chat.ts to see type of messages parameter where encodeChatMessages is called
grep -B20 "encodeChatMessages(messages)" apps/gateway/src/chat/chat.ts | head -40

# 3. Look for existing message type definitions
rg -n "type.*Message|interface.*Message" --type=ts | grep -v node_modules | head -20

# 4. Check estimate-tokens.ts
rg -B5 "encodeChatMessages" apps/gateway/src/chat/tools/estimate-tokens.ts

# 5. Check the estimateChatMessageTokens function signature
rg -A10 "function estimateChatMessageTokens" --type=ts

Repository: theopenco/llmgateway

Length of output: 5181


🏁 Script executed:

#!/bin/bash
# Get the MessageLike interface definition
cat -n packages/shared/src/token-estimate.ts | head -50

# Check if encodeChatMessages is defined in tokenizer.ts with more context
cat -n apps/gateway/src/chat/tools/tokenizer.ts

# Check the types imported/used in chat.ts for the messages parameter
grep -B10 "function.*messages" apps/gateway/src/chat/chat.ts | head -30

# Check if there are type definitions for fullOutput in costs.ts
rg -B5 "fullOutput" apps/gateway/src/lib/costs.ts | head -30

Repository: theopenco/llmgateway

Length of output: 4083


🏁 Script executed:

#!/bin/bash
# Check if MessageLike is exported from the shared package
rg -n "export.*MessageLike|export.*interface.*MessageLike" packages/shared/src/

# Check the exports in token-estimate.ts
grep -n "export" packages/shared/src/token-estimate.ts

# Check what's currently imported in tokenizer.ts
head -10 apps/gateway/src/chat/tools/tokenizer.ts

# Check if `@llmgateway/shared` is available as a dependency in gateway
grep -n "@llmgateway/shared" apps/gateway/package.json

# Check the full token-estimate.ts to see if MessageLike is exported
cat packages/shared/src/token-estimate.ts

Repository: theopenco/llmgateway

Length of output: 2476


Replace any[] in encodeChatMessages with a concrete message type.

Line 28 uses any[], which violates the coding guidelines and drops type safety on a routing/cost-critical path. The function passes messages directly to estimateChatMessageTokens from @llmgateway/shared, which expects a concrete structure with content (string, array of parts, or null) and parts with optional type and text fields.

Suggested type-safe change
+type EstimationContentPart = {
+	type?: string;
+	text?: string;
+};
+
+type EstimationMessage = {
+	content?: string | EstimationContentPart[] | null;
+};
+
-export function encodeChatMessages(messages: any[]): number {
+export function encodeChatMessages(messages: EstimationMessage[]): number {
	return estimateChatMessageTokens(messages);
}
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@apps/gateway/src/chat/tools/tokenizer.ts` around lines 28 - 30, Replace the
use of any[] in encodeChatMessages with a concrete message type that matches
what estimateChatMessageTokens expects: define or import a ChatMessage
type/interface where each message includes a content field typed as string |
Array<{ type?: string; text?: string }> | null (and any other optional fields
your codebase requires), update the function signature to
encodeChatMessages(messages: ChatMessage[]): number, and pass that typed array
into estimateChatMessageTokens so the compiler enforces the correct structure;
reference the encodeChatMessages function and the estimateChatMessageTokens call
when making the change.

Loading
Loading