Skip to content

feat(routing): cache weight via length-based estimate - #2104

Merged
steebchen merged 11 commits into
mainfrom
weight-cache-estimate
Apr 29, 2026
Merged

steebchen merged 11 commits into
mainfrom
weight-cache-estimate

Conversation

@steebchen

@steebchen steebchen commented Apr 27, 2026 •

Copy link
Copy Markdown
Member

Summary

  • Re-applies the cache-support routing weight from feat(routing): weight cache support for large prompts #2095, which was reverted in c94491c because the in-flight gpt-tokenizer call regressed gateway throughput on every chat request.
  • Replaces encodeChatMessages/encode with a cheap chars/4 length-based estimate everywhere on the gateway hot path. Accuracy is intentionally traded for throughput; routing only needs a rough threshold check, and post-call usage estimation only matters when the upstream omits token counts.
  • gpt-tokenizer is dropped from apps/gateway/package.json.

Code paths switched off the tokenizer

  • apps/gateway/src/chat/chat.ts — routing prompt-token estimate at the top of the chat handler (the call site that caused the regression), the streaming-response prompt/completion/reasoning estimates around line 6857, and the non-streaming reasoning estimate near line 8944.
  • apps/gateway/src/chat/tools/tokenizer.ts — encodeChatMessages now sums messageContentToString(...).length and divides by 4.
  • apps/gateway/src/chat/tools/estimate-tokens.ts — completion-token fallback now uses estimateTokensFromContent.
  • apps/gateway/src/lib/costs.ts — calculateCosts no longer runs encodeChat/encode to fill in missing prompt/completion tokens; it reuses encodeChatMessages (now length-based) and estimateTokensFromContent.
  • apps/gateway/src/chat/tools/calculate-prompt-tokens.ts — unchanged signature, but now backed by the length-based encodeChatMessages.

Test plan

  • pnpm format
  • pnpm build (turbo, all apps)
  • npx vitest run --no-file-parallelism packages/actions packages/models packages/db apps/gateway/src/lib apps/gateway/src/chat (485 unit tests passing)
  • Spot-check routing decisions in staging with prompts above and below 5k tokens (carry-over from the original PR)

Notes / implications

  • Estimates are now systematically lower-resolution. On English text the chars/4 heuristic is close to gpt-tokenizer (which itself is a heuristic outside OpenAI models), but for code, JSON, or non-Latin scripts the bias can be larger in either direction. The same heuristic was already used as the fallback path everywhere, so the worst case for billing accuracy is what you'd hit when the upstream provider omits token counts and the encoder fails — now that's the steady state.
  • Affected billing surface: calculateCosts only estimates prompt/completion tokens when the upstream response didn't include them. For mainstream providers (OpenAI, Anthropic, Google, etc.) the upstream usage is honored and the heuristic isn't used. Where it is used (e.g. some streaming paths or older providers), prompt token counts may now be a few percent off, which propagates into the displayed input/output cost.
  • The cacheSupported routing weight only kicks in above CACHE_PROMPT_TOKEN_THRESHOLD (5k tokens). Length-based estimation is more than accurate enough to gate that decision.

🤖 Generated with Claude Code

Summary by CodeRabbit

Release Notes

  • New Features

    • Intelligent routing now factors in provider prompt caching support for large prompts (5,000+ tokens), automatically favoring providers with caching capability.
    • Added visibility of image quality parameters and cache support status in activity logs and provider scoring details.
  • Documentation

    • Updated routing algorithm documentation to reflect new cache-support weighting for optimized prompt handling.

Re-applies #2095 (cache-support routing weight) using a chars/4
length-based prompt estimate instead of gpt-tokenizer. Replaces
all in-flight tokenizer usage in the gateway with the same cheap
heuristic, since the encoder caused a measured throughput regression.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings April 27, 2026 18:51
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@coderabbitai

coderabbitai Bot commented Apr 27, 2026 •

Copy link
Copy Markdown
Contributor

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

Walkthrough

This pull request removes the gpt-tokenizer dependency from the gateway and replaces it with a new shared length-based token estimation module. It centralizes prompt token computation, adds cache-support awareness to the routing scoring algorithm with conditional weighting for large prompts, and extends cost calculations and UI displays to include image quality parameters.

Changes

Cohort / File(s) Summary
Token Estimation Refactoring
apps/gateway/package.json, apps/gateway/src/chat/tools/tokenizer.ts, apps/gateway/src/chat/tools/estimate-tokens.ts, apps/gateway/src/chat/tools/estimate-tokens-from-content.ts, apps/gateway/src/chat/tools/types.ts
Removes gpt-tokenizer dependency and replaces with delegated calls to a new shared token estimation module. DEFAULT_TOKENIZER_MODEL constant removed; error handling and fallback logic simplified.
Shared Token Estimation Module
packages/shared/src/token-estimate.ts, packages/shared/src/token-estimate.spec.ts, packages/shared/src/index.ts
Introduces new exported functions estimateTokensFromText and estimateChatMessageTokens that perform length-based token estimation (length/4 heuristic) with proper handling for multimodal content; includes comprehensive test coverage.
Centralized Gateway Prompt Token Estimation
apps/gateway/src/chat/chat.ts, apps/gateway/src/lib/prompt-tokens.spec.ts
Centralizes routingPromptTokens computation from messages and tool size, eliminates per-auto-routing duplication, extends cost calculation to include image_config.image_quality. Updates tests to validate new length-based estimation behavior.
Cache-Support Routing
packages/actions/src/get-cheapest-from-available-providers.ts, packages/actions/src/models.spec.ts, apps/docs/content/features/routing.mdx
Adds prompt-caching awareness: ProviderSelectionOptions accepts promptTokens; scoring includes cacheSupported flag. Cache weight (20%) applied only when promptTokens ≥ 5000. New test scenarios validate cache-weighting behavior and override by cost.
Cost Calculation Updates
apps/gateway/src/lib/costs.ts
Extends calculateCosts signature with optional imageQuality parameter; removes gpt-tokenizer usage; updates image output token fallback logic to prefer totalOutputTokens when imageOutputTokensPerImage is missing.
Schema & API Updates
packages/db/src/schema.ts, apps/api/src/routes/logs.ts
Extends routingMetadata.providerScores JSON typing with optional cacheSupported boolean field; updates log schema validation.
UI Display Enhancements
apps/ui/src/app/dashboard/[orgId]/[projectId]/activity/[logId]/log-detail-client.tsx, packages/shared/src/components/log-card.tsx
Adds cache indicator rendering when score.cacheSupported is true; displays image quality metadata; improves base64 image detection and extraction heuristics with whitespace compaction.

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~60 minutes

Possibly related PRs

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 38.10% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title 'feat(routing): cache weight via length-based estimate' clearly and concisely summarizes the main change: adding cache-weight support to routing via a length-based token estimation approach instead of gpt-tokenizer.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch weight-cache-estimate

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share
Review rate limit: 7/8 reviews remaining, refill in 7 minutes and 30 seconds.

Comment @coderabbitai help to get the list of available commands and usage tips.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR reintroduces cache-aware routing for large prompts while removing gpt-tokenizer from the gateway hot path by switching token counting to a chars/4 heuristic, improving throughput at the cost of estimation precision.

Changes:

  • Add cache-support weighting to provider selection when estimated prompt tokens exceed a 5k threshold, and expose cacheSupported in routing score metadata.
  • Replace gpt-tokenizer usage across gateway routing/cost/usage estimation paths with length-based token estimates.
  • Remove gpt-tokenizer from apps/gateway dependencies and update tests/docs accordingly.

Reviewed changes

Copilot reviewed 10 out of 11 changed files in this pull request and generated 1 comment.

Show a summary per file
File Description
pnpm-lock.yaml Drops gpt-tokenizer and updates lockfile graph accordingly.
packages/actions/src/get-cheapest-from-available-providers.ts Adds cache-support scoring weight gated by estimated prompt tokens; surfaces cacheSupported in provider scores.
packages/actions/src/models.spec.ts Adds unit tests covering cache-support weighting behavior above/below threshold.
apps/gateway/src/chat/chat.ts Computes a single prompt-token estimate per request and plumbs it into routing; removes tokenizer calls in streaming/non-streaming usage estimation.
apps/gateway/src/chat/tools/tokenizer.ts Replaces chat token counting with a chars/4 length-based estimator.
apps/gateway/src/chat/tools/estimate-tokens.ts Switches completion-token fallback to length-based estimation.
apps/gateway/src/chat/tools/types.ts Removes tokenizer-specific constants/types tied to gpt-tokenizer.
apps/gateway/src/lib/costs.ts Replaces cost-time tokenization with length-based estimates for missing usage.
apps/gateway/src/lib/prompt-tokens.spec.ts Updates tests to reflect length-based estimation semantics (including empty-input behavior).
apps/gateway/package.json Removes gpt-tokenizer dependency from the gateway app.
apps/docs/content/features/routing.mdx Documents cache-support weighting behavior for large prompts.
Files not reviewed (1)
  • pnpm-lock.yaml: Language not supported
Comments suppressed due to low confidence (1)

apps/gateway/src/lib/costs.ts:193

  • calculateCosts now uses a length-based estimator that can return 0 for empty prompts/messages, but the early-return check if (!calculatedPromptTokens) treats 0 the same as null/undefined and returns all-null costs. This can incorrectly suppress costs for legitimately empty prompts (should be $0 with 0 tokens) and can also affect any caller that passes promptTokens = 0 intentionally. Consider switching these truthiness checks to explicit nullish checks (e.g., calculatedPromptTokens == null) and similarly using promptTokens == null / completionTokens == null in the estimation gate so 0 is treated as a real value.
	if ((!promptTokens || !completionTokens) && fullOutput) {
		// We're going to estimate at least some of the tokens
		isEstimated = true;
		// Calculate prompt tokens using a cheap length-based estimate.
		// Accuracy is intentionally traded for throughput so we never run
		// gpt-tokenizer on the gateway hot path.
		if (!promptTokens && fullOutput) {
			if (fullOutput.messages) {
				calculatedPromptTokens = encodeChatMessages(fullOutput.messages);
			} else if (fullOutput.prompt) {
				calculatedPromptTokens = estimateTokensFromContent(
					JSON.stringify(fullOutput.prompt),
				);
			}
		}

		// Calculate completion tokens
		if (!completionTokens && fullOutput) {
			let completionText = "";

			// Include main completion content
			if (fullOutput.completion) {
				completionText += fullOutput.completion;
			}

			// Include tool results if available
			if (fullOutput.toolResults && Array.isArray(fullOutput.toolResults)) {
				for (const toolResult of fullOutput.toolResults) {
					if (toolResult.function?.name) {
						completionText += toolResult.function.name;
					}
					if (toolResult.function?.arguments) {
						completionText += JSON.stringify(toolResult.function.arguments);
					}
				}
			}

			if (completionText) {
				calculatedCompletionTokens = estimateTokensFromContent(completionText);
			}
		}
	}

	// If we don't have prompt tokens, we can't calculate any costs
	if (!calculatedPromptTokens) {
		return {
			inputCost: null,

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.


**Cache Support for Large Prompts**:

When the estimated prompt is at least 5,000 tokens, an additional 20% weight is added to the score for whether each provider supports prompt caching (advertised via a cached input price). Providers that support caching score better than ones that do not, since caching can substantially reduce the cost of large or repeated prompts. Below the 5k threshold, this weight is dropped entirely — caching has little impact on small prompts, so cache support is ignored. The selected provider's cache support is exposed as `cacheSupported` on the routing metadata.

Copilot AI Apr 27, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The docs say cache support is exposed as cacheSupported "on the routing metadata", but in code it’s added on each entry of metadata.providerScores (and not as a top-level RoutingMetadata.cacheSupported). Please clarify the exact field path (e.g. routingMetadata.providerScores[].cacheSupported) to avoid consumers looking for a non-existent top-level property.

Suggested change
When the estimated prompt is at least 5,000 tokens, an additional 20% weight is added to the score for whether each provider supports prompt caching (advertised via a cached input price). Providers that support caching score better than ones that do not, since caching can substantially reduce the cost of large or repeated prompts. Below the 5k threshold, this weight is dropped entirely — caching has little impact on small prompts, so cache support is ignored. The selected provider's cache support is exposed as `cacheSupported` on the routing metadata.
When the estimated prompt is at least 5,000 tokens, an additional 20% weight is added to the score for whether each provider supports prompt caching (advertised via a cached input price). Providers that support caching score better than ones that do not, since caching can substantially reduce the cost of large or repeated prompts. Below the 5k threshold, this weight is dropped entirely — caching has little impact on small prompts, so cache support is ignored. Cache support is exposed on each provider score entry in the routing metadata as `routingMetadata.providerScores[].cacheSupported`.

Copilot uses AI. Check for mistakes.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 5

🧹 Nitpick comments (4)
apps/gateway/src/chat/tools/tokenizer.ts (1)

17-28: Optional: consolidate the chars/4 heuristic in one place.

estimateTokensFromLength here and estimateTokensFromContent in estimate-tokens-from-content.ts implement the same Math.max(1, Math.round(length / 4)) logic with the same CHARS_PER_TOKEN factor (currently hardcoded as / 4 there). If the heuristic ever changes (e.g. different ratio, BPE-aware adjustment), both call sites must be kept in sync. Consider having estimateTokensFromContent delegate to estimateTokensFromLength(content.length) so the constant lives in a single module.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@apps/gateway/src/chat/tools/tokenizer.ts` around lines 17 - 28, Consolidate
the chars-per-token heuristic by having estimateTokensFromContent delegate to
estimateTokensFromLength so the CHARS_PER_TOKEN constant is maintained in one
place: remove the duplicate hardcoded `/ 4` logic from estimateTokensFromContent
and replace its implementation with a call to
estimateTokensFromLength(content.length), ensuring estimateTokensFromLength (and
its CHARS_PER_TOKEN constant) remain the single source of truth for the
heuristic.
packages/actions/src/get-cheapest-from-available-providers.ts (3)

586-597: selectByPriceOnly fallback doesn't surface cacheSupported.

When metricsMap is missing/empty (e.g., metrics layer down), routing falls through to selectByPriceOnly, and the resulting providerScores entries omit the new cacheSupported field even for large prompts. This isn't a correctness bug for selection (price-only intentionally ignores cache), but it makes the routing-metadata schema inconsistent and hides whether the selected provider supports caching from downstream consumers (logs, dashboard, the cacheSupported signal documented in routing.mdx).

Consider populating cacheSupported here as well so the field is always present:

♻️ Proposed fix
 	for (const provider of stableProviders) {
 		const providerInfo = modelWithPricing.providers.find(
 			(p) =>
 				p.providerId === provider.providerId && p.region === provider.region,
 		);
 		const totalPrice = getProviderSelectionPrice(providerInfo, videoPricing);

 		// Apply provider priority: lower priority = effectively higher price
 		const providerDef = getProviderDefinition(provider.providerId);
 		const priority = providerDef?.priority ?? 1;
 		const effectivePrice = priority > 0 ? totalPrice / priority : totalPrice;

 		providerPrices.push({
 			providerId: provider.providerId,
 			region: provider.region,
 			price: totalPrice,
 			effectivePrice,
 			priority,
+			cacheSupported: providerSupportsCaching(providerInfo),
 		});

(plus a corresponding field on the local array type and on the mapped providerScores entry).

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@packages/actions/src/get-cheapest-from-available-providers.ts` around lines
586 - 597, The RoutingMetadata produced in
get-cheapest-from-available-providers.ts (within the select-by-price-only
fallback) omits the cacheSupported flag on providerScores entries; update the
local providerPrices entry type and the mapped providerScores construction (the
array built from providerPrices.map(...) that feeds RoutingMetadata) to include
cacheSupported (set appropriately from each provider's capabilities, e.g.,
provider.cacheSupported or a boolean derived from provider info) and ensure the
RoutingMetadata type/shape includes cacheSupported so downstream consumers
always see the field even when metricsMap is missing.

165-199: providerSupportsCaching triggers off any cachedInputPrice, including a 0 placeholder.

The check providerInfo.cachedInputPrice !== undefined returns true whenever the field exists, even if it's 0 or equal to inputPrice. In practice cachedInputPrice: 0 would mean "free cached input", which is still cache support — fine. But if a provider mapping ever sets cachedInputPrice equal to or higher than inputPrice as a placeholder (or sets it on a region but not at the top level for matching), it will get the cache-routing bonus without any actual cost benefit. Worth tightening the check to cachedInputPrice strictly less than inputPrice (and similarly for tiers/regions) if you want the routing weight to track an actual savings rather than the mere presence of the field.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@packages/actions/src/get-cheapest-from-available-providers.ts` around lines
165 - 199, The function providerSupportsCaching currently treats any present
cachedInputPrice (including 0 or placeholders >= inputPrice) as support; update
providerSupportsCaching to accept an inputPrice parameter and change all checks
so they only return true when cachedInputPrice is strictly less than inputPrice
(i.e., providerInfo.cachedInputPrice < inputPrice, tier.cachedInputPrice <
inputPrice, region.cachedInputPrice < inputPrice, and region.pricingTiers items
similarly) so the routing bonus only applies when there is an actual cost
savings; keep the same nested checks and early-return logic but compare values
rather than just presence.

304-307: Threshold compares estimated prompt tokens to a real-token threshold.

promptTokens flowing in here comes from the chars/4 heuristic in encodeChatMessages (per PR objectives). For English prose, chars/4 ≈ true gpt-tokenizer count, so the 5k threshold is roughly preserved. For code-heavy or non-Latin prompts (CJK, etc.), the heuristic systematically under-counts tokens, so the cache-support weight will under-trigger on exactly the prompts where caching would help most. This is a known trade-off mentioned in the PR description, but worth a brief code comment so future readers understand the threshold is approximate, and consider biasing the threshold downward (e.g. 4000) to compensate.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@packages/actions/src/get-cheapest-from-available-providers.ts` around lines
304 - 307, The comparison using promptTokens (from encodeChatMessages' chars/4
heuristic) to CACHE_PROMPT_TOKEN_THRESHOLD can under-count tokens for code-heavy
or non-Latin input, so update the comment near the cacheSupportRelevant
calculation to note the heuristic is approximate, may under-count for CJK/code,
and recommend a conservative bias (e.g. lower the effective threshold to ~4000)
or revisit CACHE_PROMPT_TOKEN_THRESHOLD; reference encodeChatMessages and
CACHE_PROMPT_TOKEN_THRESHOLD and the cacheSupportRelevant variable so
maintainers know where to adjust the threshold or heuristic.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@apps/gateway/src/chat/chat.ts`:
- Around line 1090-1100: routingPromptTokens is computed before
applyRedactions() can change messages, causing routing (cache-weight threshold
and auto-routing) to be based on pre-redaction content; move or recompute the
token estimate after redactions so selection uses the final payload.
Specifically, after applyRedactions() finishes (the code path that may mutate
messages/tools), call encodeChatMessages(messages) and re-measure tools
(JSON.stringify(tools).length / 4) to set routingPromptTokens, replacing the
earlier pre-redaction calculation in the same scope where routing decisions are
made.
- Around line 6887-6890: calculatedReasoningTokens is computed as a fallback but
never propagated; update downstream logic to use calculatedReasoningTokens
wherever reasoningTokens is currently referenced so the fallback is honoured:
pass calculatedReasoningTokens into calculateCosts(...), include it in
total-token math and any usage chunk computations (e.g. usageChunks/usage
totals), and ensure transformResponseToOpenai(...) receives the
calculatedReasoningTokens for response metadata. Locate uses of reasoningTokens
in the same scope and replace or augment them to prefer
calculatedReasoningTokens when reasoningTokens is falsy so billing and metadata
reflect the estimated reasoning usage (also apply the same change near the other
occurrence around the 8940–8942 region).

In `@apps/gateway/src/chat/tools/estimate-tokens.ts`:
- Around line 26-30: The code is unnecessarily calling JSON.stringify on content
before token estimation, which wraps the string in quotes and inflates the token
count; in the branch that checks if (!completionTokens && content) pass content
directly to estimateTokensFromContent (replace JSON.stringify(content) with
content) so calculatedCompletionTokens is computed from the raw string, matching
the behavior used elsewhere (see costs.ts usage) and keeping estimates
consistent.

In `@apps/gateway/src/lib/costs.ts`:
- Around line 156-160: The token estimation inflates prompt length because
fullOutput.prompt (typed as string) is being wrapped with JSON.stringify before
calling estimateTokensFromContent; instead, call estimateTokensFromContent with
the raw fullOutput.prompt value (no JSON.stringify) so it matches the raw-string
path used for completionText and yields consistent token estimates; update the
branch in costs.ts where fullOutput.prompt is handled to pass fullOutput.prompt
directly to estimateTokensFromContent.

In `@packages/actions/src/models.spec.ts`:
- Around line 948-963: The test relies on implicit provider priority defaults
(openai/deepseek) which makes it brittle; update the test in
packages/actions/src/models.spec.ts that calls getCheapestFromAvailableProviders
with cacheTestModel to explicitly assert cache-related behavior instead of equal
scores—e.g., check result.metadata.providerScores entries for a cacheSupported
(or equivalent) flag for each provider and assert that cacheSupported is
true/false as expected for cache-supporting providers, or run two calls (one
with providers modified to disable cache support) and assert the score delta
between runs; reference getCheapestFromAvailableProviders, cacheTestModel, and
the providerScores/openai and deepseek entries to locate and change the
assertions.

---

Nitpick comments:
In `@apps/gateway/src/chat/tools/tokenizer.ts`:
- Around line 17-28: Consolidate the chars-per-token heuristic by having
estimateTokensFromContent delegate to estimateTokensFromLength so the
CHARS_PER_TOKEN constant is maintained in one place: remove the duplicate
hardcoded `/ 4` logic from estimateTokensFromContent and replace its
implementation with a call to estimateTokensFromLength(content.length), ensuring
estimateTokensFromLength (and its CHARS_PER_TOKEN constant) remain the single
source of truth for the heuristic.

In `@packages/actions/src/get-cheapest-from-available-providers.ts`:
- Around line 586-597: The RoutingMetadata produced in
get-cheapest-from-available-providers.ts (within the select-by-price-only
fallback) omits the cacheSupported flag on providerScores entries; update the
local providerPrices entry type and the mapped providerScores construction (the
array built from providerPrices.map(...) that feeds RoutingMetadata) to include
cacheSupported (set appropriately from each provider's capabilities, e.g.,
provider.cacheSupported or a boolean derived from provider info) and ensure the
RoutingMetadata type/shape includes cacheSupported so downstream consumers
always see the field even when metricsMap is missing.
- Around line 165-199: The function providerSupportsCaching currently treats any
present cachedInputPrice (including 0 or placeholders >= inputPrice) as support;
update providerSupportsCaching to accept an inputPrice parameter and change all
checks so they only return true when cachedInputPrice is strictly less than
inputPrice (i.e., providerInfo.cachedInputPrice < inputPrice,
tier.cachedInputPrice < inputPrice, region.cachedInputPrice < inputPrice, and
region.pricingTiers items similarly) so the routing bonus only applies when
there is an actual cost savings; keep the same nested checks and early-return
logic but compare values rather than just presence.
- Around line 304-307: The comparison using promptTokens (from
encodeChatMessages' chars/4 heuristic) to CACHE_PROMPT_TOKEN_THRESHOLD can
under-count tokens for code-heavy or non-Latin input, so update the comment near
the cacheSupportRelevant calculation to note the heuristic is approximate, may
under-count for CJK/code, and recommend a conservative bias (e.g. lower the
effective threshold to ~4000) or revisit CACHE_PROMPT_TOKEN_THRESHOLD; reference
encodeChatMessages and CACHE_PROMPT_TOKEN_THRESHOLD and the cacheSupportRelevant
variable so maintainers know where to adjust the threshold or heuristic.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro

Run ID: 26804f55-5c79-4f0c-b647-d3d353443e20

📥 Commits

Reviewing files that changed from the base of the PR and between c94491c and e9ac8db.

⛔ Files ignored due to path filters (1)
  • pnpm-lock.yaml is excluded by !**/pnpm-lock.yaml
📒 Files selected for processing (10)
  • apps/docs/content/features/routing.mdx
  • apps/gateway/package.json
  • apps/gateway/src/chat/chat.ts
  • apps/gateway/src/chat/tools/estimate-tokens.ts
  • apps/gateway/src/chat/tools/tokenizer.ts
  • apps/gateway/src/chat/tools/types.ts
  • apps/gateway/src/lib/costs.ts
  • apps/gateway/src/lib/prompt-tokens.spec.ts
  • packages/actions/src/get-cheapest-from-available-providers.ts
  • packages/actions/src/models.spec.ts
💤 Files with no reviewable changes (2)
  • apps/gateway/package.json
  • apps/gateway/src/chat/tools/types.ts

Comment thread apps/gateway/src/chat/chat.ts
Comment on lines 6887 to +6890
let calculatedReasoningTokens = reasoningTokens;
if (!reasoningTokens && fullReasoningContent) {
try {
calculatedReasoningTokens = encode(fullReasoningContent).length;
} catch (error) {
// Fallback to simple estimation if encoding fails
logger.error(
"Failed to encode reasoning text in streaming",
error instanceof Error ? error : new Error(String(error)),
);
calculatedReasoningTokens =
estimateTokensFromContent(fullReasoningContent);
}
calculatedReasoningTokens =
estimateTokensFromContent(fullReasoningContent);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

Thread the estimated reasoning tokens through the downstream accounting.

calculatedReasoningTokens is populated here, but the later calculateCosts(...), total-token math, usage chunks, and transformResponseToOpenai(...) still use raw reasoningTokens. When upstream omits reasoning usage, the new fallback only affects some log fields while response metadata and billing-related totals still undercount reasoning.

Also applies to: 8940-8942

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@apps/gateway/src/chat/chat.ts` around lines 6887 - 6890,
calculatedReasoningTokens is computed as a fallback but never propagated; update
downstream logic to use calculatedReasoningTokens wherever reasoningTokens is
currently referenced so the fallback is honoured: pass calculatedReasoningTokens
into calculateCosts(...), include it in total-token math and any usage chunk
computations (e.g. usageChunks/usage totals), and ensure
transformResponseToOpenai(...) receives the calculatedReasoningTokens for
response metadata. Locate uses of reasoningTokens in the same scope and replace
or augment them to prefer calculatedReasoningTokens when reasoningTokens is
falsy so billing and metadata reflect the estimated reasoning usage (also apply
the same change near the other occurrence around the 8940–8942 region).

Comment thread apps/gateway/src/chat/tools/estimate-tokens.ts
Comment on lines 156 to 160
} else if (fullOutput.prompt) {
// For text prompt
try {
calculatedPromptTokens = encode(
JSON.stringify(fullOutput.prompt),
).length;
} catch (error) {
// If encoding fails, leave as null
logger.error(`Failed to encode prompt text: ${error}`);
}
calculatedPromptTokens = estimateTokensFromContent(
JSON.stringify(fullOutput.prompt),
);
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Drop the JSON.stringify on fullOutput.prompt.

fullOutput.prompt is typed as string (line 93). JSON.stringify wraps it in quotes and escapes specials, inflating the chars/4 estimate by at least 2 characters compared to the consistent raw-string path used at line 185 for completionText. Pass the prompt directly.

♻️ Proposed fix
 			} else if (fullOutput.prompt) {
-				calculatedPromptTokens = estimateTokensFromContent(
-					JSON.stringify(fullOutput.prompt),
-				);
+				calculatedPromptTokens = estimateTokensFromContent(fullOutput.prompt);
 			}
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@apps/gateway/src/lib/costs.ts` around lines 156 - 160, The token estimation
inflates prompt length because fullOutput.prompt (typed as string) is being
wrapped with JSON.stringify before calling estimateTokensFromContent; instead,
call estimateTokensFromContent with the raw fullOutput.prompt value (no
JSON.stringify) so it matches the raw-string path used for completionText and
yields consistent token estimates; update the branch in costs.ts where
fullOutput.prompt is handled to pass fullOutput.prompt directly to
estimateTokensFromContent.

Comment on lines +948 to +963
it("does not factor cache support when prompt is below the threshold", () => {
const result = getCheapestFromAvailableProviders(
cacheTestModel.providers,
cacheTestModel,
{ metricsMap: equalMetrics, promptTokens: 1000 },
);

const openai = result?.metadata.providerScores.find(
(p) => p.providerId === "openai",
);
const deepseek = result?.metadata.providerScores.find(
(p) => p.providerId === "deepseek",
);

expect(openai?.score).toBe(deepseek?.score);
});

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# Verify openai and deepseek share the same priority in the provider catalog
rg -nP --type=ts -C2 '(providerId|id)\s*:\s*"(openai|deepseek)"\b' packages/models/src
rg -nP --type=ts -C5 '\bpriority\s*:' packages/models/src | rg -nP -B5 '"(openai|deepseek)"'

Repository: theopenco/llmgateway

Length of output: 46


🏁 Script executed:

# First, explore the repo structure to understand where provider definitions are
find . -type f -name "*.ts" | grep -E "(provider|model)" | head -20

Repository: theopenco/llmgateway

Length of output: 1111


🏁 Script executed:

# Look at the test file itself to understand cacheTestModel and the function being tested
find . -type f -name "models.spec.ts" | head -5

Repository: theopenco/llmgateway

Length of output: 143


🏁 Script executed:

# Find getCheapestFromAvailableProviders function location
rg -l "getCheapestFromAvailableProviders" --type=ts | head -10

Repository: theopenco/llmgateway

Length of output: 226


🏁 Script executed:

# Read the test file around lines 948-963 to see cacheTestModel setup and test context
sed -n '900,1000p' packages/actions/src/models.spec.ts

Repository: theopenco/llmgateway

Length of output: 2684


🏁 Script executed:

# Find the getCheapestFromAvailableProviders implementation and look at priorityPenalty logic
sed -n '490,520p' packages/actions/src/get-cheapest-from-available-providers.ts

Repository: theopenco/llmgateway

Length of output: 1315


🏁 Script executed:

# Search for provider definitions/priorities
rg -n "priority" packages/actions/src/models.spec.ts | head -20

Repository: theopenco/llmgateway

Length of output: 376


🏁 Script executed:

# Search for provider registry or definitions
rg -n "getProviderDefinition\|provider.*priority" --type=ts | head -20

Repository: theopenco/llmgateway

Length of output: 46


🏁 Script executed:

# Find where provider definitions are stored and their priorities
rg -n "getProviderDefinition\|export.*provider.*definition" --type=ts -A5 | head -40

Repository: theopenco/llmgateway

Length of output: 46


🏁 Script executed:

# Look for provider definitions or a provider registry
find . -type f -name "*.ts" | xargs grep -l "getProviderDefinition" | head -10

Repository: theopenco/llmgateway

Length of output: 400


🏁 Script executed:

# Check if there's a provider catalog or registry with priority values
rg -n "openai.*priority\|deepseek.*priority" --type=ts | head -20

Repository: theopenco/llmgateway

Length of output: 46


🏁 Script executed:

# Check the provider definitions in packages/models/src
cat packages/models/src/providers.ts | head -150

Repository: theopenco/llmgateway

Length of output: 4078


🏁 Script executed:

# Look for getProviderDefinition function implementation
rg -n "getProviderDefinition" packages/models/src --type=ts -A10 | head -50

Repository: theopenco/llmgateway

Length of output: 2090


🏁 Script executed:

# Check the provider.ts file for priority definitions
cat packages/models/src/provider.ts

Repository: theopenco/llmgateway

Length of output: 3907


🏁 Script executed:

# Search for openai and deepseek provider definitions with priority info
rg -n "id.*openai\|id.*deepseek" packages/models/src/providers.ts -A15 | head -60

Repository: theopenco/llmgateway

Length of output: 46


🏁 Script executed:

# Get all provider definitions from providers.ts to see openai and deepseek entries
cat packages/models/src/providers.ts | grep -A10 "id.*:.*openai\|id.*:.*deepseek"

Repository: theopenco/llmgateway

Length of output: 611


🏁 Script executed:

# Alternative: extract the entire providers array
sed -n '/export const providers = \[/,/^\]/p' packages/models/src/providers.ts | head -300

Repository: theopenco/llmgateway

Length of output: 6640


🏁 Script executed:

# Verify the exact priority values for openai and deepseek by extracting their full definitions
sed -n '/id.*:.*"openai"/,/^[[:space:]]*},/p' packages/models/src/providers.ts
sed -n '/id.*:.*"deepseek"/,/^[[:space:]]*},/p' packages/models/src/providers.ts

Repository: theopenco/llmgateway

Length of output: 516


🏁 Script executed:

# Get openai provider definition with more context
rg -n "id.*:.*\"openai\"" packages/models/src/providers.ts -A20

Repository: theopenco/llmgateway

Length of output: 668


🏁 Script executed:

# Get deepseek provider definition with more context
rg -n "id.*:.*\"deepseek\"" packages/models/src/providers.ts -A20

Repository: theopenco/llmgateway

Length of output: 604


The test assertion is correct as written. Both openai and deepseek provider definitions in packages/models/src/providers.ts (lines 72–86 and 249–263 respectively) lack an explicit priority field, so both default to priority = 1. With equal metrics and pricing, their final scores will indeed be equal.

However, the test could be more resilient to future provider catalog changes. Consider explicitly asserting on cache-logic behavior alone—e.g., verify that cacheSupported is reported correctly or compare score deltas between a run with cache-supporting providers and one without, rather than relying on priority field defaults to remain unchanged.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@packages/actions/src/models.spec.ts` around lines 948 - 963, The test relies
on implicit provider priority defaults (openai/deepseek) which makes it brittle;
update the test in packages/actions/src/models.spec.ts that calls
getCheapestFromAvailableProviders with cacheTestModel to explicitly assert
cache-related behavior instead of equal scores—e.g., check
result.metadata.providerScores entries for a cacheSupported (or equivalent) flag
for each provider and assert that cacheSupported is true/false as expected for
cache-supporting providers, or run two calls (one with providers modified to
disable cache support) and assert the score delta between runs; reference
getCheapestFromAvailableProviders, cacheTestModel, and the providerScores/openai
and deepseek entries to locate and change the assertions.

steebchen and others added 8 commits April 28, 2026 18:26
content is already a non-empty string at the call site; wrapping it
in JSON.stringify just adds ~2 chars of quote/escape overhead and
diverged from the equivalent path in costs.ts.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Surfaces the new cacheSupported routing-score field on the log
detail card alongside uptime/throughput/latency/price.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
## Summary
- Add `canopywave` provider entry for the `deepseek-v4-flash` model
- Prices: $0.14/Mt input, $0.28/Mt output, $0.03/Mt cached input; 1M
context; 30% discount
- Capabilities aligned with canopywave's listed features
(function-calling, structured-outputs, no reasoning)

## Test plan
- [x] `TEST_MODELS="canopywave/deepseek-v4-flash" pnpm test:e2e` for
basic, streaming, tool calls, JSON, reasoning, and prompt-caching suites
— all pass individually
- [x] `pnpm build`
- [x] `pnpm format`

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Introduced canopywave as a new provider option for the
deepseek-v4-flash model, featuring custom pricing and capability
configurations.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Luca Steeb (bot) <contact@luca-steeb.com>
Adds the cacheSupported badge to the shared log-card component
used in log lists, mirroring the detail page.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 6

♻️ Duplicate comments (2)
apps/gateway/src/chat/chat.ts (2)

1090-1100: ⚠️ Potential issue | 🟠 Major

Recompute routingPromptTokens after guardrail redactions.

routingPromptTokens is still captured before applyRedactions() can rewrite messages (Line 1403), so the 5k cache-weight threshold and the 10k auto-routing cutoff can be evaluated against a different prompt than the one we actually send upstream.

Suggested fix
-	let routingPromptTokens = 0;
-	if (messages && messages.length > 0) {
-		routingPromptTokens = encodeChatMessages(messages);
-	}
-	if (tools && tools.length > 0) {
-		routingPromptTokens += Math.round(JSON.stringify(tools).length / 4);
-	}
+	const calculateRoutingPromptTokens = () => {
+		let promptTokens = messages.length > 0 ? encodeChatMessages(messages) : 0;
+		if (tools && tools.length > 0) {
+			promptTokens += Math.round(JSON.stringify(tools).length / 4);
+		}
+		return promptTokens;
+	};
+	let routingPromptTokens = calculateRoutingPromptTokens();
...
 		if (guardrailResult.redactions.length > 0) {
 			messages = applyRedactions(
 				messages as Parameters<typeof applyRedactions>[0],
 				guardrailResult.redactions,
 			) as typeof messages;
+			routingPromptTokens = calculateRoutingPromptTokens();
 		}
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@apps/gateway/src/chat/chat.ts` around lines 1090 - 1100, routingPromptTokens
is computed from messages and tools before applyRedactions() mutates messages,
so downstream routing thresholds (cache-weight and auto-routing) may use stale
token counts; update the code to recompute routingPromptTokens after
applyRedactions() runs (or call encodeChatMessages(messages) again on the
redacted messages and re-add the tools heuristic), ensuring the token estimate
used for decisions reflects the actual messages sent upstream and still uses
encodeChatMessages and the existing tools length/4 heuristic.

6889-6892: ⚠️ Potential issue | 🟠 Major

Use calculatedReasoningTokens in downstream usage and cost accounting.

These fallback blocks still only populate a local variable. Line 7141, Line 7200, Line 7419, Line 8961, and Lines 9004-9007 continue to read reasoningTokens, so reasoning_tokens, total_tokens, and cost metadata are still undercounted whenever the provider omits reasoning usage.

Also applies to: 8946-8948

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@apps/gateway/src/chat/chat.ts` around lines 6889 - 6892, The fallback logic
that sets calculatedReasoningTokens (via estimateTokensFromContent) is correct
but never used downstream; update all downstream usages that currently read
reasoningTokens to use calculatedReasoningTokens when reasoningTokens is
falsy—specifically wherever you compute reasoning_tokens, total_tokens, and cost
metadata ensure you default to calculatedReasoningTokens (e.g., use
(reasoningTokens ?? calculatedReasoningTokens) or similar). Make the change in
the code paths that assemble usage/cost objects so reasoning_tokens reflects the
fallback value and total_tokens/cost calculations include it, leaving existing
behavior unchanged when reasoningTokens is already present.
🧹 Nitpick comments (3)
apps/playground/src/components/playground/chat-ui.tsx (1)

457-468: Prefer shared image config over local GPT-image option duplication.

isGptImage/size/quality logic is now split between this file and apps/playground/src/lib/image-gen.ts; consolidating on the shared config would reduce drift risk.

Also applies to: 486-486

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@apps/playground/src/components/playground/chat-ui.tsx` around lines 457 -
468, The duplicate image-model logic in chat-ui.tsx (variables isGptImage and
usesPixelDimensions derived from selectedModel and used for size/quality UI)
should be removed and replaced with a call to the shared image config in
apps/playground/src/lib/image-gen.ts; locate the isGptImage and
usesPixelDimensions definitions in chat-ui.tsx and replace them with imports and
usage of the exported helpers or config (e.g., functions or flags from
image-gen.ts that determine pixel-based dimensions and quality support) so the
UI consumes the single source of truth for image model detection and
size/quality rules.
apps/gateway/src/chat/tools/transform-response-to-openai.ts (1)

463-468: Consider passing image tokens through in existing-response branches for consistency.

The applyExtendedUsageFields calls in branches that modify existing responses (e.g., lines 463-468, 749-754, and similar patterns for alibaba, bytedance, xai, zai, and the default case) don't pass imageInputTokens/imageOutputTokens.

Currently this is fine because image tokens are only relevant for OpenAI gpt-image models, which take the buildUsageObject path. However, for consistency and future-proofing (e.g., if other providers add image token reporting), you could pass them through:

 applyExtendedUsageFields(transformedResponse.usage, {
   costs,
   cachedTokens,
   cacheCreationTokens,
   reasoningTokens,
+  imageInputTokens,
+  imageOutputTokens,
 });

Also applies to: 749-754

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@apps/gateway/src/chat/tools/transform-response-to-openai.ts` around lines 463
- 468, The applyExtendedUsageFields call in branches that update an existing
response (e.g., where transformedResponse.usage is passed) omits
imageInputTokens and imageOutputTokens, so add these tokens to all such calls
for consistency and future-proofing; update each invocation of
applyExtendedUsageFields (including the OpenAI-existing-response branch shown
and the analogous alibaba, bytedance, xai, zai, and default branches) to include
imageInputTokens and imageOutputTokens alongside costs, cachedTokens,
cacheCreationTokens, and reasoningTokens, ensuring the values come from the same
buildUsageObject or upstream variables used when image tokens are available.
apps/playground/src/components/playground/chat-page-client.tsx (1)

1296-1398: Consider extracting duplicated logic into shared hooks.

The model configuration logic (isGptImage detection, usesPixelDimensions checks, sendMessageWithHeaders callback, and the useEffect reset logic) is nearly identical between the main component and ExtraChatPanel.

Consider extracting this into:

  1. A custom hook like useImageConfig(selectedModel) for state and reset logic
  2. A helper function for building the imageConfig object

This would reduce maintenance burden when adding new model types or changing behavior.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@apps/playground/src/components/playground/chat-page-client.tsx` around lines
1296 - 1398, The duplicated model/image-detection and reset logic (isGptImage,
usesPixelDimensions, the useEffect that resets sizes/quality, and the
sendMessageWithHeaders imageConfig construction) should be extracted so both
this component and ExtraChatPanel share it; create a custom hook
useImageConfig(selectedModel) that encapsulates the useEffect reset behavior and
exports computed flags (isGptImage, usesPixelDimensions, supportsImageGen,
supportsImages, default sizes/qualities and setters), and refactor the
imageConfig building into a helper buildImageConfig({isGptImage,
usesPixelDimensions, alibabaImageSize, imageSize, imageAspectRatio,
imageQuality, imageCount, useImageGen}) which returns the image_config object
used in sendMessageWithHeaders, then replace the inline logic in
sendMessageWithHeaders and the component’s useEffect with calls to the new hook
and helper so both components consume the shared behavior.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@apps/gateway/src/lib/costs.ts`:
- Around line 151-157: The prompt token estimation path overcounts multimodal
content because encodeChatMessages serializes arrays (via JSON.stringify) and
treats image blocks/data URLs as text; update the branch that handles
fullOutput.messages before calling encodeChatMessages to strip or replace
non-text message parts (e.g., objects with a type like "image", items with data:
URLs, Buffers, or known multimodal markers) with empty strings or a placeholder
text-only representation so encodeChatMessages only sees textual content; refer
to encodeChatMessages and the fullOutput.messages handling to locate the change
and ensure you do the same sanitization logic used by the tokenizer for image
detection so inputCost and promptTokens are not double-counted.

In `@apps/playground/src/components/playground/image-page-client.tsx`:
- Around line 204-232: The reset and request payload logic currently derives
config from selectedModels[0]; instead, call getModelImageConfig(modelId) for
each model when in comparison mode so each model's reset and imageConfigBody is
built from its own config (e.g., inside the loop that sends requests over
selectedModels), and only include fields like image_quality or pixel-dimension
settings when that model's config.supportsQuality or config.usesPixelDimensions
respectively; update the useEffect reset logic and the imageConfigBody
construction to compute config per model (referencing selectedModels,
getModelImageConfig, imageConfigBody, and the
imageQuality/imageSize/alibabaImageSize setters) to avoid leaking unsupported
fields to secondary models.

In `@apps/ui/src/components/models/model-card.tsx`:
- Around line 114-118: hasEstimatedImageCost currently returns true when
imageOutputTokensByResolution is an empty object; change the guard in
hasEstimatedImageCost to require both mapping.imageOutputPrice truthy and
mapping.imageOutputTokensByResolution to have at least one key (e.g.,
Object.keys(mapping.imageOutputTokensByResolution || {}).length > 0). Update the
same pattern wherever the same logic appears (the other occurrences that check
imageOutputPrice and imageOutputTokensByResolution) so empty maps do not count
as providing an estimated image cost.

In `@packages/actions/src/prepare-request-body.spec.ts`:
- Around line 207-235: Remove the unnecessary "as any" test casts: update the
three occurrences where requestBody is declared (calls to
prepareOpenAIImageRequest in the image tests) to use the actual return type
instead of "as any" (e.g., const requestBody = await
prepareOpenAIImageRequest({...})); this keeps type safety and lets existing
expect assertions work—adjust the variable declarations for the tests "should
not derive size from aspect_ratio", "should drop unsupported quality values",
and the earlier image-size test to drop the "as any" suffix.

In `@packages/actions/src/prepare-request-body.ts`:
- Around line 685-695: The code currently ignores image_config.aspect_ratio when
image_size is not set, causing silent fallback; update the image request
construction logic (involving image_config, openaiImageRequest, image_size,
aspect_ratio, normalizeImageQuality) to either map supported aspect_ratio values
to concrete OpenAI size strings or reject/throw a clear error when
image_config.aspect_ratio is present but image_config.image_size is not;
validate aspect_ratio against an explicit whitelist of supported ratios and
ensure the rejection occurs before building openaiImageRequest so callers
receive a fast, descriptive failure instead of silently using OpenAI defaults.

In `@scripts/image-edit.sh`:
- Around line 20-21: Validate the N variable immediately after it is set and
before any payload construction: check that N contains only digits (and
optionally is >=1) and exit with a clear error message if not; update the script
near where N is defined (reference variable N and TIMESTAMP/RESPONSE_FILE usage)
and also add the same validation before the payload-building section referenced
around the later payload block (the code near line ~84) so the script fails fast
with a helpful message instead of letting Python raise a traceback.

---

Duplicate comments:
In `@apps/gateway/src/chat/chat.ts`:
- Around line 1090-1100: routingPromptTokens is computed from messages and tools
before applyRedactions() mutates messages, so downstream routing thresholds
(cache-weight and auto-routing) may use stale token counts; update the code to
recompute routingPromptTokens after applyRedactions() runs (or call
encodeChatMessages(messages) again on the redacted messages and re-add the tools
heuristic), ensuring the token estimate used for decisions reflects the actual
messages sent upstream and still uses encodeChatMessages and the existing tools
length/4 heuristic.
- Around line 6889-6892: The fallback logic that sets calculatedReasoningTokens
(via estimateTokensFromContent) is correct but never used downstream; update all
downstream usages that currently read reasoningTokens to use
calculatedReasoningTokens when reasoningTokens is falsy—specifically wherever
you compute reasoning_tokens, total_tokens, and cost metadata ensure you default
to calculatedReasoningTokens (e.g., use (reasoningTokens ??
calculatedReasoningTokens) or similar). Make the change in the code paths that
assemble usage/cost objects so reasoning_tokens reflects the fallback value and
total_tokens/cost calculations include it, leaving existing behavior unchanged
when reasoningTokens is already present.

---

Nitpick comments:
In `@apps/gateway/src/chat/tools/transform-response-to-openai.ts`:
- Around line 463-468: The applyExtendedUsageFields call in branches that update
an existing response (e.g., where transformedResponse.usage is passed) omits
imageInputTokens and imageOutputTokens, so add these tokens to all such calls
for consistency and future-proofing; update each invocation of
applyExtendedUsageFields (including the OpenAI-existing-response branch shown
and the analogous alibaba, bytedance, xai, zai, and default branches) to include
imageInputTokens and imageOutputTokens alongside costs, cachedTokens,
cacheCreationTokens, and reasoningTokens, ensuring the values come from the same
buildUsageObject or upstream variables used when image tokens are available.

In `@apps/playground/src/components/playground/chat-page-client.tsx`:
- Around line 1296-1398: The duplicated model/image-detection and reset logic
(isGptImage, usesPixelDimensions, the useEffect that resets sizes/quality, and
the sendMessageWithHeaders imageConfig construction) should be extracted so both
this component and ExtraChatPanel share it; create a custom hook
useImageConfig(selectedModel) that encapsulates the useEffect reset behavior and
exports computed flags (isGptImage, usesPixelDimensions, supportsImageGen,
supportsImages, default sizes/qualities and setters), and refactor the
imageConfig building into a helper buildImageConfig({isGptImage,
usesPixelDimensions, alibabaImageSize, imageSize, imageAspectRatio,
imageQuality, imageCount, useImageGen}) which returns the image_config object
used in sendMessageWithHeaders, then replace the inline logic in
sendMessageWithHeaders and the component’s useEffect with calls to the new hook
and helper so both components consume the shared behavior.

In `@apps/playground/src/components/playground/chat-ui.tsx`:
- Around line 457-468: The duplicate image-model logic in chat-ui.tsx (variables
isGptImage and usesPixelDimensions derived from selectedModel and used for
size/quality UI) should be removed and replaced with a call to the shared image
config in apps/playground/src/lib/image-gen.ts; locate the isGptImage and
usesPixelDimensions definitions in chat-ui.tsx and replace them with imports and
usage of the exported helpers or config (e.g., functions or flags from
image-gen.ts that determine pixel-based dimensions and quality support) so the
UI consumes the single source of truth for image model detection and
size/quality rules.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro

Run ID: 26ac9f60-fef4-49fa-9054-2f4cef47191f

📥 Commits

Reviewing files that changed from the base of the PR and between 8b5379d and 0cc09ec.

⛔ Files ignored due to path filters (4)
  • apps/code/src/lib/api/v1.d.ts is excluded by !**/v1.d.ts
  • apps/playground/src/lib/api/v1.d.ts is excluded by !**/v1.d.ts
  • ee/admin/src/lib/api/v1.d.ts is excluded by !**/v1.d.ts
  • pnpm-lock.yaml is excluded by !**/pnpm-lock.yaml
📒 Files selected for processing (34)
  • .github/workflows/images.yml
  • apps/api/package.json
  • apps/docs/content/features/image-generation.mdx
  • apps/gateway/src/chat/chat.ts
  • apps/gateway/src/chat/schemas/completions.ts
  • apps/gateway/src/chat/tools/create-log-entry.ts
  • apps/gateway/src/chat/tools/parse-provider-response.ts
  • apps/gateway/src/chat/tools/resolve-provider-context.ts
  • apps/gateway/src/chat/tools/transform-response-to-openai.spec.ts
  • apps/gateway/src/chat/tools/transform-response-to-openai.ts
  • apps/gateway/src/images/images.ts
  • apps/gateway/src/lib/costs.spec.ts
  • apps/gateway/src/lib/costs.ts
  • apps/playground/package.json
  • apps/playground/src/app/api/chat/route.ts
  • apps/playground/src/app/api/image/route.ts
  • apps/playground/src/components/model-selector.tsx
  • apps/playground/src/components/playground/chat-page-client.tsx
  • apps/playground/src/components/playground/chat-ui.tsx
  • apps/playground/src/components/playground/image-controls.tsx
  • apps/playground/src/components/playground/image-page-client.tsx
  • apps/playground/src/lib/image-gen.ts
  • apps/ui/src/app/dashboard/[orgId]/[projectId]/activity/[logId]/log-detail-client.tsx
  • apps/ui/src/components/models/model-card.tsx
  • docs/providers/openai/gpt-image-2/openai-gpt-image-2.md
  • ee/admin/package.json
  • packages/actions/src/prepare-request-body.spec.ts
  • packages/actions/src/prepare-request-body.ts
  • packages/cache/src/swr.ts
  • packages/models/src/models/deepseek.ts
  • packages/models/src/models/openai.ts
  • packages/models/src/types.ts
  • packages/shared/src/components/log-card.tsx
  • scripts/image-edit.sh
✅ Files skipped from review due to trivial changes (8)
  • apps/playground/package.json
  • ee/admin/package.json
  • apps/gateway/src/chat/schemas/completions.ts
  • apps/api/package.json
  • .github/workflows/images.yml
  • apps/gateway/src/chat/tools/transform-response-to-openai.spec.ts
  • apps/gateway/src/chat/tools/resolve-provider-context.ts
  • packages/cache/src/swr.ts
🚧 Files skipped from review as they are similar to previous changes (1)
  • apps/ui/src/app/dashboard/[orgId]/[projectId]/activity/[logId]/log-detail-client.tsx

Comment thread apps/gateway/src/lib/costs.ts
Comment on lines +204 to +232
// Reset image size/quality when the selected model changes and the current
// value isn't valid for the new model. Including the value itself in deps
// would clobber the user's explicit selection on every re-render.
useEffect(() => {
const primaryModel = selectedModels[0] ?? "";
const config = getModelImageConfig(primaryModel);
if (!config.availableSizes.includes(imageSize as never)) {
if (config.usesPixelDimensions) {
if (config.isGptImage && alibabaImageSize === "1024x1024") {
setAlibabaImageSize(config.defaultSize);
} else if (
!(config.availableSizes as readonly string[]).includes(alibabaImageSize)
) {
setAlibabaImageSize(config.defaultSize);
}
} else if (
!(config.availableSizes as readonly string[]).includes(imageSize)
) {
setImageSize(config.defaultSize);
}
if (
!config.supportsQuality ||
!(config.availableQualities as readonly string[]).includes(imageQuality)
) {
setImageQuality(config.defaultQuality ?? "auto");
}
if (!isEditModel) {
setInputImages([]);
}
}, [selectedModels, imageSize, imageGenModels, isEditModel]);
}, [selectedModels, isEditModel]);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

Build image config per compared model, not from selectedModels[0] only.

In comparison mode, both the reset logic and imageConfigBody are keyed off the primary model, but that same payload is sent to every selected model. With the new image_quality field, comparing a quality-capable model with one that doesn’t support quality will leak image_quality into the secondary request; the same applies to pixel-dimension defaults. Compute getModelImageConfig(modelId) inside the request loop or restrict comparison mode to the intersection of supported settings.

Possible fix
-			const primaryModel = selectedModels[0] ?? "";
-			const config = getModelImageConfig(primaryModel);
-			// Always forward the user's quality choice (including "auto") so it
-			// shows up in the activity log; the gateway / model treat "auto" the
-			// same as omitting the field upstream.
-			const includeQuality = config.supportsQuality && !!imageQuality;
-			const imageConfigBody = config.usesPixelDimensions
-				? {
-						...(config.isGptImage
-							? alibabaImageSize !== "auto" && {
-									image_size: alibabaImageSize,
-								}
-							: alibabaImageSize !== "1024x1024" && {
-									image_size: alibabaImageSize,
-								}),
-						...(includeQuality && { image_quality: imageQuality }),
-						n: imageCount,
-					}
-				: {
-						...(imageAspectRatio !== "auto" && {
-							aspect_ratio: imageAspectRatio,
-						}),
-						...(imageSize !== "1K" && { image_size: imageSize }),
-						n: imageCount,
-					};
-
 			// Fire requests independently — each updates gallery as images stream in
 			pendingRef.current = selectedModels.length;
 
 			for (const modelId of selectedModels) {
+				const config = getModelImageConfig(modelId);
+				const includeQuality = config.supportsQuality && !!imageQuality;
+				const imageConfigBody = config.usesPixelDimensions
+					? {
+							...(config.isGptImage
+								? alibabaImageSize !== "auto" && {
+										image_size: alibabaImageSize,
+									}
+								: alibabaImageSize !== "1024x1024" && {
+										image_size: alibabaImageSize,
+									}),
+							...(includeQuality && { image_quality: imageQuality }),
+							n: imageCount,
+						}
+					: {
+							...(imageAspectRatio !== "auto" && {
+								aspect_ratio: imageAspectRatio,
+							}),
+							...(imageSize !== "1K" && { image_size: imageSize }),
+							n: imageCount,
+						};
+
 				const noFallback = shouldDisableFallback(modelId);

Also applies to: 301-315

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@apps/playground/src/components/playground/image-page-client.tsx` around lines
204 - 232, The reset and request payload logic currently derives config from
selectedModels[0]; instead, call getModelImageConfig(modelId) for each model
when in comparison mode so each model's reset and imageConfigBody is built from
its own config (e.g., inside the loop that sends requests over selectedModels),
and only include fields like image_quality or pixel-dimension settings when that
model's config.supportsQuality or config.usesPixelDimensions respectively;
update the useEffect reset logic and the imageConfigBody construction to compute
config per model (referencing selectedModels, getModelImageConfig,
imageConfigBody, and the imageQuality/imageSize/alibabaImageSize setters) to
avoid leaking unsupported fields to secondary models.

Comment on lines +114 to +118
function hasEstimatedImageCost(mapping: ApiModelProviderMapping): boolean {
return Boolean(
mapping.imageOutputPrice && mapping.imageOutputTokensByResolution,
);
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Guard hasEstimatedImageCost against empty resolution maps.

An empty {} map currently evaluates as estimated, which can suppress default token pricing until expand, despite no usable estimate rows.

Suggested fix
 function hasEstimatedImageCost(mapping: ApiModelProviderMapping): boolean {
-	return Boolean(
-		mapping.imageOutputPrice && mapping.imageOutputTokensByResolution,
-	);
+	const hasPrice =
+		mapping.imageOutputPrice !== null &&
+		mapping.imageOutputPrice !== undefined;
+	const byResolution = mapping.imageOutputTokensByResolution;
+	const hasEstimateRows =
+		!!byResolution && Object.keys(byResolution).length > 0;
+	return hasPrice && hasEstimateRows;
 }

Also applies to: 516-519, 958-958

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@apps/ui/src/components/models/model-card.tsx` around lines 114 - 118,
hasEstimatedImageCost currently returns true when imageOutputTokensByResolution
is an empty object; change the guard in hasEstimatedImageCost to require both
mapping.imageOutputPrice truthy and mapping.imageOutputTokensByResolution to
have at least one key (e.g., Object.keys(mapping.imageOutputTokensByResolution
|| {}).length > 0). Update the same pattern wherever the same logic appears (the
other occurrences that check imageOutputPrice and imageOutputTokensByResolution)
so empty maps do not count as providing an estimated image cost.

Comment on lines +207 to +235
const requestBody = (await prepareOpenAIImageRequest({
image_size: size,
image_quality: "high",
n: 1,
})) as any;

expect(requestBody).toMatchObject({
model: "gpt-image-2",
prompt: "Generate a cinematic landscape",
size,
quality: "high",
n: 1,
});
});

test("should not derive size from aspect_ratio", async () => {
const requestBody = (await prepareOpenAIImageRequest({
aspect_ratio: "16:9",
})) as any;

expect(requestBody.size).toBeUndefined();
});

test("should drop unsupported quality values", async () => {
const requestBody = (await prepareOpenAIImageRequest({
image_size: "1024x1024",
image_quality: "standard",
})) as any;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Remove new as any casts in these tests.

Line 211, Line 225, and Line 234 introduce new as any casts that are not necessary here. A typed helper return keeps assertions readable without dropping type safety.

♻️ Suggested fix
+type OpenAIImagePreparedBody = {
+	model: string;
+	prompt: string;
+	size?: string;
+	quality?: "low" | "medium" | "high" | "auto";
+	n?: number;
+};
+
 async function prepareOpenAIImageRequest(imageConfig: {
 	aspect_ratio?: string;
 	image_size?: string;
 	image_quality?: string;
 	n?: number;
-}) {
-	return await prepareRequestBody(
+}): Promise<OpenAIImagePreparedBody> {
+	return (await prepareRequestBody(
 		"openai",
 		"gpt-image-2",
 		[{ role: "user", content: "Generate a cinematic landscape" }],
 		false,
 		undefined,
@@
 		imageConfig,
 		undefined,
 		true,
-	);
+	)) as OpenAIImagePreparedBody;
 }
@@
-		const requestBody = (await prepareOpenAIImageRequest({
+		const requestBody = await prepareOpenAIImageRequest({
 			image_size: size,
 			image_quality: "high",
 			n: 1,
-		})) as any;
+		});
@@
-		const requestBody = (await prepareOpenAIImageRequest({
+		const requestBody = await prepareOpenAIImageRequest({
 			aspect_ratio: "16:9",
-		})) as any;
+		});
@@
-		const requestBody = (await prepareOpenAIImageRequest({
+		const requestBody = await prepareOpenAIImageRequest({
 			image_size: "1024x1024",
 			image_quality: "standard",
-		})) as any;
+		});

As per coding guidelines, **/*.{ts,tsx}: Never use any or as any in TypeScript unless absolutely necessary.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@packages/actions/src/prepare-request-body.spec.ts` around lines 207 - 235,
Remove the unnecessary "as any" test casts: update the three occurrences where
requestBody is declared (calls to prepareOpenAIImageRequest in the image tests)
to use the actual return type instead of "as any" (e.g., const requestBody =
await prepareOpenAIImageRequest({...})); this keeps type safety and lets
existing expect assertions work—adjust the variable declarations for the tests
"should not derive size from aspect_ratio", "should drop unsupported quality
values", and the earlier image-size test to drop the "as any" suffix.

Comment on lines +685 to 695
// Pass image_size straight through to OpenAI as `WxH` (or `auto`).
// OpenAI returns a 4xx for unsupported sizes, which we propagate.
const openaiSize = image_config?.image_size;
const openaiQuality = normalizeImageQuality(image_config?.image_quality);

const openaiImageRequest: OpenAIImageRequest = {
model: usedModel,
prompt,
...(openaiSize && { size: openaiSize }),
...(openaiQuality && { quality: openaiQuality }),
...(image_config?.n && { n: image_config.n }),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

Reject unsupported OpenAI aspect_ratio values instead of silently dropping them.

apps/gateway/src/images/images.ts still populates image_config.aspect_ratio, but this branch now forwards only image_size/quality. A request that sets only aspect_ratio will quietly fall back to OpenAI’s default size, which is a behavior regression and very hard for callers to detect. Either map the supported ratios here or fail fast when aspect_ratio is present without a concrete image_size.

Possible fix
 		// Pass image_size straight through to OpenAI as `WxH` (or `auto`).
 		// OpenAI returns a 4xx for unsupported sizes, which we propagate.
 		const openaiSize = image_config?.image_size;
+		if (image_config?.aspect_ratio && !openaiSize) {
+			throw new Error(
+				"OpenAI image generation requires image_size; aspect_ratio alone is not supported for this provider.",
+			);
+		}
 		const openaiQuality = normalizeImageQuality(image_config?.image_quality);
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@packages/actions/src/prepare-request-body.ts` around lines 685 - 695, The
code currently ignores image_config.aspect_ratio when image_size is not set,
causing silent fallback; update the image request construction logic (involving
image_config, openaiImageRequest, image_size, aspect_ratio,
normalizeImageQuality) to either map supported aspect_ratio values to concrete
OpenAI size strings or reject/throw a clear error when image_config.aspect_ratio
is present but image_config.image_size is not; validate aspect_ratio against an
explicit whitelist of supported ratios and ensure the rejection occurs before
building openaiImageRequest so callers receive a fast, descriptive failure
instead of silently using OpenAI defaults.

Comment thread scripts/image-edit.sh
Comment on lines +20 to 21
N=${N:-1}
RESPONSE_FILE=${RESPONSE_FILE:-.context/image-edit-response-${TIMESTAMP}.json}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Validate N before building the payload.

N is converted with int(...) later, so non-numeric values fail late with a Python traceback. Fail fast in bash with a clearer message.

Suggested fix
 N=${N:-1}
+if ! [[ "$N" =~ ^[1-9][0-9]*$ ]]; then
+	echo "N must be a positive integer (got: $N)" >&2
+	exit 1
+fi
 RESPONSE_FILE=${RESPONSE_FILE:-.context/image-edit-response-${TIMESTAMP}.json}

Also applies to: 84-84

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@scripts/image-edit.sh` around lines 20 - 21, Validate the N variable
immediately after it is set and before any payload construction: check that N
contains only digits (and optionally is >=1) and exit with a clear error message
if not; update the script near where N is defined (reference variable N and
TIMESTAMP/RESPONSE_FILE usage) and also add the same validation before the
payload-building section referenced around the later payload block (the code
near line ~84) so the script fails fast with a helpful message instead of
letting Python raise a traceback.

# Conflicts:
#	apps/docs/content/features/image-generation.mdx
#	apps/playground/src/components/playground/chat-page-client.tsx
Move the chars/4 estimator into @llmgateway/shared with explicit
multimodal handling: only text content is counted; image/file/audio
parts are skipped. The previous JSON.stringify path on multimodal
content double-counted image bytes against costs.ts image-input
billing for vision/image requests when upstream usage was missing.

Tracked follow-up for multimodal-aware estimation: #2112.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
@steebchen
steebchen enabled auto-merge April 29, 2026 11:54

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (1)
packages/shared/src/token-estimate.spec.ts (1)

26-91: Add boundary tests around the 5k-token routing cutoff.

Given cache-weight routing depends on a 5k threshold, please add explicit cases near the boundary (for example, 19,999 vs 20,000 chars) to lock expected behavior.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@packages/shared/src/token-estimate.spec.ts` around lines 26 - 91, Add
explicit boundary tests in token-estimate.spec.ts around the 5k-token cutoff
using estimateChatMessageTokens: create two tests that build messages with a
single string content of length 19,999 and 20,000 characters respectively, call
estimateChatMessageTokens([{ content: longString }]) for each, and assert the
token estimate is below 5000 for the 19,999-char case and at or above 5000 for
the 20,000-char case so the routing cutoff behavior is locked in.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@apps/gateway/src/chat/tools/tokenizer.ts`:
- Around line 28-30: Replace the use of any[] in encodeChatMessages with a
concrete message type that matches what estimateChatMessageTokens expects:
define or import a ChatMessage type/interface where each message includes a
content field typed as string | Array<{ type?: string; text?: string }> | null
(and any other optional fields your codebase requires), update the function
signature to encodeChatMessages(messages: ChatMessage[]): number, and pass that
typed array into estimateChatMessageTokens so the compiler enforces the correct
structure; reference the encodeChatMessages function and the
estimateChatMessageTokens call when making the change.

---

Nitpick comments:
In `@packages/shared/src/token-estimate.spec.ts`:
- Around line 26-91: Add explicit boundary tests in token-estimate.spec.ts
around the 5k-token cutoff using estimateChatMessageTokens: create two tests
that build messages with a single string content of length 19,999 and 20,000
characters respectively, call estimateChatMessageTokens([{ content: longString
}]) for each, and assert the token estimate is below 5000 for the 19,999-char
case and at or above 5000 for the 20,000-char case so the routing cutoff
behavior is locked in.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro

Run ID: a2228502-e2e9-4740-bb05-d42ad6a18f83

📥 Commits

Reviewing files that changed from the base of the PR and between 0cc09ec and 11d6eb1.

⛔ Files ignored due to path filters (4)
  • apps/code/src/lib/api/v1.d.ts is excluded by !**/v1.d.ts
  • apps/playground/src/lib/api/v1.d.ts is excluded by !**/v1.d.ts
  • apps/ui/src/lib/api/v1.d.ts is excluded by !**/v1.d.ts
  • ee/admin/src/lib/api/v1.d.ts is excluded by !**/v1.d.ts
📒 Files selected for processing (6)
  • apps/gateway/src/chat/tools/estimate-tokens-from-content.ts
  • apps/gateway/src/chat/tools/tokenizer.ts
  • apps/gateway/src/lib/prompt-tokens.spec.ts
  • packages/shared/src/index.ts
  • packages/shared/src/token-estimate.spec.ts
  • packages/shared/src/token-estimate.ts
✅ Files skipped from review due to trivial changes (1)
  • packages/shared/src/index.ts
🚧 Files skipped from review as they are similar to previous changes (1)
  • apps/gateway/src/lib/prompt-tokens.spec.ts

Comment on lines 28 to 30
export function encodeChatMessages(messages: any[]): number {
try {
const chatMessages = messages.map((m) => ({
role: m.role as "user" | "assistant" | "system" | undefined,
content: messageContentToString(m.content),
...(m.name !== null && m.name !== undefined && { name: m.name }),
}));
return encodeChat(chatMessages, DEFAULT_TOKENIZER_MODEL).length;
} catch (error) {
logger.error("Failed to encode chat messages", {
error: error instanceof Error ? error.message : String(error),
messageCount: messages.length,
messageRoles: messages.map((m) => m.role),
messageContentTypes: messages.map((m) => typeof m.content),
});
// Fallback: rough 4-chars-per-token estimate
return Math.max(
1,
Math.round(
messages.reduce(
(acc: number, m: any) =>
acc + messageContentToString(m.content).length,
0,
) / 4,
),
);
}
return estimateChatMessageTokens(messages);
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# Verify current `any[]` usage and inspect encodeChatMessages call sites before tightening types.
rg -n --type=ts -C2 'export function encodeChatMessages\(messages:\s*any\[\]\)'
rg -n --type=ts -C2 '\bencodeChatMessages\('

Repository: theopenco/llmgateway

Length of output: 2749


🏁 Script executed:

#!/bin/bash
# Check the actual types of messages at call sites
# 1. Check costs.ts to see type of fullOutput.messages
head -160 apps/gateway/src/lib/costs.ts | tail -20

# 2. Check chat.ts to see type of messages parameter where encodeChatMessages is called
grep -B20 "encodeChatMessages(messages)" apps/gateway/src/chat/chat.ts | head -40

# 3. Look for existing message type definitions
rg -n "type.*Message|interface.*Message" --type=ts | grep -v node_modules | head -20

# 4. Check estimate-tokens.ts
rg -B5 "encodeChatMessages" apps/gateway/src/chat/tools/estimate-tokens.ts

# 5. Check the estimateChatMessageTokens function signature
rg -A10 "function estimateChatMessageTokens" --type=ts

Repository: theopenco/llmgateway

Length of output: 5181


🏁 Script executed:

#!/bin/bash
# Get the MessageLike interface definition
cat -n packages/shared/src/token-estimate.ts | head -50

# Check if encodeChatMessages is defined in tokenizer.ts with more context
cat -n apps/gateway/src/chat/tools/tokenizer.ts

# Check the types imported/used in chat.ts for the messages parameter
grep -B10 "function.*messages" apps/gateway/src/chat/chat.ts | head -30

# Check if there are type definitions for fullOutput in costs.ts
rg -B5 "fullOutput" apps/gateway/src/lib/costs.ts | head -30

Repository: theopenco/llmgateway

Length of output: 4083


🏁 Script executed:

#!/bin/bash
# Check if MessageLike is exported from the shared package
rg -n "export.*MessageLike|export.*interface.*MessageLike" packages/shared/src/

# Check the exports in token-estimate.ts
grep -n "export" packages/shared/src/token-estimate.ts

# Check what's currently imported in tokenizer.ts
head -10 apps/gateway/src/chat/tools/tokenizer.ts

# Check if `@llmgateway/shared` is available as a dependency in gateway
grep -n "@llmgateway/shared" apps/gateway/package.json

# Check the full token-estimate.ts to see if MessageLike is exported
cat packages/shared/src/token-estimate.ts

Repository: theopenco/llmgateway

Length of output: 2476


Replace any[] in encodeChatMessages with a concrete message type.

Line 28 uses any[], which violates the coding guidelines and drops type safety on a routing/cost-critical path. The function passes messages directly to estimateChatMessageTokens from @llmgateway/shared, which expects a concrete structure with content (string, array of parts, or null) and parts with optional type and text fields.

Suggested type-safe change
+type EstimationContentPart = {
+	type?: string;
+	text?: string;
+};
+
+type EstimationMessage = {
+	content?: string | EstimationContentPart[] | null;
+};
+
-export function encodeChatMessages(messages: any[]): number {
+export function encodeChatMessages(messages: EstimationMessage[]): number {
	return estimateChatMessageTokens(messages);
}
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@apps/gateway/src/chat/tools/tokenizer.ts` around lines 28 - 30, Replace the
use of any[] in encodeChatMessages with a concrete message type that matches
what estimateChatMessageTokens expects: define or import a ChatMessage
type/interface where each message includes a content field typed as string |
Array<{ type?: string; text?: string }> | null (and any other optional fields
your codebase requires), update the function signature to
encodeChatMessages(messages: ChatMessage[]): number, and pass that typed array
into estimateChatMessageTokens so the compiler enforces the correct structure;
reference the encodeChatMessages function and the estimateChatMessageTokens call
when making the change.

@steebchen
steebchen added this pull request to the merge queue Apr 29, 2026
Merged via the queue into main with commit 7ff05f5 Apr 29, 2026
33 of 34 checks passed
@steebchen
steebchen deleted the weight-cache-estimate branch April 29, 2026 12:08
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants