feat: image-aware token estimate for auto-routing (#2112) - #2475
Conversation
Extend estimateChatMessageTokens with opt-in, model-parameterized image awareness. When a model id is passed, image_url/image parts contribute the model's per-image token count (from imageInputTokensByResolution, default 560) and file/audio/video parts a flat default — mirroring the cost tables without serializing the raw blob. Without a model id the estimate stays text-only, so billing-path callers can opt out and avoid double-counting against the separate imageInputCost in costs.ts. Adds unit tests for text-only, mixed text+image, image-only, multi-image, unknown-model fallback, and file/audio cases. Refs theopenco#2112 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository UI Review profile: CHILL Plan: Pro Run ID: 📒 Files selected for processing (1)
🚧 Files skipped from review as they are similar to previous changes (1)
WalkthroughRecomputes routing input-token estimates using an image-aware encoder: adds multimodal token constants and classifiers, extends estimator to accept modelId and fallback reporting, forwards modelId through the gateway tokenizer with fallback logging, adds tests, and re-exports the fallback-reporting type. ChangesMultimodal Token Estimation
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~20 minutes Possibly related PRs
Suggested reviewers
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@apps/gateway/src/chat/tools/tokenizer.ts`:
- Around line 31-32: Replace the loose any[] in encodeChatMessages with the
shared concrete type: import MessageLike (and ContentPart if required) from the
token-estimate module and change the signature to messages: MessageLike[];
update any downstream callers if types flow through, and ensure the shared
package exports MessageLike/ContentPart so estimateChatMessageTokens can be used
with the exact multimodal contract.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository UI
Review profile: CHILL
Plan: Pro
Run ID: bfb0e3c0-4db4-4a7e-bd43-2fa8cd810601
📒 Files selected for processing (4)
apps/gateway/src/chat/chat.tsapps/gateway/src/chat/tools/tokenizer.tspackages/shared/src/token-estimate.spec.tspackages/shared/src/token-estimate.ts
| export function encodeChatMessages(messages: any[], modelId?: string): number { | ||
| return estimateChatMessageTokens(messages, modelId); |
There was a problem hiding this comment.
🧩 Analysis chain
🏁 Script executed:
#!/bin/bash
# Verify whether a reusable message type already exists before exporting a new one.
rg -n -C2 'export (interface|type) .*MessageLike|interface MessageLike|type MessageLike' packages/shared/src apps/gateway/src
rg -n -C2 'encodeChatMessages\(' apps/gateway/srcRepository: theopenco/llmgateway
Length of output: 3069
🏁 Script executed:
#!/bin/bash
set -euo pipefail
# Inspect shared token-estimate types around MessageLike
sed -n '1,120p' packages/shared/src/token-estimate.ts
sed -n '120,240p' packages/shared/src/token-estimate.ts
# Inspect tokenizer + estimator signatures
sed -n '1,120p' apps/gateway/src/chat/tools/tokenizer.ts
sed -n '1,200p' apps/gateway/src/chat/tools/estimate-tokens.ts
sed -n '1,200p' apps/gateway/src/chat/tools/calculate-prompt-tokens.tsRepository: theopenco/llmgateway
Length of output: 7073
🏁 Script executed:
#!/bin/bash
set -euo pipefail
rg -n 'messages:\s*any\[\]' apps/gateway/src/chat/tools
rg -n 'encodeChatMessages\(messages:\s*any' -S apps/gateway/src/chat/tools || true
rg -n 'any\[\]' apps/gateway/src/chat/tools/tokenizer.ts apps/gateway/src/chat/tools/estimate-tokens.ts apps/gateway/src/chat/tools/calculate-prompt-tokens.tsRepository: theopenco/llmgateway
Length of output: 1254
🏁 Script executed:
#!/bin/bash
set -euo pipefail
rg -n 'messages:\s*any\[\]' apps/gateway/src/chat/tools
rg -n 'any\[\]' apps/gateway/src/chat/tools/tokenizer.ts apps/gateway/src/chat/tools/estimate-tokens.ts apps/gateway/src/chat/tools/calculate-prompt-tokens.tsRepository: theopenco/llmgateway
Length of output: 1130
🏁 Script executed:
#!/bin/bash
set -euo pipefail
rg -n 'messages:\s*any\[\]' apps/gateway/src/chat/tools
rg -n 'any\[\]' apps/gateway/src/chat/tools/tokenizer.ts apps/gateway/src/chat/tools/estimate-tokens.ts apps/gateway/src/chat/tools/calculate-prompt-tokens.tsRepository: theopenco/llmgateway
Length of output: 1130
Replace any[] in encodeChatMessages with a concrete message type.
apps/gateway/src/chat/tools/tokenizer.ts currently declares encodeChatMessages(messages: any[], ...), but the shared token estimator in packages/shared/src/token-estimate.ts already models this as MessageLike[] (with content?: string | ContentPart[] | null). Export MessageLike (and ContentPart if needed) from shared and use it here to preserve the multimodal contract.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@apps/gateway/src/chat/tools/tokenizer.ts` around lines 31 - 32, Replace the
loose any[] in encodeChatMessages with the shared concrete type: import
MessageLike (and ContentPart if required) from the token-estimate module and
change the signature to messages: MessageLike[]; update any downstream callers
if types flow through, and ensure the shared package exports
MessageLike/ContentPart so estimateChatMessageTokens can be used with the exact
multimodal contract.
|
@Lawson-Darrow thanks for this. on unknown cases where we don't have information, we should at least log a logger.warn() so that we can collect unknown cases where the estimation is triggered and continuously improve these methods in the future. additionally we should check if there is information on how input tokens on document/audio/video input are tracked for the flagship models/providers and check if we can integrate that also in the models package and track it there if it's available. |
|
Sounds good. I'll add a logger.warn on the fallback cases so we capture when the rough estimate kicks in. I'll also look into whether the flagship providers publish input-token counts for document/audio/video and add them to the models package if they do, probably as a follow-up PR. |
…eopenco#2112) estimateChatMessageTokens now reports via an optional onFallback callback when it had to use a rough default: the model has no per-image token table, or there are file/audio/video parts which have no per-model data yet. encodeChatMessages logs these via logger.warn so the unknown cases can be collected and the per-model token data improved over time, per maintainer feedback. Adds unit tests for the fallback reporting.
|
Done. estimateChatMessageTokens now reports fallbacks through an onFallback callback and encodeChatMessages logs them with logger.warn at the gateway layer. Each warn fires once per estimation (not per part) and includes the unknown modelId plus the count of image parts and file/audio/video parts that hit the rough default, so we can collect the unknown cases and improve the per-model data over time. On your second point about tracking document/audio/video input tokens for the flagship providers in the models package, that is a good follow-up and I will look into it separately so this PR can land on the routing-estimate change. |
There was a problem hiding this comment.
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (2)
apps/gateway/src/chat/chat.ts (2)
2553-2603:⚠️ Potential issue | 🟠 Major | ⚡ Quick winCarry the
supportsNfilter into the requested-provider rate-limit fallback.This fallback path rebuilds candidates manually and never excludes mappings with
supportsN !== true. Forn > 1, a rate-limited requested provider can be rerouted onto an incompatible mapping and then fail later at Line 4619 with a 400, even when another valid fallback exists. ReusefilterEligibleModelProviders(..., { n, ... })here or mirror itssupportsNpredicate.Also applies to: 4611-4623
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@apps/gateway/src/chat/chat.ts` around lines 2553 - 2603, The fallback candidate rebuild currently omits the supportsN check so mappings with supportsN !== true can be chosen for requests with n>1; update the filter used to construct availableModelProviders to include the same supportsN predicate used by filterEligibleModelProviders (or simply call filterEligibleModelProviders with the same parameters, e.g. passing n) so that when n > 1 you only allow providers/mappings where provider.supportsN === true; keep all existing predicates (usedProvider/usedRegion, webSearchTool, response_format/jsonOutput/jsonOutputSchema, hasImages, hasAudio + googleProviderSupportsAudioFormat, hasDocuments) and add the supportsN guard tied to the request's n value.
1926-1935:⚠️ Potential issue | 🟠 Major | ⚡ Quick winUse the multimodal estimate for provider scoring too.
estimatedInputTokensonly feedsrequiredContextSize; the routing helpers right below still receive text-onlyroutingPromptTokens. Large image/file/audio requests can therefore pass the new context-size gate but still be ranked with the old text-only weight, which keeps cache/price-based provider selection blind to the multimodal payload.Also applies to: 2201-2222
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@apps/gateway/src/chat/chat.ts` around lines 1926 - 1935, The routing/scoring still uses a text-only routingPromptTokens while requiredContextSize is computed with multimodal inputs (estimatedInputTokens includes tools and encodeChatMessages); update the provider scoring call(s) to use the multimodal estimate (requiredContextSize or a new variable computed from estimatedInputTokens including tools/files/images/audio) instead of routingPromptTokens so cache/price-based provider selection accounts for multimodal payloads (apply the same change in the other block around the 2201-2222 region); locate uses of routingPromptTokens passed into the routing helpers and substitute or augment them with requiredContextSize/estimatedInputTokens.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Outside diff comments:
In `@apps/gateway/src/chat/chat.ts`:
- Around line 2553-2603: The fallback candidate rebuild currently omits the
supportsN check so mappings with supportsN !== true can be chosen for requests
with n>1; update the filter used to construct availableModelProviders to include
the same supportsN predicate used by filterEligibleModelProviders (or simply
call filterEligibleModelProviders with the same parameters, e.g. passing n) so
that when n > 1 you only allow providers/mappings where provider.supportsN ===
true; keep all existing predicates (usedProvider/usedRegion, webSearchTool,
response_format/jsonOutput/jsonOutputSchema, hasImages, hasAudio +
googleProviderSupportsAudioFormat, hasDocuments) and add the supportsN guard
tied to the request's n value.
- Around line 1926-1935: The routing/scoring still uses a text-only
routingPromptTokens while requiredContextSize is computed with multimodal inputs
(estimatedInputTokens includes tools and encodeChatMessages); update the
provider scoring call(s) to use the multimodal estimate (requiredContextSize or
a new variable computed from estimatedInputTokens including
tools/files/images/audio) instead of routingPromptTokens so cache/price-based
provider selection accounts for multimodal payloads (apply the same change in
the other block around the 2201-2222 region); locate uses of routingPromptTokens
passed into the routing helpers and substitute or augment them with
requiredContextSize/estimatedInputTokens.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository UI
Review profile: CHILL
Plan: Pro
Run ID: 8e264764-6fef-47e8-b219-a189baf7874d
📒 Files selected for processing (2)
apps/gateway/src/chat/chat.tspackages/shared/src/index.ts
💤 Files with no reviewable changes (1)
- packages/shared/src/index.ts
|
Quick context on the two CodeRabbit flags so they don't hold things up. The supportsN rate-limit fallback one is pre-existing and unrelated to token estimation, so it is out of scope for this PR. Happy to file it as its own issue if it looks like a real gap. The provider-scoring one is intentional. routingPromptTokens stays text-only so billing is unchanged, and the new multimodal estimate only gates requiredContextSize. The diff has a comment calling that out. I can extend scoring to the multimodal estimate in a follow-up if you would rather it factor into cache and price selection. |
…heopenco#2475) Closes theopenco#2112. `estimateChatMessageTokens` now takes an optional `modelId`. Pass it and image/file/audio parts get counted on top of the text chars/4 estimate. Images use the model's per-image token table and fall back to 560. Leave it off and the estimate stays text-only, so the billing path in `costs.ts` is unchanged and images don't get double counted (`imageInputCost` already handles them). Auto-routing now calls `encodeChatMessages(messages, requestedModel)` so model selection and the context-size check react to image payloads. `routingPromptTokens` stays text-only since it backs the billing fallback. No tokenizer on the hot path. It's the same chars/4 heuristic plus one table lookup per image. I tested this against the running gateway. An `auto` request with an image estimates 560 more tokens than the same request without one, and 0 difference when there's no image. Unit tests are in `packages/shared`. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved token estimation for auto-routing/model selection so images and other multimodal content are properly counted, producing more accurate context-window sizing and better model/provider choices for chats with rich media. * **Tests** * Expanded test coverage for multimodal token estimation, per-image counting, mixed-content behavior, and fallback-reporting when model-specific data is absent. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: Luca Steeb <contact@luca-steeb.com>
…heopenco#2475) Closes theopenco#2112. `estimateChatMessageTokens` now takes an optional `modelId`. Pass it and image/file/audio parts get counted on top of the text chars/4 estimate. Images use the model's per-image token table and fall back to 560. Leave it off and the estimate stays text-only, so the billing path in `costs.ts` is unchanged and images don't get double counted (`imageInputCost` already handles them). Auto-routing now calls `encodeChatMessages(messages, requestedModel)` so model selection and the context-size check react to image payloads. `routingPromptTokens` stays text-only since it backs the billing fallback. No tokenizer on the hot path. It's the same chars/4 heuristic plus one table lookup per image. I tested this against the running gateway. An `auto` request with an image estimates 560 more tokens than the same request without one, and 0 difference when there's no image. Unit tests are in `packages/shared`. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved token estimation for auto-routing/model selection so images and other multimodal content are properly counted, producing more accurate context-window sizing and better model/provider choices for chats with rich media. * **Tests** * Expanded test coverage for multimodal token estimation, per-image counting, mixed-content behavior, and fallback-reporting when model-specific data is absent. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: Luca Steeb <contact@luca-steeb.com>
Closes #2112.
estimateChatMessageTokensnow takes an optionalmodelId. Pass it and image/file/audio parts get counted on top of the text chars/4 estimate. Images use the model's per-image token table and fall back to 560. Leave it off and the estimate stays text-only, so the billing path incosts.tsis unchanged and images don't get double counted (imageInputCostalready handles them).Auto-routing now calls
encodeChatMessages(messages, requestedModel)so model selection and the context-size check react to image payloads.routingPromptTokensstays text-only since it backs the billing fallback.No tokenizer on the hot path. It's the same chars/4 heuristic plus one table lookup per image.
I tested this against the running gateway. An
autorequest with an image estimates 560 more tokens than the same request without one, and 0 difference when there's no image. Unit tests are inpackages/shared.Summary by CodeRabbit