feat(devpass): require cached-input pricing for coding plans - #2150
Conversation
Coding plans (personal orgs on a dev plan with devPlanAllowAllModels=false) gate models via isCodingModel, but a request like groq/gpt-oss-120b passes the model-level check while routing to an uncached mapping. Deny specific provider requests without cached pricing, and filter routing/auto-routing candidates so canonical mappings only resolve to cached providers. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
WalkthroughAdds a cached-input capability check and gates coding-plan/dev-plan enforcement in the gateway. Introduces ChangesCached-Input Provider Filtering for Development Plans
Sequence Diagram(s)sequenceDiagram
autonumber
participant Client
participant Gateway
participant IAM
participant ProviderRegistry
Client->>Gateway: Send chat request (model, optional provider mapping, region)
Gateway->>ProviderRegistry: Lookup model/provider mappings
Gateway->>IAM: Validate caller permissions for provider/model
alt isDevPlanRestricted
Gateway->>ProviderRegistry: Filter mappings where providerSupportsCachedInput == true
ProviderRegistry-->>Gateway: Filtered provider list
Gateway->>Gateway: If no providers remain → 403 return
end
opt Auto-routing
Gateway->>ProviderRegistry: Select suitable provider from filtered candidates
end
Gateway->>IAM: Revalidate IAM for chosen provider/model
Gateway->>Provider: Forward request (or 403 if checks fail)
Provider-->>Gateway: Response
Gateway-->>Client: Return response or 403
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~25 minutes Possibly related PRs
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Review rate limit: 7/8 reviews remaining, refill in 7 minutes and 30 seconds.Comment |
There was a problem hiding this comment.
Pull request overview
This PR tightens “coding plan” model/provider eligibility in the gateway to ensure dev-plan personal orgs only route to provider mappings that have cached-input pricing, preventing requests from slipping to uncached upstream mappings and consuming credits unexpectedly.
Changes:
- Adds a reusable
providerSupportsCachedInputhelper and uses it inisCodingModel. - Enforces cached-input pricing at request time in
chat.tsfor dev-plan-restricted orgs (both for explicitly requested providers and for routed candidates). - Adds unit tests for
providerSupportsCachedInputandisCodingModel.
Reviewed changes
Copilot reviewed 3 out of 3 changed files in this pull request and generated 2 comments.
| File | Description |
|---|---|
| apps/gateway/src/lib/coding-models.ts | Adds providerSupportsCachedInput helper and reuses it in isCodingModel. |
| apps/gateway/src/lib/coding-models.spec.ts | Adds unit tests for cached-input detection and coding-model qualification. |
| apps/gateway/src/chat/chat.ts | Applies dev-plan cached-input enforcement for explicit provider requests and routing candidate filtering. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| import type { ModelDefinition, ProviderModelMapping } from "@llmgateway/models"; | ||
|
|
||
| export function providerSupportsCachedInput( | ||
| p: Pick<ProviderModelMapping, "cachedInputPrice">, |
| it("returns false when cachedInputPrice is null", () => { | ||
| expect( | ||
| providerSupportsCachedInput({ | ||
| cachedInputPrice: null as unknown as undefined, | ||
| }), | ||
| ).toBe(false); |
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Inline comments:
In `@apps/gateway/src/chat/chat.ts`:
- Around line 1431-1449: The current check only ensures there exists some
cached-input mapping for requestedProvider but not that the specific mapping
chosen will support cached input; update the region-aware provider selection to
apply the providerSupportsCachedInput filter (or else validate the exact
resolved mapping) so the final candidate set built from modelInfo.providers
excludes region-specific entries lacking cached input. Concretely, when
resolving provider mappings for requestedProvider and requestedRegion (the code
that builds requestedProviderMappings and the later direct-provider resolution),
filter by providerSupportsCachedInput(p) and/or re-check the selected mapping
against providerSupportsCachedInput before proceeding, ensuring the chosen
provider+region pair is confirmed cached-input-capable.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository UI
Review profile: CHILL
Plan: Pro
Run ID: 3f8da705-6d19-4a20-acbe-f2c00f95b310
📒 Files selected for processing (3)
apps/gateway/src/chat/chat.tsapps/gateway/src/lib/coding-models.spec.tsapps/gateway/src/lib/coding-models.ts
The direct-provider region picker built candidates straight from modelInfo.providers, so a provider-key-locked uncached region could still resolve to an uncached upstream after the model-level gate passed. Filter sameProviderMappings by cached pricing for restricted dev plans, and 403 when a locked region has no cached mapping. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
There was a problem hiding this comment.
Actionable comments posted: 2
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Inline comments:
In `@apps/gateway/src/chat/chat.ts`:
- Around line 1902-1909: After narrowing sameProviderMappings to cached-only
entries, ensure you run filterEligibleModelProviders(...) against that cached
list and, if isDevPlanRestricted is true and the filtered result is empty, throw
the 403 HTTPException (same message used for cached-input restriction) instead
of falling through; do the same guard where you later compute
sameProviderRoutingMappings so you don't accidentally pick a cached region that
violates providerLockedRegions, webSearchTool, response_format, max_tokens, or
reasoningEffort.
- Around line 1876-1882: After applying the cached-input filter when
isDevPlanRestricted, add the same abort/exit logic used earlier: check
iamFilteredModelProviders and expandedIamFilteredModelProviders for being empty
and, if so, stop the request (throw/return an error or set the same error
response) rather than allowing processing to continue with a stale usedProvider;
use the same behavior/response as the earlier check (lines around 1508–1519) so
that when providerSupportsCachedInput filtering yields [] the flow is aborted
and no request proceeds with an invalid usedProvider.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository UI
Review profile: CHILL
Plan: Pro
Run ID: 2542e843-15f6-4241-acd1-3d9b70b98bb5
📒 Files selected for processing (1)
apps/gateway/src/chat/chat.ts
| if (isDevPlanRestricted) { | ||
| iamFilteredModelProviders = iamFilteredModelProviders.filter( | ||
| providerSupportsCachedInput, | ||
| ); | ||
| expandedIamFilteredModelProviders = | ||
| expandedIamFilteredModelProviders.filter(providerSupportsCachedInput); | ||
| } |
There was a problem hiding this comment.
Abort when post-auto filtering removes every cached-capable provider.
This block re-filters the resolved model after IAM, but unlike the earlier check at Lines 1508-1519 it never stops on []. If auto-routing fell back to the default claude-haiku-4-5/anthropic path, usedProvider stays set and the request can continue even though no cached-input-capable provider is actually allowed for the resolved model.
Suggested fix
if (isDevPlanRestricted) {
iamFilteredModelProviders = iamFilteredModelProviders.filter(
providerSupportsCachedInput,
);
expandedIamFilteredModelProviders =
expandedIamFilteredModelProviders.filter(providerSupportsCachedInput);
+ if (iamFilteredModelProviders.length === 0) {
+ throw new HTTPException(403, {
+ message: `No provider with cached input pricing is available for model ${modelInfo.id}. Coding plans require providers with prompt caching support; enable access to all models in your dashboard settings at code.llmgateway.io/dashboard to use this model.`,
+ });
+ }
}📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| if (isDevPlanRestricted) { | |
| iamFilteredModelProviders = iamFilteredModelProviders.filter( | |
| providerSupportsCachedInput, | |
| ); | |
| expandedIamFilteredModelProviders = | |
| expandedIamFilteredModelProviders.filter(providerSupportsCachedInput); | |
| } | |
| if (isDevPlanRestricted) { | |
| iamFilteredModelProviders = iamFilteredModelProviders.filter( | |
| providerSupportsCachedInput, | |
| ); | |
| expandedIamFilteredModelProviders = | |
| expandedIamFilteredModelProviders.filter(providerSupportsCachedInput); | |
| if (iamFilteredModelProviders.length === 0) { | |
| throw new HTTPException(403, { | |
| message: `No provider with cached input pricing is available for model ${modelInfo.id}. Coding plans require providers with prompt caching support; enable access to all models in your dashboard settings at code.llmgateway.io/dashboard to use this model.`, | |
| }); | |
| } | |
| } |
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@apps/gateway/src/chat/chat.ts` around lines 1876 - 1882, After applying the
cached-input filter when isDevPlanRestricted, add the same abort/exit logic used
earlier: check iamFilteredModelProviders and expandedIamFilteredModelProviders
for being empty and, if so, stop the request (throw/return an error or set the
same error response) rather than allowing processing to continue with a stale
usedProvider; use the same behavior/response as the earlier check (lines around
1508–1519) so that when providerSupportsCachedInput filtering yields [] the flow
is aborted and no request proceeds with an invalid usedProvider.
| const sameProviderMappings = isDevPlanRestricted | ||
| ? allSameProviderMappings.filter(providerSupportsCachedInput) | ||
| : allSameProviderMappings; | ||
| if (isDevPlanRestricted && sameProviderMappings.length === 0) { | ||
| throw new HTTPException(403, { | ||
| message: `Provider ${usedProvider} does not offer cached input pricing for model ${modelInfo.id}. Coding plans require providers with prompt caching support; choose another provider or enable access to all models in your dashboard settings at code.llmgateway.io/dashboard.`, | ||
| }); | ||
| } |
There was a problem hiding this comment.
Reject the request when no cached mapping also satisfies the final provider constraints.
After sameProviderMappings is narrowed to cached-only entries, filterEligibleModelProviders(...) can legitimately return no matches for the locked region or requested capabilities. In that case this code still falls through, and the later sameProviderRoutingMappings fallback can pick an arbitrary cached region that ignores providerLockedRegions, webSearchTool, response_format, max_tokens, or reasoningEffort.
Suggested fix
const eligibleMappings = filterEligibleModelProviders(
sameProviderRoutingMappings,
{
allProviderVariants: modelInfo.providers,
providerLockedRegions,
webSearchTool,
responseFormatType: response_format?.type,
hasImages,
maxTokens: max_tokens,
reasoningEffort: reasoning_effort,
},
);
+
+ if (isDevPlanRestricted && eligibleMappings.length === 0) {
+ throw new HTTPException(403, {
+ message: lockedRegion
+ ? `Region '${lockedRegion}' for provider ${usedProvider} does not have a cached-input-capable mapping that satisfies this request for model ${modelInfo.id}.`
+ : `Provider ${usedProvider} has no cached-input-capable mapping that satisfies this request for model ${modelInfo.id}.`,
+ });
+ }
if (eligibleMappings.length > 0) {
let selectedMapping = eligibleMappings[0];Also applies to: 1938-1946
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@apps/gateway/src/chat/chat.ts` around lines 1902 - 1909, After narrowing
sameProviderMappings to cached-only entries, ensure you run
filterEligibleModelProviders(...) against that cached list and, if
isDevPlanRestricted is true and the filtered result is empty, throw the 403
HTTPException (same message used for cached-input restriction) instead of
falling through; do the same guard where you later compute
sameProviderRoutingMappings so you don't accidentally pick a cached region that
violates providerLockedRegions, webSearchTool, response_format, max_tokens, or
reasoningEffort.
Summary
devPlan != "none",!devPlanAllowAllModels) gate models viaisCodingModel, but the check passes if any provider mapping hascachedInputPrice— so a request likegroq/gpt-oss-120bslips through and hits an uncached upstream, blowing through credit allowance.cachedInputPrice, and filter post-IAM / auto-routing candidates so canonical mappings only resolve to cached providers.providerSupportsCachedInputhelper plus unit tests (apps/gateway/src/lib/coding-models.spec.ts).Test plan
pnpm vitest run apps/gateway/src/lib/coding-models.spec.ts— 13 new tests passpnpm --filter gateway build— cleangroq/gpt-oss-120breturns 403; canonicalgpt-oss-120broutes to bytedance only🤖 Generated with Claude Code
Summary by CodeRabbit
New Features
Bug Fixes
Tests