Skip to content

feat(devpass): require cached-input pricing for coding plans - #2150

Merged
steebchen merged 2 commits into
mainfrom
devpass-cached-input-only
May 4, 2026
Merged

steebchen merged 2 commits into
mainfrom
devpass-cached-input-only

Conversation

@steebchen

@steebchen steebchen commented May 4, 2026 •

Copy link
Copy Markdown
Member

Summary

  • Coding-plan personal orgs (devPlan != "none", !devPlanAllowAllModels) gate models via isCodingModel, but the check passes if any provider mapping has cachedInputPrice — so a request like groq/gpt-oss-120b slips through and hits an uncached upstream, blowing through credit allowance.
  • Deny specific provider requests whose mapping lacks cachedInputPrice, and filter post-IAM / auto-routing candidates so canonical mappings only resolve to cached providers.
  • Adds a providerSupportsCachedInput helper plus unit tests (apps/gateway/src/lib/coding-models.spec.ts).

Test plan

  • pnpm vitest run apps/gateway/src/lib/coding-models.spec.ts — 13 new tests pass
  • pnpm --filter gateway build — clean
  • Manual: dev-plan org + groq/gpt-oss-120b returns 403; canonical gpt-oss-120b routes to bytedance only

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Stricter coding-plan validation that enforces cached-input pricing support when restricted.
  • Bug Fixes

    • Provider and model selection now consistently filter out providers lacking cached-input support, preventing incompatible routing or direct-selection failures.
  • Tests

    • Added tests covering cached-input support checks and coding-model qualification scenarios.

Coding plans (personal orgs on a dev plan with devPlanAllowAllModels=false)
gate models via isCodingModel, but a request like groq/gpt-oss-120b passes
the model-level check while routing to an uncached mapping. Deny specific
provider requests without cached pricing, and filter routing/auto-routing
candidates so canonical mappings only resolve to cached providers.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings May 4, 2026 08:08
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@coderabbitai

coderabbitai Bot commented May 4, 2026 •

Copy link
Copy Markdown
Contributor

Walkthrough

Adds a cached-input capability check and gates coding-plan/dev-plan enforcement in the gateway. Introduces providerSupportsCachedInput, refactors isCodingModel to use it, and applies cached-input filtering and 403 rejections at multiple provider-selection stages in the chat gateway. Tests for both helpers added.

Changes

Cached-Input Provider Filtering for Development Plans

Layer / File(s) Summary
Data Shape / Helper
apps/gateway/src/lib/coding-models.ts
Adds exported providerSupportsCachedInput(p: Pick<ProviderModelMapping,"cachedInputPrice">): boolean and refactors isCodingModel to call it.
Gateway Restriction Flag & Early Rejection
apps/gateway/src/chat/chat.ts
Imports providerSupportsCachedInput; introduces isDevPlanRestricted gating. When enabled, rejects requests (403) if model fails isCodingModel or if a specifically requested provider has mappings but none support cached input.
IAM-filtering / Candidate Lists
apps/gateway/src/chat/chat.ts
After IAM filtering, iamFilteredModelProviders and expandedIamFilteredModelProviders are filtered to mappings where providerSupportsCachedInput is true; request fails (403) if empty.
Auto-routing / Suitable Providers
apps/gateway/src/chat/chat.ts
During auto-routing, availableModelProviders -> candidate set is filtered to cached-input-supporting mappings when isDevPlanRestricted is true; selection uses that filtered set.
Post-selection IAM Revalidation
apps/gateway/src/chat/chat.ts
After auto-routing resolves a concrete provider/model and IAM is revalidated, IAM-filtered provider lists are again constrained to cached-input supporters under the restriction.
Direct Provider Selection & Region Locking
apps/gateway/src/chat/chat.ts
For direct provider selection, filters sameProviderMappings to cached-input-supporting mappings when restricted; rejects if none qualify. When a region is locked, enforces that at least one cached-input-supporting mapping matches the region.
Tests
apps/gateway/src/lib/coding-models.spec.ts
Adds Vitest tests for providerSupportsCachedInput (positive, zero, undefined, null cases) and isCodingModel (model-level flags and various provider capability permutations).

Sequence Diagram(s)

sequenceDiagram
  autonumber
  participant Client
  participant Gateway
  participant IAM
  participant ProviderRegistry
  Client->>Gateway: Send chat request (model, optional provider mapping, region)
  Gateway->>ProviderRegistry: Lookup model/provider mappings
  Gateway->>IAM: Validate caller permissions for provider/model
  alt isDevPlanRestricted
    Gateway->>ProviderRegistry: Filter mappings where providerSupportsCachedInput == true
    ProviderRegistry-->>Gateway: Filtered provider list
    Gateway->>Gateway: If no providers remain → 403 return
  end
  opt Auto-routing
    Gateway->>ProviderRegistry: Select suitable provider from filtered candidates
  end
  Gateway->>IAM: Revalidate IAM for chosen provider/model
  Gateway->>Provider: Forward request (or 403 if checks fail)
  Provider-->>Gateway: Response
  Gateway-->>Client: Return response or 403
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Possibly related PRs

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 50.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately describes the main change: adding a requirement for cached-input pricing support when using coding plans in dev environments.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch devpass-cached-input-only

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share
Review rate limit: 7/8 reviews remaining, refill in 7 minutes and 30 seconds.

Comment @coderabbitai help to get the list of available commands and usage tips.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR tightens “coding plan” model/provider eligibility in the gateway to ensure dev-plan personal orgs only route to provider mappings that have cached-input pricing, preventing requests from slipping to uncached upstream mappings and consuming credits unexpectedly.

Changes:

  • Adds a reusable providerSupportsCachedInput helper and uses it in isCodingModel.
  • Enforces cached-input pricing at request time in chat.ts for dev-plan-restricted orgs (both for explicitly requested providers and for routed candidates).
  • Adds unit tests for providerSupportsCachedInput and isCodingModel.

Reviewed changes

Copilot reviewed 3 out of 3 changed files in this pull request and generated 2 comments.

File Description
apps/gateway/src/lib/coding-models.ts Adds providerSupportsCachedInput helper and reuses it in isCodingModel.
apps/gateway/src/lib/coding-models.spec.ts Adds unit tests for cached-input detection and coding-model qualification.
apps/gateway/src/chat/chat.ts Applies dev-plan cached-input enforcement for explicit provider requests and routing candidate filtering.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

import type { ModelDefinition, ProviderModelMapping } from "@llmgateway/models";

export function providerSupportsCachedInput(
p: Pick<ProviderModelMapping, "cachedInputPrice">,
Comment on lines +38 to +43
it("returns false when cachedInputPrice is null", () => {
expect(
providerSupportsCachedInput({
cachedInputPrice: null as unknown as undefined,
}),
).toBe(false);

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@apps/gateway/src/chat/chat.ts`:
- Around line 1431-1449: The current check only ensures there exists some
cached-input mapping for requestedProvider but not that the specific mapping
chosen will support cached input; update the region-aware provider selection to
apply the providerSupportsCachedInput filter (or else validate the exact
resolved mapping) so the final candidate set built from modelInfo.providers
excludes region-specific entries lacking cached input. Concretely, when
resolving provider mappings for requestedProvider and requestedRegion (the code
that builds requestedProviderMappings and the later direct-provider resolution),
filter by providerSupportsCachedInput(p) and/or re-check the selected mapping
against providerSupportsCachedInput before proceeding, ensuring the chosen
provider+region pair is confirmed cached-input-capable.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro

Run ID: 3f8da705-6d19-4a20-acbe-f2c00f95b310

📥 Commits

Reviewing files that changed from the base of the PR and between 9f74835 and c629910.

📒 Files selected for processing (3)
  • apps/gateway/src/chat/chat.ts
  • apps/gateway/src/lib/coding-models.spec.ts
  • apps/gateway/src/lib/coding-models.ts

Comment thread apps/gateway/src/chat/chat.ts
The direct-provider region picker built candidates straight from
modelInfo.providers, so a provider-key-locked uncached region could still
resolve to an uncached upstream after the model-level gate passed. Filter
sameProviderMappings by cached pricing for restricted dev plans, and 403
when a locked region has no cached mapping.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
@steebchen
steebchen enabled auto-merge May 4, 2026 11:06

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@apps/gateway/src/chat/chat.ts`:
- Around line 1902-1909: After narrowing sameProviderMappings to cached-only
entries, ensure you run filterEligibleModelProviders(...) against that cached
list and, if isDevPlanRestricted is true and the filtered result is empty, throw
the 403 HTTPException (same message used for cached-input restriction) instead
of falling through; do the same guard where you later compute
sameProviderRoutingMappings so you don't accidentally pick a cached region that
violates providerLockedRegions, webSearchTool, response_format, max_tokens, or
reasoningEffort.
- Around line 1876-1882: After applying the cached-input filter when
isDevPlanRestricted, add the same abort/exit logic used earlier: check
iamFilteredModelProviders and expandedIamFilteredModelProviders for being empty
and, if so, stop the request (throw/return an error or set the same error
response) rather than allowing processing to continue with a stale usedProvider;
use the same behavior/response as the earlier check (lines around 1508–1519) so
that when providerSupportsCachedInput filtering yields [] the flow is aborted
and no request proceeds with an invalid usedProvider.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro

Run ID: 2542e843-15f6-4241-acd1-3d9b70b98bb5

📥 Commits

Reviewing files that changed from the base of the PR and between c629910 and 43d11f2.

📒 Files selected for processing (1)
  • apps/gateway/src/chat/chat.ts

Comment on lines +1876 to +1882
if (isDevPlanRestricted) {
iamFilteredModelProviders = iamFilteredModelProviders.filter(
providerSupportsCachedInput,
);
expandedIamFilteredModelProviders =
expandedIamFilteredModelProviders.filter(providerSupportsCachedInput);
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major | ⚡ Quick win

Abort when post-auto filtering removes every cached-capable provider.

This block re-filters the resolved model after IAM, but unlike the earlier check at Lines 1508-1519 it never stops on []. If auto-routing fell back to the default claude-haiku-4-5/anthropic path, usedProvider stays set and the request can continue even though no cached-input-capable provider is actually allowed for the resolved model.

Suggested fix
 		if (isDevPlanRestricted) {
 			iamFilteredModelProviders = iamFilteredModelProviders.filter(
 				providerSupportsCachedInput,
 			);
 			expandedIamFilteredModelProviders =
 				expandedIamFilteredModelProviders.filter(providerSupportsCachedInput);
+			if (iamFilteredModelProviders.length === 0) {
+				throw new HTTPException(403, {
+					message: `No provider with cached input pricing is available for model ${modelInfo.id}. Coding plans require providers with prompt caching support; enable access to all models in your dashboard settings at code.llmgateway.io/dashboard to use this model.`,
+				});
+			}
 		}
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
if (isDevPlanRestricted) {
iamFilteredModelProviders = iamFilteredModelProviders.filter(
providerSupportsCachedInput,
);
expandedIamFilteredModelProviders =
expandedIamFilteredModelProviders.filter(providerSupportsCachedInput);
}
if (isDevPlanRestricted) {
iamFilteredModelProviders = iamFilteredModelProviders.filter(
providerSupportsCachedInput,
);
expandedIamFilteredModelProviders =
expandedIamFilteredModelProviders.filter(providerSupportsCachedInput);
if (iamFilteredModelProviders.length === 0) {
throw new HTTPException(403, {
message: `No provider with cached input pricing is available for model ${modelInfo.id}. Coding plans require providers with prompt caching support; enable access to all models in your dashboard settings at code.llmgateway.io/dashboard to use this model.`,
});
}
}
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@apps/gateway/src/chat/chat.ts` around lines 1876 - 1882, After applying the
cached-input filter when isDevPlanRestricted, add the same abort/exit logic used
earlier: check iamFilteredModelProviders and expandedIamFilteredModelProviders
for being empty and, if so, stop the request (throw/return an error or set the
same error response) rather than allowing processing to continue with a stale
usedProvider; use the same behavior/response as the earlier check (lines around
1508–1519) so that when providerSupportsCachedInput filtering yields [] the flow
is aborted and no request proceeds with an invalid usedProvider.

Comment on lines +1902 to +1909
const sameProviderMappings = isDevPlanRestricted
? allSameProviderMappings.filter(providerSupportsCachedInput)
: allSameProviderMappings;
if (isDevPlanRestricted && sameProviderMappings.length === 0) {
throw new HTTPException(403, {
message: `Provider ${usedProvider} does not offer cached input pricing for model ${modelInfo.id}. Coding plans require providers with prompt caching support; choose another provider or enable access to all models in your dashboard settings at code.llmgateway.io/dashboard.`,
});
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major | ⚡ Quick win

Reject the request when no cached mapping also satisfies the final provider constraints.

After sameProviderMappings is narrowed to cached-only entries, filterEligibleModelProviders(...) can legitimately return no matches for the locked region or requested capabilities. In that case this code still falls through, and the later sameProviderRoutingMappings fallback can pick an arbitrary cached region that ignores providerLockedRegions, webSearchTool, response_format, max_tokens, or reasoningEffort.

Suggested fix
 			const eligibleMappings = filterEligibleModelProviders(
 				sameProviderRoutingMappings,
 				{
 					allProviderVariants: modelInfo.providers,
 					providerLockedRegions,
 					webSearchTool,
 					responseFormatType: response_format?.type,
 					hasImages,
 					maxTokens: max_tokens,
 					reasoningEffort: reasoning_effort,
 				},
 			);
+
+			if (isDevPlanRestricted && eligibleMappings.length === 0) {
+				throw new HTTPException(403, {
+					message: lockedRegion
+						? `Region '${lockedRegion}' for provider ${usedProvider} does not have a cached-input-capable mapping that satisfies this request for model ${modelInfo.id}.`
+						: `Provider ${usedProvider} has no cached-input-capable mapping that satisfies this request for model ${modelInfo.id}.`,
+				});
+			}
 
 			if (eligibleMappings.length > 0) {
 				let selectedMapping = eligibleMappings[0];

Also applies to: 1938-1946

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@apps/gateway/src/chat/chat.ts` around lines 1902 - 1909, After narrowing
sameProviderMappings to cached-only entries, ensure you run
filterEligibleModelProviders(...) against that cached list and, if
isDevPlanRestricted is true and the filtered result is empty, throw the 403
HTTPException (same message used for cached-input restriction) instead of
falling through; do the same guard where you later compute
sameProviderRoutingMappings so you don't accidentally pick a cached region that
violates providerLockedRegions, webSearchTool, response_format, max_tokens, or
reasoningEffort.

@steebchen
steebchen added this pull request to the merge queue May 4, 2026
Merged via the queue into main with commit 337a298 May 4, 2026
17 checks passed
@steebchen
steebchen deleted the devpass-cached-input-only branch May 4, 2026 11:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants