Skip to content

fix: stop imposing catalog assumptions on custom providers - #2667

Merged
steebchen merged 3 commits into
mainfrom
fix/custom-provider-max-tokens-cap
Jun 13, 2026
Merged

steebchen merged 3 commits into
mainfrom
fix/custom-provider-max-tokens-cap

Conversation

@steebchen

@steebchen steebchen commented Jun 13, 2026 •

Copy link
Copy Markdown
Member

Problem

Requests to a custom provider were rejected when max_tokens exceeded 4096, even when the underlying model supports far more:

{"error":{"message":"The requested max_tokens (32000) exceeds the maximum output tokens allowed for model qwen3.6-plus (4096)","type":"invalid_request_error"}}

Custom providers have no entry in the model catalog, so the gateway synthesizes a mock ModelDefinition. That mock hardcoded placeholder values it cannot actually know — maxOutput: 4096, contextSize: 8192, vision: false, jsonOutput: true. These guesses both imposed false limits (the maxOutput cap above) and gated capability rejections (vision/JSON/tools) on values the gateway has no way to determine.

Fix

Treat custom providers as fully opaque — the upstream provider is the authority on limits and capabilities:

  • Drop the guessed fields (maxOutput, contextSize, vision, jsonOutput) from the synthesized custom-provider mapping in both places it is built (resolve-model-info.ts and chat.ts). The max_tokens validation guards on maxOutput !== undefined, so it is now skipped for custom providers; contextSize only feeds auto-routing (which never applies to a pinned custom provider) and has a fallback.
  • Skip capability validation for custom providers in validateModelCapabilities (early return when requestedProvider === "custom"), consistent with how it already skips the bare auto/custom model strings. This covers vision, documents, JSON output, JSON schema, reasoning, and tools — none of which the gateway can know for a custom endpoint.

streaming: true is kept only because the type requires it; it is never read for custom providers (streaming support is computed from the catalog via getModelStreamingSupport, which returns null for a non-catalog model and therefore never rejects). Pricing stays "0" (BYOK, not billed by the gateway).

Tests

  • chat-custom-provider.e2e.ts: added regression tests for max_tokens: 32000 and response_format: json_object against a custom provider, both asserting 200. Full suite (8 tests) passes.
  • validate-model-capabilities.spec.ts: added a test asserting all capability checks are skipped when the provider is custom (12 tests pass).

pnpm format and pnpm build are green.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • Bug Fixes

    • Gateway no longer imposes artificial limits or capability checks on custom providers; requested max_tokens and provider capabilities are honored/handled upstream. Streaming support preserved.
  • Tests

    • Added E2E tests verifying large max_tokens and JSON-object response_format work with custom providers.
    • Added unit tests ensuring capability validation is skipped for custom providers.

Custom providers synthesized a mock model definition with a hardcoded
maxOutput of 4096, which made the gateway reject requests whose
max_tokens exceeded 4096 even when the upstream model supports far more.
The gateway has no catalog knowledge of a custom model's real limits, so
leave maxOutput uncapped and let the upstream provider enforce it.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Jun 13, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

Caution

Review failed

The pull request is closed.

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro

Run ID: 544c12da-9448-4542-af44-629c4bb5d2b2

📥 Commits

Reviewing files that changed from the base of the PR and between 7ab36c7 and 40d3424.

📒 Files selected for processing (5)
  • apps/gateway/src/chat-custom-provider.e2e.ts
  • apps/gateway/src/chat/chat.ts
  • apps/gateway/src/chat/tools/resolve-model-info.ts
  • apps/gateway/src/chat/tools/validate-model-capabilities.spec.ts
  • apps/gateway/src/chat/tools/validate-model-capabilities.ts

Walkthrough

Skip synthesizing catalog-derived limits and capability validation for custom providers; retain only type-required streaming in mocked modelInfo. Add tests verifying a large max_tokens (32000) and response_format: { type: "json_object" } are accepted and return mock completions (HTTP 200).

Changes

Custom provider handling

Layer / File(s) Summary
Mock modelInfo for custom providers
apps/gateway/src/chat/tools/resolve-model-info.ts
Remove hard-coded contextSize, maxOutput, and capability flags from the mocked modelInfo for requestedProvider === "custom"; leave streaming only to satisfy the type and document that upstream enforces limits.
Skip capability validation for custom providers
apps/gateway/src/chat/tools/validate-model-capabilities.ts, apps/gateway/src/chat/tools/validate-model-capabilities.spec.ts
Add an early-return when requestedProvider === "custom" so validateModelCapabilities performs no checks; add a test ensuring mixed capability fields do not cause a throw for custom providers.
Synthesize finalModelInfo and E2E tests
apps/gateway/src/chat/chat.ts, apps/gateway/src/chat-custom-provider.e2e.ts
Stop injecting catalog-derived contextSize, maxOutput, and vision into the synthesized finalModelInfo for usedProvider === "custom" (retain streaming only). Add E2E tests: POST /v1/chat/completions with max_tokens: 32000 and with response_format: { type: "json_object" }, asserting HTTP 200 and expected mock completion content.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

  • theopenco/llmgateway#2230: Overlaps in capability-validation logic and prior changes to vision/hasImages checks affecting custom/auto bypass.

Suggested reviewers

  • smakosh
  • proxysoul
🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately summarizes the main change: removing hardcoded capability constraints for custom providers in the gateway.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/custom-provider-max-tokens-cap

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@apps/gateway/src/chat-custom-provider.e2e.ts`:
- Around line 200-221: The test "should not cap max_tokens for custom providers"
currently only checks for a 200 and response body; update it to assert the
upstream request actually received max_tokens: 32000 so the value isn't being
rewritten/omitted. After sending the POST to "/v1/chat/completions" (the
app.request call that sets max_tokens: 32000), read the mock upstream's captured
request (e.g., the recorded request array or spy used by your test harness) and
add an expectation that the parsed upstream request body has max_tokens ===
32000; keep the existing response assertions intact.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro

Run ID: 4334ae5e-8a51-4ac7-a622-ed6870bd73cc

📥 Commits

Reviewing files that changed from the base of the PR and between 7c466c9 and db94d82.

📒 Files selected for processing (3)
  • apps/gateway/src/chat-custom-provider.e2e.ts
  • apps/gateway/src/chat/chat.ts
  • apps/gateway/src/chat/tools/resolve-model-info.ts

Comment on lines +200 to +221
test("should not cap max_tokens for custom providers", async () => {
await setupTestData({ mode: "api-keys", includeProviderKey: true });

const res = await app.request("/v1/chat/completions", {
method: "POST",
headers: {
"Content-Type": "application/json",
Authorization: "Bearer real-token",
},
body: JSON.stringify({
model: "my-custom/qwen3.6-plus",
max_tokens: 32000,
messages: [{ role: "user", content: "hello" }],
}),
});

const json = await res.json();
expect(res.status).toBe(200);
expect(json.choices[0].message.content).toBe(
"Hello from custom provider!",
);
});

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor | ⚡ Quick win

This regression test doesn’t verify the token value is actually uncapped/forwarded.

Right now it only asserts a 200 response. Since the mock server always returns success, this can still pass if max_tokens is silently rewritten or omitted before upstream dispatch. Add an assertion on the received upstream request body (e.g., max_tokens === 32000).

Suggested test hardening
+let lastCustomProviderRequestBody: unknown = null;
+
 mockServer.post("/v1/chat/completions", async (c) => {
+	lastCustomProviderRequestBody = await c.req.json();
 	return c.json({
 		id: "chatcmpl-mock-custom",
 		object: "chat.completion",
 		created: Math.floor(Date.now() / 1000),
@@
 		test("should not cap max_tokens for custom providers", async () => {
 			await setupTestData({ mode: "api-keys", includeProviderKey: true });
@@
 			const json = await res.json();
 			expect(res.status).toBe(200);
+			expect((lastCustomProviderRequestBody as { max_tokens?: number }).max_tokens).toBe(32000);
 			expect(json.choices[0].message.content).toBe(
 				"Hello from custom provider!",
 			);
 		});
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@apps/gateway/src/chat-custom-provider.e2e.ts` around lines 200 - 221, The
test "should not cap max_tokens for custom providers" currently only checks for
a 200 and response body; update it to assert the upstream request actually
received max_tokens: 32000 so the value isn't being rewritten/omitted. After
sending the POST to "/v1/chat/completions" (the app.request call that sets
max_tokens: 32000), read the mock upstream's captured request (e.g., the
recorded request array or spy used by your test harness) and add an expectation
that the parsed upstream request body has max_tokens === 32000; keep the
existing response assertions intact.

contextSize, maxOutput, and vision cannot be known for a custom provider
since it has no catalog entry. Leave them unset instead of guessing
placeholder values; the upstream provider enforces its own limits and
capabilities.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@steebchen steebchen changed the title fix: do not cap max_tokens for custom providers fix: stop hardcoding unknown custom provider capabilities Jun 13, 2026
Custom providers have no catalog entry, so the gateway cannot know their
capabilities (vision, jsonOutput, etc). The synthesized model definition
previously guessed these flags, which both asserted false limits and
risked rejecting valid requests. Skip capability validation entirely for
custom providers and let the upstream provider be the authority. Drop the
now-unused jsonOutput flag from the synthesized mapping.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@steebchen steebchen changed the title fix: stop hardcoding unknown custom provider capabilities fix: stop imposing catalog assumptions on custom providers Jun 13, 2026
@steebchen
steebchen merged commit 27ef11e into main Jun 13, 2026
11 of 12 checks passed
@steebchen
steebchen deleted the fix/custom-provider-max-tokens-cap branch June 13, 2026 10:56
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant