Skip to content

feat(gateway): gate routes by model output capability - #2828

Merged
steebchen merged 3 commits into
mainfrom
filter-models-by-output-capability
Jun 25, 2026
Merged

steebchen merged 3 commits into
mainfrom
filter-models-by-output-capability

Conversation

@steebchen

@steebchen steebchen commented Jun 25, 2026 •

Copy link
Copy Markdown
Member

Summary

Distinguishes model types by their declared output capability instead of ad-hoc, per-route signals, and enforces it uniformly across routes.

Background: OCR models (e.g. mistral-ocr-latest) genuinely return text, so they can't be told apart from chat models by output modality alone — that's why they previously needed a special-case check. Rather than a pricing heuristic, this PR makes the model's output capability the single source of truth: OCR becomes a first-class output type, and every endpoint checks it.

Changes

  • output: ["ocr"] — added "ocr" to the ModelDefinition.output union and tagged mistral-ocr-latest, so OCR is a declared output capability rather than something only the per-provider ocr flag knows about.
  • Shared validateModelOutput helper (apps/gateway/src/lib/) — rejects (400) a model whose declared outputs don't intersect what the endpoint serves, and points the caller at the right endpoint (e.g. "Model X is an OCR model … Use the /v1/ocr endpoint instead.").
  • Chat — validateModelCapabilities now accepts only text/image output, replacing the embedding + OCR special cases. Image output is allowed because image generation routes through /v1/chat/completions (so text-only, image-only, and text+image all pass). This also closes the gap that previously let video / audio / image-only models through. /v1/responses and /v1/messages forward to chat, so they inherit this.
  • Images — guards that the requested model actually produces image output before forwarding to chat completions (auto/custom and unknown models pass through as before).
  • /v1/models — surfaces output_modalities 1:1 with the model catalog, including ["ocr"] (added to the public schema enum), so third-party clients can reference the same modality taxonomy. Per-page pricing (ocr_page) remains alongside.

Routes that resolve models through a per-mapping capability flag (embeddings, speechGenerations, videoGenerations, ocr) already reject mismatched models at resolution, so they were left as-is.

Defaults / safety

A missing output defaults to ["text"], so the ~290 existing chat models need no change — only non-text models opt in. A new model-metadata invariant test asserts that any model carrying a non-text capability flag (imageGenerations/embeddings/speechGenerations/videoGenerations/ocr) declares the matching output, so a non-text model can't silently default to text and get wrongly accepted on chat. (This invariant would have caught the original OCR bug.)

Tests

  • New validateModelCapabilities - output capability cases: rejects OCR/video/audio on chat, allows image-output models.
  • model-metadata invariant: output ⇄ capability-flag consistency.
  • /v1/models spec asserts mistral-ocr-latest → output_modalities: ["ocr"].
  • Existing OCR, models, and image API specs pass; pnpm build and pnpm format are green.

🤖 Generated with Claude Code

Distinguish model types by their declared `output` capability instead of
ad-hoc, per-route signals, and enforce it uniformly.

- Add `"ocr"` to the model `output` union and tag `mistral-ocr-latest`
  with `output: ["ocr"]`, so OCR is a first-class output capability
  rather than something only the per-provider `ocr` flag knows about.
- New shared `validateModelOutput` helper: rejects (400) a model whose
  declared outputs don't intersect what the endpoint serves, pointing the
  caller at the right endpoint.
- Chat (and the `/v1/responses` + `/v1/messages` routes that forward to
  it) now accept only text/image output, replacing the embedding+OCR
  special cases. This also closes the gap that let video/audio/image-only
  models through.
- Images route guards that the requested model actually produces image
  output before forwarding to chat completions.
- `/v1/models` surfaces OCR models as `output_modalities: ["text"]`
  (they return text; per-page pricing already distinguishes them).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Jun 25, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

Walkthrough

Adds OCR as a model output, introduces shared output-capability validation, maps OCR to text in the models API response, and applies the validator to chat completions and image generation/edit request flows.

Changes

Model output capability routing

Layer / File(s) Summary
Output capability contract
packages/models/src/models.ts, packages/models/src/models/mistral.ts, apps/gateway/src/lib/validate-model-output.ts
ModelDefinition.output accepts ocr, mistral-ocr-latest declares ["ocr"], and validateModelOutput derives accepted outputs and raises endpoint-specific HTTP 400 errors.
Public OCR response mapping
apps/gateway/src/models/models.ts
The models API maps ocr output modalities to text in architecture.output_modalities while keeping the public modality union unchanged.
Chat completions validation
apps/gateway/src/chat/tools/validate-model-capabilities.ts, apps/gateway/src/chat/tools/validate-model-capabilities.spec.ts
Chat completions now call validateModelOutput(..., ["text", "image"]), and the spec adds OCR/video/audio rejects plus image-output allow cases.
Image request validation
apps/gateway/src/images/images.ts
The image generation and edit flows add assertImageModel, resolve model metadata, and reject non-image models before forwarding requests.
Provider metadata consistency
packages/models/src/model-metadata.spec.ts
The model metadata test maps provider capability flags to required outputs and asserts that all enabled non-text capabilities align with each model’s declared output value.

Estimated review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly related PRs

  • theopenco/llmgateway#2230: Both PRs change apps/gateway/src/chat/tools/validate-model-capabilities.ts to alter how image/vision-capable models are accepted or rejected for chat completions.
  • theopenco/llmgateway#2793: Both PRs add early chat-completions validation that rejects models whose output is incompatible with the chat endpoint.
  • theopenco/llmgateway#2818: Both PRs add OCR-specific rejection for /v1/chat/completions, with this PR generalizing that logic through validateModelOutput.
🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Title check ✅ Passed The title accurately summarizes the main change: gateway routes are now gated by model output capability.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch filter-models-by-output-capability

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Guards against a non-text model (embeddings, OCR, speech, video, image
gen) being added without a matching `output`, which chat completions
would wrongly accept since a missing `output` defaults to ["text"].

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 816b49c617

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

// handler to reject as "model not found".
function assertImageModel(model: string): void {
const slashIdx = model.indexOf("/");
const modelKey = slashIdx > 0 ? model.slice(slashIdx + 1) : model;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Mirror chat model parsing before image gating

When callers use provider-qualified names, this local parser does not mirror the chat parser: it leaves region suffixes in modelKey and treats unknown provider prefixes as built-in catalog lookups. For example, alibaba/qwen-plus:cn-beijing will not match the catalog and bypasses the new non-image check entirely, while a custom-provider request like mycompany/gpt-4o can be rejected against the built-in text-only gpt-4o before the custom provider path runs. Normalize with the same provider/model/region rules (and skip custom prefixes) before looking up models.

Useful? React with 👍 / 👎.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@apps/gateway/src/images/images.ts`:
- Around line 563-576: `assertImageModel()` only skips the bare `custom`
sentinel, so `custom/...` still reaches the `models.find(...)` catalog lookup
and may be misclassified. Update `assertImageModel` to detect and short-circuit
any custom-provider image model before the catalog search, alongside the
existing `auto` and `custom` handling, so `validateModelOutput(...)` is not
applied to `custom/<model>` values.

In `@apps/gateway/src/models/models.ts`:
- Around line 204-216: The `/v1/models` output fallback in `models.ts` is
inconsistent with router validation because `outputModalities` only defaults
when `model.output` is undefined, while `getModelOutputs()` also treats an empty
array as text. Update the `outputModalities` mapping in the models list logic to
fall back to ["text"] for both undefined and empty `model.output`, keeping the
advertised outputs aligned with routing behavior and preserving the OCR-to-text
normalization.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro

Run ID: f5e4ed42-7bf0-4a7e-9ddc-d8efb0e36c8e

📥 Commits

Reviewing files that changed from the base of the PR and between 975a594 and 816b49c.

📒 Files selected for processing (7)
  • apps/gateway/src/chat/tools/validate-model-capabilities.spec.ts
  • apps/gateway/src/chat/tools/validate-model-capabilities.ts
  • apps/gateway/src/images/images.ts
  • apps/gateway/src/lib/validate-model-output.ts
  • apps/gateway/src/models/models.ts
  • packages/models/src/models.ts
  • packages/models/src/models/mistral.ts

Comment on lines +563 to +576
function assertImageModel(model: string): void {
const slashIdx = model.indexOf("/");
const modelKey = slashIdx > 0 ? model.slice(slashIdx + 1) : model;
if (modelKey === "auto" || modelKey === "custom") {
return;
}
const modelInfo = models.find(
(m) =>
m.id === model ||
m.id === modelKey ||
m.providers.some((p) => p.externalId === modelKey),
);
if (modelInfo) {
validateModelOutput(modelInfo, modelKey, ["image"]);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟡 Minor

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
set -euo pipefail

# Find how custom providers/models are represented and parsed.
rg -n --type=ts 'requestedProvider\s*===\s*"custom"|["'\'']custom/|indexOf\("/"\)|slice\(slashIdx \+ 1\)|split\("/"\)' apps packages

# Compare the image-route guard with the chat-route custom-provider bypass.
rg -n -C3 --type=ts 'assertImageModel|validateModelCapabilities|requestedProvider === "custom"' apps/gateway packages/models

Repository: theopenco/llmgateway

Length of output: 25369


🏁 Script executed:

#!/bin/bash
set -euo pipefail

# Inspect the image route and how request.model is parsed/resolved.
ast-grep outline apps/gateway/src/images/images.ts --view expanded

# Show the relevant image-route sections around parsing and the assert call.
sed -n '620,700p' apps/gateway/src/images/images.ts
sed -n '1060,1105p' apps/gateway/src/images/images.ts

# Inspect chat parsing for custom-provider model syntax as a comparison point.
sed -n '1,120p' apps/gateway/src/chat/tools/parse-model-input.ts
sed -n '1,120p' apps/gateway/src/chat/tools/resolve-model-info.ts

Repository: theopenco/llmgateway

Length of output: 13316


🏁 Script executed:

#!/bin/bash
set -euo pipefail

# Inspect the image route and how request.model is parsed/resolved.
ast-grep outline apps/gateway/src/images/images.ts --view expanded

# Show the relevant image-route sections around parsing and the assert call.
sed -n '620,700p' apps/gateway/src/images/images.ts
sed -n '1060,1105p' apps/gateway/src/images/images.ts

# Inspect chat parsing for custom-provider model syntax as a comparison point.
sed -n '1,120p' apps/gateway/src/chat/tools/parse-model-input.ts
sed -n '1,120p' apps/gateway/src/chat/tools/resolve-model-info.ts

Repository: theopenco/llmgateway

Length of output: 13316


🏁 Script executed:

#!/bin/bash
set -euo pipefail

# Inspect how the image route resolves/forwards the requested model.
sed -n '340,390p' apps/gateway/src/images/images.ts
sed -n '570,620p' apps/gateway/src/images/images.ts

# Find any parsing or special handling of custom/provider-scoped models in images routes.
rg -n --type=ts 'parseModelInput|resolveImageRequestModel|requestedProvider === "custom"|custom/' apps/gateway/src/images apps/gateway/src/chat

Repository: theopenco/llmgateway

Length of output: 4037


🏁 Script executed:

#!/bin/bash
set -euo pipefail

# Inspect how the image route resolves/forwards the requested model.
sed -n '340,390p' apps/gateway/src/images/images.ts
sed -n '570,620p' apps/gateway/src/images/images.ts

# Find any parsing or special handling of custom/provider-scoped models in images routes.
rg -n --type=ts 'parseModelInput|resolveImageRequestModel|requestedProvider === "custom"|custom/' apps/gateway/src/images apps/gateway/src/chat

Repository: theopenco/llmgateway

Length of output: 4037


🏁 Script executed:

#!/bin/bash
set -euo pipefail

# Inspect the image request schema and route docs/comments to see what model strings are allowed.
sed -n '1,120p' apps/gateway/src/images/images.ts
rg -n --type=ts --type=md 'images\.generations|imageGenerationsRequestSchema|request\.model|model:' apps/gateway/src/images apps/gateway/src -g '!**/*.spec.ts'

Repository: theopenco/llmgateway

Length of output: 37981


🏁 Script executed:

#!/bin/bash
set -euo pipefail

# Inspect the image request schema and route docs/comments to see what model strings are allowed.
sed -n '1,120p' apps/gateway/src/images/images.ts
rg -n --type=ts --type=md 'images\.generations|imageGenerationsRequestSchema|request\.model|model:' apps/gateway/src/images apps/gateway/src -g '!**/*.spec.ts'

Repository: theopenco/llmgateway

Length of output: 37981


Skip custom/... before catalog lookup apps/gateway/src/images/images.ts:563-576
assertImageModel() only bypasses bare custom, so custom/<model> still falls through to models.find(...) and can be classified using unrelated catalog metadata. Short-circuit custom-provider image models first.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@apps/gateway/src/images/images.ts` around lines 563 - 576,
`assertImageModel()` only skips the bare `custom` sentinel, so `custom/...`
still reaches the `models.find(...)` catalog lookup and may be misclassified.
Update `assertImageModel` to detect and short-circuit any custom-provider image
model before the catalog search, alongside the existing `auto` and `custom`
handling, so `validateModelOutput(...)` is not applied to `custom/<model>`
values.

Comment thread apps/gateway/src/models/models.ts Outdated
Mirror the model catalog 1:1 instead of collapsing OCR to "text", so
third-party clients see output_modalities: ["ocr"] and can reference the
same taxonomy. Adds "ocr" to the public schema enum.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
packages/models/src/model-metadata.spec.ts (1)

14-23: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Preserve the ModelDefinition.output union here.

REQUIRED_OUTPUT_BY_FLAG.output and outputs are both widened to string, so a typo in this contract test would still compile and quietly stop validating the intended modality. Keep them typed from ModelDefinition["output"] so output-vocabulary drift fails at compile time.

Suggested change
+type ModelOutput = NonNullable<ModelDefinition["output"]>[number];
+
 const REQUIRED_OUTPUT_BY_FLAG: {
 	flag: keyof ProviderModelMapping;
-	output: string;
+	output: ModelOutput;
 }[] = [
 	{ flag: "imageGenerations", output: "image" },
 	{ flag: "embeddings", output: "embedding" },
 	{ flag: "speechGenerations", output: "audio" },
 	{ flag: "videoGenerations", output: "video" },
 	{ flag: "ocr", output: "ocr" },
 ];
...
-			const outputs: string[] = model.output ?? ["text"];
+			const outputs: readonly ModelOutput[] = model.output ?? ["text"];

Also applies to: 63-63

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@packages/models/src/model-metadata.spec.ts` around lines 14 - 23, Keep the
contract test tied to the ModelDefinition.output union instead of widening
REQUIRED_OUTPUT_BY_FLAG.output and outputs to string; update the
REQUIRED_OUTPUT_BY_FLAG declaration and the related outputs usage in
model-metadata.spec to reference ModelDefinition["output"] so any typo or new
modality mismatch fails at compile time. Locate the check by the
REQUIRED_OUTPUT_BY_FLAG constant and the outputs assertion, and preserve the
existing modality mapping while tightening the types to the source-of-truth
union.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Nitpick comments:
In `@packages/models/src/model-metadata.spec.ts`:
- Around line 14-23: Keep the contract test tied to the ModelDefinition.output
union instead of widening REQUIRED_OUTPUT_BY_FLAG.output and outputs to string;
update the REQUIRED_OUTPUT_BY_FLAG declaration and the related outputs usage in
model-metadata.spec to reference ModelDefinition["output"] so any typo or new
modality mismatch fails at compile time. Locate the check by the
REQUIRED_OUTPUT_BY_FLAG constant and the outputs assertion, and preserve the
existing modality mapping while tightening the types to the source-of-truth
union.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro

Run ID: 83ccd9da-5699-4c5f-b3f3-0cb44cbd60ee

📥 Commits

Reviewing files that changed from the base of the PR and between 816b49c and 2179ff1.

📒 Files selected for processing (1)
  • packages/models/src/model-metadata.spec.ts

@steebchen
steebchen merged commit f193f2b into main Jun 25, 2026
17 checks passed
@steebchen
steebchen deleted the filter-models-by-output-capability branch June 25, 2026 20:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant