Skip to content

fix(gateway): reject TTS/video models in chat - #2793

Merged
steebchen merged 1 commit into
mainfrom
fix/reject-tts-models-in-chat
Jun 22, 2026
Merged

steebchen merged 1 commit into
mainfrom
fix/reject-tts-models-in-chat

Conversation

@steebchen

@steebchen steebchen commented Jun 22, 2026 •

Copy link
Copy Markdown
Member

What

The production error log showed:

Error: Could not use provider: elevenlabs. Provider elevenlabs requires a baseUrl
    at <anonymous> (/app/src/chat/chat.ts:5085:9)

ElevenLabs (and OpenAI TTS) models are speech-only (output: ["audio"], speechGenerations: true) and are served by the dedicated /v1/audio/speech endpoint, which has its own provider base-URL map. They have no chat base URL, so when a request hit /v1/chat/completions with a model like elevenlabs/eleven-multilingual-v2, endpoint resolution fell through to the default case in getProviderEndpoint and threw a confusing requires a baseUrl 500.

Fix

Add an early guard in the chat-completions handler (right after model resolution) that rejects models whose output is audio or video — they belong to dedicated endpoints — with a clear 400 pointing the caller to the right endpoint:

  • audio → /v1/audio/speech
  • video → /v1/videos

Image-only models (e.g. reve, grok-image) are intentionally allowed, since they are served by the chat-completions image-generation flow. This mirrors the existing convention in chat-helpers.e2e.ts that filters out audio/video-only (but not image) models from chat e2e tests.

Test

Added a unit test in apps/gateway/src/api.spec.ts asserting elevenlabs/eleven-multilingual-v2 on /v1/chat/completions returns 400 with a pointer to /v1/audio/speech.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • Bug Fixes

    • Chat completions endpoint now validates that models support text output before processing requests. Models that only support audio or video are rejected with helpful error messages guiding users to the correct endpoint (/v1/audio/speech for audio, /v1/videos for video).
  • Tests

    • Added test coverage for model compatibility validation in chat completions.

Text-to-speech models (e.g. elevenlabs/eleven-multilingual-v2) and video
models are served by dedicated endpoints and have no chat base URL, so
routing them through /v1/chat/completions fell through to a confusing
"Provider elevenlabs requires a baseUrl" 500. Reject them early with a
400 pointing to /v1/audio/speech or /v1/videos. Image-only models stay
allowed since they use the chat-completions image flow.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Jun 22, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

Walkthrough

The chat-completions handler gains an early output-type guard that throws HTTP 400 when the resolved model's output does not include "text", directing callers to /v1/audio/speech for audio models and /v1/videos for video models. An integration test validates this behaviour for the elevenlabs/eleven-multilingual-v2 speech model.

Changes

Non-text model rejection in chat completions

Layer / File(s) Summary
Non-text model guard and integration test
apps/gateway/src/chat/chat.ts, apps/gateway/src/api.spec.ts
Inserts an early guard after modelInfo.output is resolved that throws HTTPException(400) with targeted endpoint hints when the model lacks text output. The integration test provisions an API key, posts with elevenlabs/eleven-multilingual-v2, and asserts a 400 status and /v1/audio/speech in the error message.

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~8 minutes

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title 'fix(gateway): reject TTS/video models in chat' accurately and concisely describes the main change: adding validation to reject text-to-speech and video models in the chat completions endpoint with a clear error message directing users to appropriate endpoints.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/reject-tts-models-in-chat

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (2)
apps/gateway/src/api.spec.ts (1)

84-86: 🧹 Nitpick | 🔵 Trivial | ⚡ Quick win

Drop the historical context comment from the test body.

Line 84 through Line 86 contain narrative context that is not needed for understanding the assertion path and conflicts with the no-unnecessary-comments rule.

As per coding guidelines, **/*.{ts,tsx,js,jsx}: "No unnecessary code comments."

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@apps/gateway/src/api.spec.ts` around lines 84 - 86, Remove the three-line
comment block starting with "ElevenLabs models are speech-only" from the test
body in the api.spec.ts file. This comment provides historical context about
previous behavior rather than explaining the current test assertion, and
therefore violates the no-unnecessary-comments rule. Delete the entire comment
block that discusses the fallthrough behavior and 500 error to keep the test
focused on the actual assertion being tested.

Source: Coding guidelines

apps/gateway/src/chat/chat.ts (1)

1763-1768: 🧹 Nitpick | 🔵 Trivial | ⚡ Quick win

Remove the verbose inline rationale block and keep code self-explanatory.

Line 1763 through Line 1768 add a long explanatory comment that mostly restates behavior already evident in the guard and test coverage. Please trim/remove it to follow repository comment policy.

As per coding guidelines, **/*.{ts,tsx,js,jsx}: "No unnecessary code comments."

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@apps/gateway/src/chat/chat.ts` around lines 1763 - 1768, Remove the verbose
multi-line comment block that begins with "Text-to-speech and video models are
served by dedicated endpoints" in the chat.ts file. This lengthy rationale
restates behavior that is already evident from the guard logic and test
coverage, violating the repository's "No unnecessary code comments" policy.
Either delete the comment entirely or replace it with a single concise line if
additional context is truly needed, allowing the code logic itself to be
self-explanatory.

Source: Coding guidelines

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Nitpick comments:
In `@apps/gateway/src/api.spec.ts`:
- Around line 84-86: Remove the three-line comment block starting with
"ElevenLabs models are speech-only" from the test body in the api.spec.ts file.
This comment provides historical context about previous behavior rather than
explaining the current test assertion, and therefore violates the
no-unnecessary-comments rule. Delete the entire comment block that discusses the
fallthrough behavior and 500 error to keep the test focused on the actual
assertion being tested.

In `@apps/gateway/src/chat/chat.ts`:
- Around line 1763-1768: Remove the verbose multi-line comment block that begins
with "Text-to-speech and video models are served by dedicated endpoints" in the
chat.ts file. This lengthy rationale restates behavior that is already evident
from the guard logic and test coverage, violating the repository's "No
unnecessary code comments" policy. Either delete the comment entirely or replace
it with a single concise line if additional context is truly needed, allowing
the code logic itself to be self-explanatory.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro

Run ID: ed700fe0-df46-478d-99e1-b9bdb849d395

📥 Commits

Reviewing files that changed from the base of the PR and between 66f1045 and caa8692.

📒 Files selected for processing (2)
  • apps/gateway/src/api.spec.ts
  • apps/gateway/src/chat/chat.ts

@steebchen
steebchen added this pull request to the merge queue Jun 22, 2026
@steebchen
steebchen removed this pull request from the merge queue due to a manual request Jun 22, 2026
@steebchen
steebchen merged commit fd996fc into main Jun 22, 2026
17 of 18 checks passed
@steebchen
steebchen deleted the fix/reject-tts-models-in-chat branch June 22, 2026 22:30
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant