Skip to content

fix: derive OpenAI audio providers from declared supported_endpoints - #43784

Open
jiweiyeah wants to merge 2 commits into
BerriAI:mainfrom
jiweiyeah:fix/derive-audio-providers-from-declared-endpoints
Open

jiweiyeah wants to merge 2 commits into
BerriAI:mainfrom
jiweiyeah:fix/derive-audio-providers-from-declared-endpoints

Conversation

@jiweiyeah

Copy link
Copy Markdown

TLDR

Problem this solves:

  • providers.json providers became audio-capable without asking for audio
  • transcription() and speech() sent OPENAI_API_KEY to third-party base urls
  • 17 of the 30 JSON providers were reachable this way, none declares an audio endpoint
  • speech() bypassed the audio allow-list entirely and read openai_compatible_providers

How it solves it:

  • The audio allow-list is derived from each provider's declared supported_endpoints
  • Providers with no audio declaration are excluded, and a missing key means no audio
  • speech() and transcription() now consult the same derived set
  • The 53 python-configured OpenAI-compatible providers keep routing exactly as before

Intentional product change: transcription("<json-provider>/whisper-1") and speech("<json-provider>/tts-1") now fail instead of sending a request the provider never advertised. Operators who do want audio through one of these providers add /v1/audio/transcriptions or /v1/audio/speech to that entry in providers.json, which is the point: audio support becomes a claim the provider config makes rather than a side effect of sharing a base class

User Flow

Before: a proxy operator who set only an OpenAI key can have that key forwarded to an unrelated third party

  1. Operator starts the gateway with OPENAI_API_KEY set and no PUBLICAI_API_KEY, no POE_API_KEY
  2. Any key holder sends POST https://gateway.example/v1/audio/transcriptions with "model": "publicai/whisper-1" and an audio file
  3. The gateway returns 200 with whatever api.publicai.co replied, and the Admin UI logs page shows the call as a normal transcription
  4. api.publicai.co received Authorization: Bearer <the operator's OpenAI key>, and POST /v1/audio/speech with "model": "poe/tts-1" behaves the same way
  5. Any authenticated caller could spend the operator's OpenAI quota and expose the credential to a service they picked

After: the same request is refused before a socket opens

  1. Operator runs the same gateway with the same environment
  2. Same POST https://gateway.example/v1/audio/transcriptions with "model": "publicai/whisper-1"
  3. The gateway returns an error naming the provider as unmapped, and POST /v1/audio/speech with "model": "poe/tts-1" is refused the same way
  4. The logs page shows the rejected call with no upstream request recorded
  5. No caller can push the operator's OpenAI credential to a JSON provider that does not declare audio

Relevant issues

Surfaced by a security review on a provider-addition PR, #41559, which added a JSON-configured provider and inherited this exposure the same way the other 16 did. The fix here is repo-wide and does not depend on that PR merging

Pre-Submission checklist

  • I have added meaningful tests
  • The handful of test files covering my change pass locally
  • My PR passes all required CI/CD checks
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5

tests/unit/llms/openai_like was fully green before this change (196 passed) and green after it (206 passed, 10 new). The wider suites I ran, tests/unit/llms/{openai,groq,publicai,azure,deepgram,soniox,edenai} plus every unit test file matching *audio*, *speech* or *transcription*, produce byte-identical failure lists before and after: 41 failures from optional packages missing in my local venv (azure-identity, websockets, the guardrail deps) and 1 pre-existing test_bytesio_at_end_position. No commit on this branch adds a failure, and I am not claiming a green board

Screenshots / Proof of Fix

Deliberately not a live-API run: the bug is a credential leaving litellm, so reproducing it against a real provider would disclose a real key to a real third party. Both revisions stand up a plain http.server on loopback, point the provider's api_base at it, and print every request that arrives together with its Authorization header. Nothing inside litellm is mocked, the only substitution is the upstream host

Shared setup for both revisions, /tmp/leak_probe.py:

os.environ["OPENAI_API_KEY"] = "sk-OPENAI-canary-must-not-leave-litellm"
os.environ.pop("PUBLICAI_API_KEY", None)
os.environ.pop("POE_API_KEY", None)
# plain http.server on a free loopback port, recording (path, Authorization) per POST

litellm.transcription(model="publicai/whisper-1", file=audio_file(), api_base=f"http://127.0.0.1:{port}/v1")
litellm.speech(model="poe/tts-1", input="hello", voice="alloy", api_base=f"http://127.0.0.1:{port}/v1")

Before (ffb15f9)

publicai transcription

  1. git stash push -- litellm/ && PYTHONPATH=$PWD .venv/bin/python /tmp/leak_probe.py
  2. outcome: returned without error
  3. requests: 1 -> [{'path': '/v1/audio/transcriptions', 'authorization': 'Bearer sk-OPENAI-canary-must-not-leave-litellm'}]

poe speech

  1. Same command, second case
  2. outcome: returned without error
  3. requests: 1 -> [{'path': '/v1/audio/speech', 'authorization': 'Bearer sk-OPENAI-canary-must-not-leave-litellm'}]

After (184a2b6)

publicai transcription

  1. PYTHONPATH=$PWD .venv/bin/python /tmp/leak_probe.py
  2. outcome: ValueError: Unmapped provider passed in. Unable to get the response.
  3. requests: 0

poe speech

  1. Same command, second case
  2. outcome: Exception: Unable to map the custom llm provider=poe to a known provider=[...]
  3. requests: 0

groq transcription, unchanged

  1. tests/unit/llms/openai_like/test_json_loader.py::test_transcription_still_routes_a_python_provider_with_its_own_credential
  2. POST https://api.groq.com/openai/v1/audio/transcriptions goes out with authorization: Bearer gsk-provider-key
  3. response.text == "hello world", so providers that were always allowed to use the OpenAI audio transport still do, with their own credential

Type

🐛 Bug Fix

Caveats (if any)

High

  • The other half is not fixed here. completion("<json-provider>/<model>") where the provider's own *_API_KEY is unset still falls back to OPENAI_API_KEY in OpenAIGPTConfig.get_api_key (litellm/llms/openai/chat/gpt_transformation.py:793) while get_complete_url sends it to that provider's base url. Same disclosure on the chat path for all 17 providers, and no allow-list is involved. I can send that separately, since it changes credential resolution rather than a derived set

Medium

  • Membership moved out of litellm/constants.py. constants.py cannot import litellm.llms.openai_like.json_loader, because litellm/llms/__init__.py imports litellm._logging, which imports litellm.constants, and that edge would close a cycle. Reading providers.json from constants.py instead would parse untyped JSON under reportAny: error. The set now sits next to the registry that already owns the file, litellm.OPENAI_AUDIO_TRANSCRIPTION_PROVIDERS still resolves through litellm.main, and litellm.constants.OPENAI_AUDIO_TRANSCRIPTION_PROVIDERS no longer exists
  • No provider in providers.json declares an audio endpoint today, so the "still audio-capable once it declares audio" branch has no production example. It is covered by parametrized cases against the derivation function, which is why the derivation takes an injected mapping instead of being an inline expression

Low

  • speech() surfaces the block as Unable to map the custom llm provider=<slug> to a known provider=[...], the existing fall-through message. Accurate, but not a friendly hint that the provider config needs an audio endpoint
  • A supported_endpoints key that omits audio and a missing key both land on excluded, so the conservative default answers the same way either way

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

@greptile-apps

greptile-apps Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 2/5

[High risk] Moves audio provider list from static constant to dynamic derivation.

The PR is not safe to merge until declared JSON audio routes enforce the requested operation and speech uses the resolved provider credential.

Findings

  1. P1 Either declaration enables both operations ▶
  2. P1 Security Speech ignores provider credentials ▶

Summary

The PR derives OpenAI-compatible audio routing from JSON provider endpoint declarations and applies it to speech as well as transcription. Undeclared JSON providers are excluded, but the declared-audio path still needs operation-specific checks and correct speech credential handling.

Reviews (1) · Last reviewed commit: "fix: derive OpenAI audio providers from ..."

declaring_providers: Final = frozenset(
slug
for slug, endpoints in declared_endpoints.items()
if any(endpoint in OPENAI_AUDIO_ENDPOINTS for endpoint in endpoints)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Either declaration enables both operations
If a JSON provider declares only /v1/audio/transcriptions, this set also permits speech() requests to that provider. Declaring only /v1/audio/speech likewise permits transcription() requests. Both functions use this set, so either declaration sends the other, undeclared operation upstream rather than rejecting it.

Knowledge Base Used: Provider adapters and capabilities

Comment thread litellm/main.py
)
elif custom_llm_provider == "openai" or (
custom_llm_provider in litellm.openai_compatible_providers
custom_llm_provider in OPENAI_AUDIO_TRANSCRIPTION_PROVIDERS

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 security Speech ignores provider credentials
When a JSON provider declares audio speech, get_llm_provider() resolves its provider-specific key into dynamic_api_key, but this route does not use that value. Even if the provider key is configured, speech() can instead fall back to OPENAI_API_KEY and send it to the provider's base URL. How this was verified: JSON provider resolution returns its configured key separately, while speech dispatch ignores that value and falls back to the OpenAI key.

Knowledge Base Used: Provider adapters and capabilities

@codecov

codecov Bot commented Sep 30, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@codspeed

codspeed Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing jiweiyeah:fix/derive-audio-providers-from-declared-endpoints (e3a2947) with main (8520626)1

Open in CodSpeed

Footnotes

  1. No successful run was found on main (5d513d8) during the generation of this report, so 8520626 was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩

jiweiyeah and others added 2 commits October 4, 2026 11:02
OPENAI_AUDIO_TRANSCRIPTION_PROVIDERS took the whole openai_compatible_providers
list, so every JSON-configured provider in providers.json became audio-capable
whether or not it advertised an audio endpoint. 17 of the 30 JSON providers are
in that list and none of them declares /v1/audio/transcriptions or
/v1/audio/speech, yet litellm.transcription() and litellm.speech() would send
them the caller's OPENAI_API_KEY against their third-party base url.

speech() had the same hole through a second path: it tested membership in
litellm.openai_compatible_providers directly, bypassing the audio constant.
Both dispatch sites now read one derived set.

Derivation lives in litellm/llms/openai_like/json_loader.py because constants.py
cannot import it: litellm/llms/__init__.py imports litellm._logging, which
imports litellm.constants. A provider with no supported_endpoints key is treated
as not audio-capable.

Tests: tests/unit/llms/openai_like/test_json_loader.py
Signed-off-by: jiweiyeah <yeahjiwei@163.com>
@jiweiyeah
jiweiyeah force-pushed the fix/derive-audio-providers-from-declared-endpoints branch from b717835 to e3a2947 Compare October 4, 2026 03:05

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant