Repository navigation
Conversation
|
| declaring_providers: Final = frozenset( | ||
| slug | ||
| for slug, endpoints in declared_endpoints.items() | ||
| if any(endpoint in OPENAI_AUDIO_ENDPOINTS for endpoint in endpoints) |
There was a problem hiding this comment.
Either declaration enables both operations
If a JSON provider declares only /v1/audio/transcriptions, this set also permits speech() requests to that provider. Declaring only /v1/audio/speech likewise permits transcription() requests. Both functions use this set, so either declaration sends the other, undeclared operation upstream rather than rejecting it.
Knowledge Base Used: Provider adapters and capabilities
| ) | ||
| elif custom_llm_provider == "openai" or ( | ||
| custom_llm_provider in litellm.openai_compatible_providers | ||
| custom_llm_provider in OPENAI_AUDIO_TRANSCRIPTION_PROVIDERS |
There was a problem hiding this comment.
Speech ignores provider credentials
When a JSON provider declares audio speech, get_llm_provider() resolves its provider-specific key into dynamic_api_key, but this route does not use that value. Even if the provider key is configured, speech() can instead fall back to OPENAI_API_KEY and send it to the provider's base URL. How this was verified: JSON provider resolution returns its configured key separately, while speech dispatch ignores that value and falls back to the OpenAI key.
Knowledge Base Used: Provider adapters and capabilities
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
OPENAI_AUDIO_TRANSCRIPTION_PROVIDERS took the whole openai_compatible_providers list, so every JSON-configured provider in providers.json became audio-capable whether or not it advertised an audio endpoint. 17 of the 30 JSON providers are in that list and none of them declares /v1/audio/transcriptions or /v1/audio/speech, yet litellm.transcription() and litellm.speech() would send them the caller's OPENAI_API_KEY against their third-party base url. speech() had the same hole through a second path: it tested membership in litellm.openai_compatible_providers directly, bypassing the audio constant. Both dispatch sites now read one derived set. Derivation lives in litellm/llms/openai_like/json_loader.py because constants.py cannot import it: litellm/llms/__init__.py imports litellm._logging, which imports litellm.constants. A provider with no supported_endpoints key is treated as not audio-capable. Tests: tests/unit/llms/openai_like/test_json_loader.py Signed-off-by: jiweiyeah <yeahjiwei@163.com>
b717835 to
e3a2947
Compare
TLDR
Problem this solves:
providers.jsonproviders became audio-capable without asking for audiotranscription()andspeech()sentOPENAI_API_KEYto third-party base urlsspeech()bypassed the audio allow-list entirely and readopenai_compatible_providersHow it solves it:
supported_endpointsspeech()andtranscription()now consult the same derived setIntentional product change:
transcription("<json-provider>/whisper-1")andspeech("<json-provider>/tts-1")now fail instead of sending a request the provider never advertised. Operators who do want audio through one of these providers add/v1/audio/transcriptionsor/v1/audio/speechto that entry inproviders.json, which is the point: audio support becomes a claim the provider config makes rather than a side effect of sharing a base classUser Flow
Before: a proxy operator who set only an OpenAI key can have that key forwarded to an unrelated third party
OPENAI_API_KEYset and noPUBLICAI_API_KEY, noPOE_API_KEYhttps://gateway.example/v1/audio/transcriptionswith"model": "publicai/whisper-1"and an audio fileAuthorization: Bearer <the operator's OpenAI key>, and POST/v1/audio/speechwith"model": "poe/tts-1"behaves the same wayAfter: the same request is refused before a socket opens
https://gateway.example/v1/audio/transcriptionswith"model": "publicai/whisper-1"/v1/audio/speechwith"model": "poe/tts-1"is refused the same wayRelevant issues
Surfaced by a security review on a provider-addition PR, #41559, which added a JSON-configured provider and inherited this exposure the same way the other 16 did. The fix here is repo-wide and does not depend on that PR merging
Pre-Submission checklist
tests/unit/llms/openai_likewas fully green before this change (196 passed) and green after it (206 passed, 10 new). The wider suites I ran,tests/unit/llms/{openai,groq,publicai,azure,deepgram,soniox,edenai}plus every unit test file matching*audio*,*speech*or*transcription*, produce byte-identical failure lists before and after: 41 failures from optional packages missing in my local venv (azure-identity,websockets, the guardrail deps) and 1 pre-existingtest_bytesio_at_end_position. No commit on this branch adds a failure, and I am not claiming a green boardScreenshots / Proof of Fix
Deliberately not a live-API run: the bug is a credential leaving litellm, so reproducing it against a real provider would disclose a real key to a real third party. Both revisions stand up a plain
http.serveron loopback, point the provider'sapi_baseat it, and print every request that arrives together with itsAuthorizationheader. Nothing inside litellm is mocked, the only substitution is the upstream hostShared setup for both revisions,
/tmp/leak_probe.py:Before (ffb15f9)
publicai transcription
git stash push -- litellm/ && PYTHONPATH=$PWD .venv/bin/python /tmp/leak_probe.pyoutcome: returned without errorrequests: 1 -> [{'path': '/v1/audio/transcriptions', 'authorization': 'Bearer sk-OPENAI-canary-must-not-leave-litellm'}]poe speech
outcome: returned without errorrequests: 1 -> [{'path': '/v1/audio/speech', 'authorization': 'Bearer sk-OPENAI-canary-must-not-leave-litellm'}]After (184a2b6)
publicai transcription
PYTHONPATH=$PWD .venv/bin/python /tmp/leak_probe.pyoutcome: ValueError: Unmapped provider passed in. Unable to get the response.requests: 0poe speech
outcome: Exception: Unable to map the custom llm provider=poe to a known provider=[...]requests: 0groq transcription, unchanged
tests/unit/llms/openai_like/test_json_loader.py::test_transcription_still_routes_a_python_provider_with_its_own_credentialhttps://api.groq.com/openai/v1/audio/transcriptionsgoes out withauthorization: Bearer gsk-provider-keyresponse.text == "hello world", so providers that were always allowed to use the OpenAI audio transport still do, with their own credentialType
🐛 Bug Fix
Caveats (if any)
High
completion("<json-provider>/<model>")where the provider's own*_API_KEYis unset still falls back toOPENAI_API_KEYinOpenAIGPTConfig.get_api_key(litellm/llms/openai/chat/gpt_transformation.py:793) whileget_complete_urlsends it to that provider's base url. Same disclosure on the chat path for all 17 providers, and no allow-list is involved. I can send that separately, since it changes credential resolution rather than a derived setMedium
litellm/constants.py.constants.pycannot importlitellm.llms.openai_like.json_loader, becauselitellm/llms/__init__.pyimportslitellm._logging, which importslitellm.constants, and that edge would close a cycle. Readingproviders.jsonfromconstants.pyinstead would parse untyped JSON underreportAny: error. The set now sits next to the registry that already owns the file,litellm.OPENAI_AUDIO_TRANSCRIPTION_PROVIDERSstill resolves throughlitellm.main, andlitellm.constants.OPENAI_AUDIO_TRANSCRIPTION_PROVIDERSno longer existsproviders.jsondeclares an audio endpoint today, so the "still audio-capable once it declares audio" branch has no production example. It is covered by parametrized cases against the derivation function, which is why the derivation takes an injected mapping instead of being an inline expressionLow
speech()surfaces the block asUnable to map the custom llm provider=<slug> to a known provider=[...], the existing fall-through message. Accurate, but not a friendly hint that the provider config needs an audio endpointsupported_endpointskey that omits audio and a missing key both land on excluded, so the conservative default answers the same way either wayFinal Attestation