Repository navigation
Conversation
Greptile SummaryAdds DashScope realtime WebSocket support through a provider-specific handler built on the shared OpenAI realtime transport.
Confidence Score: 5/5The PR appears safe to merge; the previous findings are resolved or withdrawn, and the latest revision introduces no actionable regression. The credential-rejection path now uses a valid WebSocket close code, environment-only health-check credentials are resolved, and the provider-boundary concern was correctly withdrawn after confirming the implementation follows the existing realtime architecture. No new blocking or non-blocking findings remain. Important Files Changed
Reviews (3): Last reviewed commit: "feat(dashscope): add realtime websocket ..." | Re-trigger Greptile |
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
9627d63 to
29470ba
Compare
29470ba to
d9313df
Compare
706d378 to
6d18f6e
Compare
6d18f6e to
52e511b
Compare
Add a handler for DashScope's /api-ws/v1/realtime WebSocket API, wire it into the realtime proxy path and the realtime health check, and register the Qwen-Omni-Realtime family in the cost map. DashScope deviates from OpenAI's realtime endpoint in three ways that matter: the path is /api-ws/v1/realtime rather than /v1/realtime, auth is a bearer token with no OpenAI-Beta header, and the session shape is the flat beta one. That last point needs OpenAIRealtime to stop remapping the client's session.update into GA's nested output_modalities / audio.input form, which DashScope silently drops and which broke audio sessions. It sits behind an opt-in override so no other provider's behavior changes.
52e511b to
ddd35a2
Compare
|
@mateo-berri ready for review. Rebased onto main, CLA signed, everything green except one check that also fails on main |
TLDR
Problem this solves:
How it solves it:
User Flow
Before: a developer pointing a realtime client at a DashScope model never gets a session
model: dashscope/qwen3.5-omni-plus-realtimewss://<litellm-host>/v1/realtime?model=qwen3.5-omni-plus-realtimesession.createdarrives and the socket closes, so there is nothing to talk toAfter: the same client gets a real DashScope session and a reply
wss://<litellm-host>/v1/realtime?model=qwen3.5-omni-plus-realtimesession.createdarrives andsession.updatedshows the modalities the client asked forresponse.donereports the token usageRelevant issues
Linear ticket
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Screenshots / Proof of Fix
Setup:
ds_probe.pyenters LiteLLM at its realtime entry point and talks to the real DashScopeendpoint
wss://dashscope.aliyuncs.com/api-ws/v1/realtimewith a real API key, so there is nomocked upstream and the runs cost real money. It sends
session.update(flat,modalities: ["text"]), thenconversation.item.createwith aninput_textpart, thenresponse.create.Before (e5da593)
case 1: text session round trip
DS_KEY=<key> python ds_probe.pyraised: ValueError : Unsupported model: qwen3.5-omni-plus-realtimecase 2: rejected credentials
DS_KEY=<key> python ds_probe.pyraised: ValueError : Unsupported model: qwen3.5-omni-plus-realtimeAfter (ddd35a2)
case 1: text session round trip
DS_KEY=<key> python ds_probe.pysession.created,session.updated,conversation.item.created,response.created,response.output_item.added,conversation.item.created,response.content_part.added,response.text.delta,response.text.done,response.content_part.done,response.output_item.done,response.donemodalities=['text'] input_audio_format='pcm' output_modalities=None nested_audio=False'pong'case 2: rejected credentials
DS_KEY=sk-definitely-invalid python ds_probe.pyclient close_code=1011 reason='Internal server error: server rejected WebSocket connection: HTTP 401'case 3: real audio turn (a real voice question, real spoken answer)
Input is
tests/e2e/llm_translation/realtime/fixtures/weather_question_24k.wav, the 24 kHz PCM16 fixture the repo's own realtime e2e suite uses. It is streamed up asinput_audio_buffer.appendthroughws://localhost:4000/v1/realtime, exactly what a voice client does.python ds_audio_probe.pysession.created,input_audio_buffer.speech_stopped,input_audio_buffer.committed,conversation.item.input_audio_transcription.completed,response.created, 40xresponse.audio.delta,response.audio_transcript.done,response.audio.done,response.done"I don't have access to real-time weather data, so you'll need to check a local forecast or a weather website for the current conditions in Paris."{"input_tokens": 505, "output_tokens": 130, "input_tokens_details": {"text_tokens": 484, "audio_tokens": 21}, "output_tokens_details": {"text_tokens": 33, "audio_tokens": 97}}That usage also confirms the audio rates in the cost map do the work they claim. Priced with litellm's own
generic_cost_per_token:Both match the computed values exactly. Audio is 93.6% of the output bill, so a session billed at text rates alone would undercount the output side by roughly 4x.
A workspace-scoped deployment is also exercised: with
api_base: https://<workspace>.cn-beijing.maas.aliyuncs.com/compatible-mode/v1, the handler resolves the backend towss://<workspace>.cn-beijing.maas.aliyuncs.com/api-ws/v1/realtime. Probing that host directly shows/api-ws/v1/realtimeanswers withsession.createdwhile/compatible-mode/v1/realtimeand/v1/realtimeboth return HTTP 404, so the path rewrite in_construct_urlis what makes the OpenAI-compatible base usable at all.Type
🆕 New Feature
Caveats (if any)
Medium
Low
input_cost_per_tokenomittedFinal Attestation
Known limitation, not introduced here
A realtime health check resolves
api_keyonly from the deployment's resolvedlitellm_params, so a deployment that supplies credentials solely through the process environment reports unhealthy. Verified on this branch: withapi_key: os.environ/DASHSCOPE_API_KEYthe check returnshealthy_count: 1, and with noapi_keyline at all it returnsHTTP 401from an empty bearer token, even though the request path works fine in both cases.This is not specific to DashScope.
_realtime_health_check_auth_headersreturns empty headers for every provider whenapi_keyisNone, and the OpenAI, xAI, Bedrock and Vertex branches behave the same way. Flagging it rather than fixing it here to keep this PR scoped to DashScope; happy to open a separate PR for the shared helper if that is preferred.