Skip to content

feat(dashscope): add realtime websocket support - #40579

Open
kuun993 wants to merge 1 commit into
BerriAI:mainfrom
kuun993:litellm_dashscope_realtime
Open

kuun993 wants to merge 1 commit into
BerriAI:mainfrom
kuun993:litellm_dashscope_realtime

Conversation

@kuun993

@kuun993 kuun993 commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • LiteLLM cannot serve DashScope realtime models at all
  • DashScope speaks a different wire format than OpenAI

How it solves it:

  • Adds a DashScopeRealtime handler on the shared OpenAI one
  • Keeps DashScope's flat session.update out of the GA remap
  • Catches the handshake exception class websockets 15 raises
  • Registers the Qwen-Omni-Realtime models in the cost map

User Flow

Before: a developer pointing a realtime client at a DashScope model never gets a session

  1. They add a deployment with model: dashscope/qwen3.5-omni-plus-realtime
  2. Their client opens wss://<litellm-host>/v1/realtime?model=qwen3.5-omni-plus-realtime
  3. No session.created arrives and the socket closes, so there is nothing to talk to

After: the same client gets a real DashScope session and a reply

  1. They add the same deployment
  2. Their client opens wss://<litellm-host>/v1/realtime?model=qwen3.5-omni-plus-realtime
  3. session.created arrives and session.updated shows the modalities the client asked for
  4. The client sends a message and response.done reports the token usage
  5. If the key is wrong, the socket closes with 1011 and the reason names the upstream 401

Relevant issues

Linear ticket

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Screenshots / Proof of Fix

Setup: ds_probe.py enters LiteLLM at its realtime entry point and talks to the real DashScope
endpoint wss://dashscope.aliyuncs.com/api-ws/v1/realtime with a real API key, so there is no
mocked upstream and the runs cost real money. It sends session.update (flat, modalities: ["text"]), then conversation.item.create with an input_text part, then response.create.

Before (e5da593)

case 1: text session round trip

  1. DS_KEY=<key> python ds_probe.py
  2. Observed: raised: ValueError : Unsupported model: qwen3.5-omni-plus-realtime

case 2: rejected credentials

  1. DS_KEY=<key> python ds_probe.py
  2. Observed: raised: ValueError : Unsupported model: qwen3.5-omni-plus-realtime

After (ddd35a2)

case 1: text session round trip

  1. DS_KEY=<key> python ds_probe.py
  2. Observed events: session.created, session.updated, conversation.item.created, response.created, response.output_item.added, conversation.item.created, response.content_part.added, response.text.delta, response.text.done, response.content_part.done, response.output_item.done, response.done
  3. Observed session echo: modalities=['text'] input_audio_format='pcm' output_modalities=None nested_audio=False
  4. Observed assistant text: 'pong'

case 2: rejected credentials

  1. DS_KEY=sk-definitely-invalid python ds_probe.py
  2. Observed: client close_code=1011 reason='Internal server error: server rejected WebSocket connection: HTTP 401'

case 3: real audio turn (a real voice question, real spoken answer)

Input is tests/e2e/llm_translation/realtime/fixtures/weather_question_24k.wav, the 24 kHz PCM16 fixture the repo's own realtime e2e suite uses. It is streamed up as input_audio_buffer.append through ws://localhost:4000/v1/realtime, exactly what a voice client does.

  1. python ds_audio_probe.py
  2. Observed events: session.created, input_audio_buffer.speech_stopped, input_audio_buffer.committed, conversation.item.input_audio_transcription.completed, response.created, 40x response.audio.delta, response.audio_transcript.done, response.audio.done, response.done
  3. Observed input transcription: the upstream transcribed the spoken question correctly
  4. Observed answer text: "I don't have access to real-time weather data, so you'll need to check a local forecast or a weather website for the current conditions in Paris."
  5. Observed audio returned: 372480 bytes, 7.76 s at 24 kHz, peak amplitude 13823, RMS 1130 (real speech, not an empty stream)
  6. Observed usage: {"input_tokens": 505, "output_tokens": 130, "input_tokens_details": {"text_tokens": 484, "audio_tokens": 21}, "output_tokens_details": {"text_tokens": 33, "audio_tokens": 97}}

That usage also confirms the audio rates in the cost map do the work they claim. Priced with litellm's own generic_cost_per_token:

484 * input_cost_per_token        + 21 * input_cost_per_audio_token  = $0.0013629
 33 * output_cost_per_token       + 97 * output_cost_per_audio_token = $0.0064232

Both match the computed values exactly. Audio is 93.6% of the output bill, so a session billed at text rates alone would undercount the output side by roughly 4x.

A workspace-scoped deployment is also exercised: with api_base: https://<workspace>.cn-beijing.maas.aliyuncs.com/compatible-mode/v1, the handler resolves the backend to wss://<workspace>.cn-beijing.maas.aliyuncs.com/api-ws/v1/realtime. Probing that host directly shows /api-ws/v1/realtime answers with session.created while /compatible-mode/v1/realtime and /v1/realtime both return HTTP 404, so the path rewrite in _construct_url is what makes the OpenAI-compatible base usable at all.

Type

🆕 New Feature

Caveats (if any)

Medium

  • DashScope audio is covered by the real audio turn above on Qwen-Omni-Realtime; the livetranslate family still has no audio turn
  • qwen-audio-3.0-realtime-* publish no token limits; max tokens omitted

Low

  • qwen3.5-livetranslate-* take no text input, so input_cost_per_token omitted

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Known limitation, not introduced here

A realtime health check resolves api_key only from the deployment's resolved litellm_params, so a deployment that supplies credentials solely through the process environment reports unhealthy. Verified on this branch: with api_key: os.environ/DASHSCOPE_API_KEY the check returns healthy_count: 1, and with no api_key line at all it returns HTTP 401 from an empty bearer token, even though the request path works fine in both cases.

This is not specific to DashScope. _realtime_health_check_auth_headers returns empty headers for every provider when api_key is None, and the OpenAI, xAI, Bedrock and Vertex branches behave the same way. Flagging it rather than fixing it here to keep this PR scoped to DashScope; happy to open a separate PR for the shared helper if that is preferred.

@kuun993
kuun993 requested a review from a team September 10, 2026 12:38
@CLAassistant

CLAassistant commented Sep 10, 2026 •

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

@codspeed

codspeed Bot commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing kuun993:litellm_dashscope_realtime (ddd35a2) with main (90e4962)

Open in CodSpeed

@greptile-apps

greptile-apps Bot commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

Adds DashScope realtime WebSocket support through a provider-specific handler built on the shared OpenAI realtime transport.

  • Resolves DashScope regional and workspace WebSocket endpoints and bearer authentication.
  • Preserves DashScope’s beta-style realtime session event format.
  • Integrates realtime request dispatch and deployment health checks.
  • Registers supported Qwen realtime models, capabilities, and token pricing.
  • Adds unit and provider translation coverage for URL construction, event forwarding, authentication failures, and audio-token billing.

Confidence Score: 5/5

The PR appears safe to merge; the previous findings are resolved or withdrawn, and the latest revision introduces no actionable regression.

The credential-rejection path now uses a valid WebSocket close code, environment-only health-check credentials are resolved, and the provider-boundary concern was correctly withdrawn after confirming the implementation follows the existing realtime architecture. No new blocking or non-blocking findings remain.

Important Files Changed
Filename Overview
litellm/llms/dashscope/realtime/handler.py Defines DashScope endpoint resolution, authentication, URL construction, and beta-protocol behavior.
litellm/llms/openai/realtime/handler.py Adds a provider-overridable upstream protocol indicator to shared realtime forwarding.
litellm/realtime_api/main.py Dispatches DashScope realtime sessions and constructs authenticated health checks through the provider handler.
tests/test_litellm/llms/dashscope/test_dashscope_realtime_handler.py Covers the DashScope wire contract, endpoint selection, rejected handshakes, and realtime token billing.
model_prices_and_context_window.json Registers DashScope realtime model capabilities, limits, and separate text and audio pricing.
provider_endpoints_support.json Advertises realtime endpoint support for DashScope.

Reviews (3): Last reviewed commit: "feat(dashscope): add realtime websocket ..." | Re-trigger Greptile

Comment thread litellm/llms/openai/realtime/handler.py Outdated
Comment thread litellm/realtime_api/main.py Outdated
Comment thread litellm/realtime_api/main.py
@codecov

codecov Bot commented Sep 10, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 91.89189% with 3 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
litellm/llms/dashscope/realtime/handler.py 92.00% 2 Missing ⚠️
litellm/realtime_api/main.py 87.50% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

@kuun993
kuun993 force-pushed the litellm_dashscope_realtime branch from 9627d63 to 29470ba Compare September 10, 2026 12:59
@kuun993
kuun993 force-pushed the litellm_dashscope_realtime branch from 29470ba to d9313df Compare September 11, 2026 01:28
@yuneng-berri
yuneng-berri deleted the branch BerriAI:main September 13, 2026 04:51
@yuneng-berri yuneng-berri reopened this Sep 13, 2026
@kuun993
kuun993 changed the base branch from litellm_internal_staging to main September 15, 2026 03:39
@kuun993 kuun993 closed this Sep 16, 2026
@kuun993
kuun993 deleted the litellm_dashscope_realtime branch September 16, 2026 09:52
@kuun993
kuun993 restored the litellm_dashscope_realtime branch September 16, 2026 10:09
@kuun993 kuun993 reopened this Sep 16, 2026
@kuun993
kuun993 force-pushed the litellm_dashscope_realtime branch 2 times, most recently from 706d378 to 6d18f6e Compare September 16, 2026 11:24
@kuun993
kuun993 force-pushed the litellm_dashscope_realtime branch from 6d18f6e to 52e511b Compare September 28, 2026 09:06
Add a handler for DashScope's /api-ws/v1/realtime WebSocket API, wire it into the
realtime proxy path and the realtime health check, and register the
Qwen-Omni-Realtime family in the cost map.

DashScope deviates from OpenAI's realtime endpoint in three ways that matter: the
path is /api-ws/v1/realtime rather than /v1/realtime, auth is a bearer token with no
OpenAI-Beta header, and the session shape is the flat beta one. That last point needs
OpenAIRealtime to stop remapping the client's session.update into GA's nested
output_modalities / audio.input form, which DashScope silently drops and which broke
audio sessions. It sits behind an opt-in override so no other provider's behavior
changes.
@kuun993
kuun993 force-pushed the litellm_dashscope_realtime branch from 52e511b to ddd35a2 Compare September 28, 2026 10:09
@kuun993

kuun993 commented Sep 28, 2026

Copy link
Copy Markdown
Contributor Author

@mateo-berri ready for review. Rebased onto main, CLA signed, everything green except one check that also fails on main

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants