Repository navigation
feat(tium): add Tium as a JSON-configured OpenAI-compatible provider - #40367
curiousbox wants to merge 1 commit into
Conversation
Greptile SummaryAdds Tium as a JSON-configured OpenAI-compatible provider.
Confidence Score: 5/5The PR appears safe to merge; no actionable regression remains in the changes since the previous review. Both previous findings were resolved, and the only subsequent change corrects a stale auto-router test expectation while preserving its substantive reasoning-effort assertion.
|
| Filename | Overview |
|---|---|
| litellm/llms/openai_like/providers.json | Registers Tium’s base URL, environment variables, parameter mapping, and chat-completions endpoint. |
| model_prices_and_context_window.json | Adds pricing, limits, and capability metadata for five Tium models. |
| tests/test_litellm/llms/openai_like/test_tium_provider.py | Covers provider registration and mocked completion behavior, including URL, authentication, tools, and parameter mapping. |
| ui/litellm-dashboard/src/components/provider_info_helpers.tsx | Adds Tium’s dashboard display name, provider slug, logo, and model placeholder. |
| ui/litellm-dashboard/src/components/add_model/add_auto_router_tab.test.tsx | Aligns the expected reasoning-tier model with the current Anthropic preset while retaining the reasoning-effort contract. |
Reviews (3): Last reviewed commit: "test(auto-router): derive the REASONING ..." | Re-trigger Greptile
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
Tium serves open-weight models (GLM, DeepSeek, Kimi) over an OpenAI-compatible API from a gateway operated in Germany. Callers currently have to route through openai/ with a custom api_base Registered in providers.json with TIUM_API_KEY, the TIUM_API_BASE override, the max_completion_tokens to max_tokens mapping, and chat completions as the only supported endpoint, since POST /v1/responses answers 404 on this host. The slug is also added to openai_compatible_providers and LlmProviders, and the base URL to openai_compatible_endpoints for api_base autodetection The openai_compatible_providers entry is load-bearing. JSONProviderRegistry does not register into that list, and a provider missing from it takes the non-OpenAI-compatible branch of add_provider_specific_params_to_optional_params, so provider-specific body params arrive as top-level kwargs instead of nesting in extra_body, and tools are stripped before the request is sent. See BerriAI#26443, still open against three JSON-only providers Five chat models go into the cost map, with each capability flag measured rather than defaulted. All five support function calling and prompt caching, so each carries a cache read rate. The two DeepSeek models refuse response_format with a JSON schema; image input works on glm-5.3-flash and kimi-k3. max_input_tokens and max_output_tokens are the host's policy ceilings. No constraints block is needed, since all five accept temperature from 0.0 to 2.0 Costs are USD per token, derived from the published subscription rate of $2.83 per million weighted tokens times each model's multiplier and weights. Tium bills weighted tokens rather than per-token, and prepaid packs cost more; the docs page states both The provider is registered in the Add Model form and the Admin UI provider list, which test_every_backend_provider_is_listed_in_add_model_or_frozen_as_unlisted requires of any new entry in LlmProviders Tests drive real completion calls through respx, asserting the request URL, the bearer auth header, the message body, that tools and tool_choice reach the wire and tool calls come back, and that max_completion_tokens is sent as max_tokens. They use the local_model_cost_map fixture, since the network-fetched cost map lags this branch until merge and capability lookups would otherwise read stale data
60b4e2c to
c59b321
Compare
18b9cc0 to
c59b321
Compare
|
That is now fixed upstream by #40456, merged into Worth recording why this PR surfaced it at all, in case it helps someone else: |
TLDR
Problem this solves:
openai/with a customapi_baseHow it solves it:
tiuminproviders.json, so no Python module is neededUser Flow
Before: a developer routing to Tium has to describe it as a generic OpenAI server, and every request is billed at zero, so budgets and spend dashboards never see the traffic
openai/glm-5.3, withapi_base: https://api.tium.ai/v1and theirsk-tium-...key, then start the proxy on http://localhost:4000"model": "glm-5.3", a user message asking for the weather in Paris, and atoolsarray holding aget_weatherfunctionx-litellm-response-costheader at all, and everyx-litellm-response-cost-*breakdown header reads0.0After: the same developer names the provider directly and the identical request is priced, so budgets and dashboards work
tium/glm-5.3, setTIUM_API_KEYin the environment, and start the proxy on http://localhost:4000"model": "glm-5.3", the same user message and the sametoolsarrayx-litellm-response-cost: 0.000537138, broken out into input, output, cache read and reasoning headersRelevant issues
Docs live in the other repo, so the provider page goes up separately in BerriAI/litellm-docs#1388. That one is not a prerequisite for merging this:
documentationandcode-qualityboth pass here already. It should still land so the provider has a page.The
openai_compatible_providersentry is deliberate rather than incidental, and relates to #26443.JSONProviderRegistrydoes not register into that list, and a provider missing from it takes the non-OpenAI-compatible branch ofadd_provider_specific_params_to_optional_params, so provider-specific body params arrive as top-level kwargs instead of nesting inextra_body, andtoolsare stripped before the request is sent. That issue is still open against three JSON-only providers.Linear ticket
Pre-Submission checklist
osv-scan, is a base-branch finding this PR does not cause:smol-toml1.6.1 (GHSA-7w5x-hrqm-74c2), a dev-only transitive dependency ofknippinned inui/litellm-dashboard/package-lock.jsononlitellm_internal_staging. This PR does not touch the lockfile. The advisory is newly published, so the base is affected independently of this PR. Happy to bump it here if you would rather that than a separate fixType
New Feature
Screenshots / Proof of Fix
Shared setup for both runs, against the live Tium API with a real key.
tium_config.yaml, Before:tium_config.yaml, After:Proxy started with
LITELLM_LOCAL_MODEL_COST_MAP=True python litellm/proxy/proxy_cli.py --config tium_config.yaml --port 4000payload.json, one user turn and one tool:{ "model": "glm-5.3", "messages": [{"role": "user", "content": "What is the weather in Paris?"}], "tools": [{"type": "function", "function": {"name": "get_weather", "description": "Get current weather for a location", "parameters": {"type": "object", "properties": {"location": {"type": "string"}}, "required": ["location"]}}}] }Before (ee7c7e1)
Cost of a tool-calling chat completion
There is no
x-litellm-response-costheader, and every breakdown header is zerogrep -i x-litellm-model-name headers.txtreportsopenai/glm-5.3, the workaround routingAfter (c59b321)
Cost of a tool-calling chat completion
grep -i x-litellm-model-name headers.txtreportstium/glm-5.3The total matches the published rates by hand, so the cost map entry is right rather than merely non-zero: 34 uncached input tokens at $2.830/M is $0.00009622, 128 cached at $0.526/M is $0.000067328, 42 output at $8.895/M is $0.00037359, summing to $0.000537138
Caveats (if any)
Low
toolstotium/*fails with a 400 if the code ships ahead of the cost mapmain, so the two normally land togethermainUnsupportedParamsError: tium does not support parameters: ['tools']LITELLM_LOCAL_MODEL_COST_MAP=Trueorallowed_openai_params=['tools']/v1/chat/completionsis wired, since the host answers 404 on/v1/responsesreasoning_effort=nonetool_choice="required"while reasoning is onFinal Attestation