Skip to content

feat(tium): add Tium as a JSON-configured OpenAI-compatible provider - #40367

Open
curiousbox wants to merge 1 commit into
BerriAI:mainfrom
curiousbox:feat/add-tium-provider
Open

curiousbox wants to merge 1 commit into
BerriAI:mainfrom
curiousbox:feat/add-tium-provider

Conversation

@curiousbox

@curiousbox curiousbox commented Sep 9, 2026 •

Copy link
Copy Markdown

TLDR

Problem this solves:

  • Tium is not a supported provider, so callers have to route through openai/ with a custom api_base
  • Under that workaround every request is logged at zero cost
  • Budgets, spend limits and cost alerts are blind to that traffic

How it solves it:

  • Registers tium in providers.json, so no Python module is needed
  • Adds five chat models to the cost map and the endpoint matrix
  • Registers the slug for the proxy Admin UI provider picker

User Flow

Before: a developer routing to Tium has to describe it as a generic OpenAI server, and every request is billed at zero, so budgets and spend dashboards never see the traffic

  1. They add a model to their proxy config called openai/glm-5.3, with api_base: https://api.tium.ai/v1 and their sk-tium-... key, then start the proxy on http://localhost:4000
  2. They POST http://localhost:4000/v1/chat/completions with "model": "glm-5.3", a user message asking for the weather in Paris, and a tools array holding a get_weather function
  3. A 200 comes back with the tool call they wanted, so the request itself works
  4. The response carries no x-litellm-response-cost header at all, and every x-litellm-response-cost-* breakdown header reads 0.0
  5. http://localhost:4000/ui/?page=logs shows the request at $0 spend, so a key with a budget can spend against Tium forever without ever hitting it
  6. http://localhost:4000/ui/?page=new_model offers no Tium entry in the provider dropdown, so the model can only be added by hand as an OpenAI one

After: the same developer names the provider directly and the identical request is priced, so budgets and dashboards work

  1. They add a model to their proxy config called tium/glm-5.3, set TIUM_API_KEY in the environment, and start the proxy on http://localhost:4000
  2. They POST http://localhost:4000/v1/chat/completions with "model": "glm-5.3", the same user message and the same tools array
  3. A 200 comes back with the same tool call
  4. The response now carries x-litellm-response-cost: 0.000537138, broken out into input, output, cache read and reasoning headers
  5. http://localhost:4000/ui/?page=logs shows that request at non-zero spend, so key and team budgets apply to Tium traffic
  6. http://localhost:4000/ui/?page=new_model lists Tium in the provider dropdown, asking for the API key and an optional base URL

Relevant issues

Docs live in the other repo, so the provider page goes up separately in BerriAI/litellm-docs#1388. That one is not a prerequisite for merging this: documentation and code-quality both pass here already. It should still land so the provider has a page.

The openai_compatible_providers entry is deliberate rather than incidental, and relates to #26443. JSONProviderRegistry does not register into that list, and a provider missing from it takes the non-OpenAI-compatible branch of add_provider_specific_params_to_optional_params, so provider-specific body params arrive as top-level kwargs instead of nesting in extra_body, and tools are stripped before the request is sent. That issue is still open against three JSON-only providers.

Linear ticket

Pre-Submission checklist

  • I have added meaningful tests
  • The handful of test files covering my change pass locally
  • My PR passes all required CI/CD checks. The one red check, osv-scan, is a base-branch finding this PR does not cause: smol-toml 1.6.1 (GHSA-7w5x-hrqm-74c2), a dev-only transitive dependency of knip pinned in ui/litellm-dashboard/package-lock.json on litellm_internal_staging. This PR does not touch the lockfile. The advisory is newly published, so the base is affected independently of this PR. Happy to bump it here if you would rather that than a separate fix
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5: scored 5/5, but on the previous revision, so a re-review of the current commit is pending

Type

New Feature

Screenshots / Proof of Fix

Shared setup for both runs, against the live Tium API with a real key.

tium_config.yaml, Before:

model_list:
  - model_name: glm-5.3
    litellm_params:
      model: openai/glm-5.3
      api_base: https://api.tium.ai/v1
      api_key: os.environ/TIUM_API_KEY

tium_config.yaml, After:

model_list:
  - model_name: glm-5.3
    litellm_params:
      model: tium/glm-5.3
      api_key: os.environ/TIUM_API_KEY

Proxy started with LITELLM_LOCAL_MODEL_COST_MAP=True python litellm/proxy/proxy_cli.py --config tium_config.yaml --port 4000

payload.json, one user turn and one tool:

{
  "model": "glm-5.3",
  "messages": [{"role": "user", "content": "What is the weather in Paris?"}],
  "tools": [{"type": "function", "function": {"name": "get_weather", "description": "Get current weather for a location", "parameters": {"type": "object", "properties": {"location": {"type": "string"}}, "required": ["location"]}}}]
}

Before (ee7c7e1)

Cost of a tool-calling chat completion

  1. Send the request and print the cost headers:
$ curl -sS -D headers.txt -o body.json -X POST http://localhost:4000/v1/chat/completions \
    -H "Content-Type: application/json" -H "Authorization: Bearer sk-1234" --data @payload.json
$ grep -i "x-litellm-response-cost" headers.txt
x-litellm-response-cost-original: 0.0
x-litellm-response-cost-discount-amount: 0.0
x-litellm-response-cost-margin-amount: 0.0
x-litellm-response-cost-margin-percent: 0.0
x-litellm-response-cost-input: 0.0
x-litellm-response-cost-output: 0.0
x-litellm-response-cost-tool-usage: 0.0

There is no x-litellm-response-cost header, and every breakdown header is zero

  1. The call itself succeeded, so the zero is a pricing gap and not a failed request:
$ python -c "import json;d=json.load(open('body.json'));print(d['choices'][0]['message']['tool_calls']);print(d['usage'])"
[{'index': 0, 'function': {'arguments': '{"location":"Paris"}', 'name': 'get_weather'}, 'id': 'call_-7273053054465734513', 'type': 'function'}]
{'completion_tokens': 38, 'prompt_tokens': 162, 'total_tokens': 200, 'completion_tokens_details': {'reasoning_tokens': 26}, 'prompt_tokens_details': {'cached_tokens': 128}}
  1. grep -i x-litellm-model-name headers.txt reports openai/glm-5.3, the workaround routing

After (c59b321)

Cost of a tool-calling chat completion

  1. Send the identical request and print the cost headers:
$ curl -sS -D headers.txt -o body.json -X POST http://localhost:4000/v1/chat/completions \
    -H "Content-Type: application/json" -H "Authorization: Bearer sk-1234" --data @payload.json
$ grep -i "x-litellm-response-cost" headers.txt
x-litellm-response-cost: 0.000537138
x-litellm-response-cost-original: 0.000537138
x-litellm-response-cost-discount-amount: 0.0
x-litellm-response-cost-margin-amount: 0.0
x-litellm-response-cost-margin-percent: 0.0
x-litellm-response-cost-input: 9.622000000000001e-05
x-litellm-response-cost-output: 0.00037359000000000003
x-litellm-response-cost-cache-read: 6.7328e-05
x-litellm-response-cost-reasoning: 0.00026685
x-litellm-response-cost-tool-usage: 0.0
  1. The same tool call comes back, now with the usage the price was computed from:
$ python -c "import json;d=json.load(open('body.json'));print(d['choices'][0]['message']['tool_calls']);print(d['usage'])"
[{'index': 0, 'function': {'arguments': '{"location":"Paris"}', 'name': 'get_weather'}, 'id': 'call_-7273047419468643660', 'type': 'function'}]
{'completion_tokens': 42, 'prompt_tokens': 162, 'total_tokens': 204, 'completion_tokens_details': {'reasoning_tokens': 30}, 'prompt_tokens_details': {'cached_tokens': 128}}
  1. grep -i x-litellm-model-name headers.txt reports tium/glm-5.3

  2. The total matches the published rates by hand, so the cost map entry is right rather than merely non-zero: 34 uncached input tokens at $2.830/M is $0.00009622, 128 cached at $0.526/M is $0.000067328, 42 output at $8.895/M is $0.00037359, summing to $0.000537138

Caveats (if any)

Low

  • Sending tools to tium/* fails with a 400 if the code ships ahead of the cost map
    • Capability lookups fetch the cost map from main, so the two normally land together
    • Only bites a build cut from this branch before it reaches main
    • The error is UnsupportedParamsError: tium does not support parameters: ['tools']
    • Workarounds are LITELLM_LOCAL_MODEL_COST_MAP=True or allowed_openai_params=['tools']
    • Offline installs are unaffected, since the bundled backup map is updated here too
  • Only /v1/chat/completions is wired, since the host answers 404 on /v1/responses
  • Context and output limits are the host's ceilings, not the models' own
  • Prices are the subscription rate; prepaid credit packs cost more
  • Both GLM models are thinking-only and refuse reasoning_effort=none
  • DeepSeek models refuse tool_choice="required" while reasoning is on
  • Effort levels are graded only on the GLM models, so only the baseline set is published

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

@curiousbox
curiousbox requested a review from a team September 9, 2026 05:32
@CLAassistant

CLAassistant commented Sep 9, 2026 •

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

@codspeed

codspeed Bot commented Sep 9, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing curiousbox:feat/add-tium-provider (c59b321) with litellm_internal_staging (8ec2f00)1

Open in CodSpeed

Footnotes

  1. No successful run was found on litellm_internal_staging (cb9c3a9) during the generation of this report, so 8ec2f00 was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩

@greptile-apps

greptile-apps Bot commented Sep 9, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

Adds Tium as a JSON-configured OpenAI-compatible provider.

  • Registers provider resolution, endpoint support, credentials, and dashboard metadata.
  • Adds pricing and capability metadata for five Tium chat models.
  • Adds mocked completion coverage for authentication, tools, parameter mapping, and custom API bases.
  • Updates an unrelated auto-router assertion to follow the current Anthropic preset.

Confidence Score: 5/5

The PR appears safe to merge; no actionable regression remains in the changes since the previous review.

Both previous findings were resolved, and the only subsequent change corrects a stale auto-router test expectation while preserving its substantive reasoning-effort assertion.

Important Files Changed

Filename Overview
litellm/llms/openai_like/providers.json Registers Tium’s base URL, environment variables, parameter mapping, and chat-completions endpoint.
model_prices_and_context_window.json Adds pricing, limits, and capability metadata for five Tium models.
tests/test_litellm/llms/openai_like/test_tium_provider.py Covers provider registration and mocked completion behavior, including URL, authentication, tools, and parameter mapping.
ui/litellm-dashboard/src/components/provider_info_helpers.tsx Adds Tium’s dashboard display name, provider slug, logo, and model placeholder.
ui/litellm-dashboard/src/components/add_model/add_auto_router_tab.test.tsx Aligns the expected reasoning-tier model with the current Anthropic preset while retaining the reasoning-effort contract.

Reviews (3): Last reviewed commit: "test(auto-router): derive the REASONING ..." | Re-trigger Greptile

Comment thread tests/test_litellm/llms/openai_like/test_tium_provider.py
Comment thread tests/test_litellm/llms/openai_like/test_tium_provider.py
@codecov

codecov Bot commented Sep 9, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

Tium serves open-weight models (GLM, DeepSeek, Kimi) over an OpenAI-compatible
API from a gateway operated in Germany. Callers currently have to route through
openai/ with a custom api_base

Registered in providers.json with TIUM_API_KEY, the TIUM_API_BASE override, the
max_completion_tokens to max_tokens mapping, and chat completions as the only
supported endpoint, since POST /v1/responses answers 404 on this host. The slug
is also added to openai_compatible_providers and LlmProviders, and the base URL
to openai_compatible_endpoints for api_base autodetection

The openai_compatible_providers entry is load-bearing. JSONProviderRegistry does
not register into that list, and a provider missing from it takes the
non-OpenAI-compatible branch of add_provider_specific_params_to_optional_params,
so provider-specific body params arrive as top-level kwargs instead of nesting
in extra_body, and tools are stripped before the request is sent. See BerriAI#26443,
still open against three JSON-only providers

Five chat models go into the cost map, with each capability flag measured rather
than defaulted. All five support function calling and prompt caching, so each
carries a cache read rate. The two DeepSeek models refuse response_format with a
JSON schema; image input works on glm-5.3-flash and kimi-k3. max_input_tokens and
max_output_tokens are the host's policy ceilings. No constraints block is needed,
since all five accept temperature from 0.0 to 2.0

Costs are USD per token, derived from the published subscription rate of $2.83
per million weighted tokens times each model's multiplier and weights. Tium bills
weighted tokens rather than per-token, and prepaid packs cost more; the docs page
states both

The provider is registered in the Add Model form and the Admin UI provider list,
which test_every_backend_provider_is_listed_in_add_model_or_frozen_as_unlisted
requires of any new entry in LlmProviders

Tests drive real completion calls through respx, asserting the request URL, the
bearer auth header, the message body, that tools and tool_choice reach the wire
and tool calls come back, and that max_completion_tokens is sent as max_tokens.
They use the local_model_cost_map fixture, since the network-fetched cost map
lags this branch until merge and capability lookups would otherwise read stale
data
@curiousbox

Copy link
Copy Markdown
Author

@greptileai

@curiousbox
curiousbox force-pushed the feat/add-tium-provider branch from 18b9cc0 to c59b321 Compare September 10, 2026 02:47
@curiousbox

curiousbox commented Sep 10, 2026 •

Copy link
Copy Markdown
Author

ui-unit-tests was failing here on a break that predated this PR: add_auto_router_tab.test.tsx:1047 expected claude-opus-5 where the Anthropic Family preset's REASONING tier resolves claude-fable-5-1. It went stale in 314e573529 (#40341), which moved the tier and updated autorouter_presets.test.ts but not this file.

That is now fixed upstream by #40456, merged into litellm_internal_staging. CI picks it up through the PR merge ref, so the job is green here without any change to this branch.

Worth recording why this PR surfaced it at all, in case it helps someone else: .github/scripts/select_ui_test_scope.sh runs only related tests when every changed UI file is under src/, and the full suite otherwise. Two of the three UI files here are under src/, but the provider logo lives at ui/litellm-dashboard/public/assets/logos/tium.svg, alongside the other 84 logos, so the full suite ran and reached the broken assertion. Recent merges touching only src/ ran the related subset in about 90 seconds and never got there.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants