Skip to content

chore(cost-map): sync openrouter prices from the models API - #43950

Merged
kerry-berri merged 1 commit into
mainfrom
litellm-providers/price-sync-openrouter
Oct 1, 2026
Merged

kerry-berri merged 1 commit into
mainfrom
litellm-providers/price-sync-openrouter

Conversation

@berriai-litellm-provider-info-sync

@berriai-litellm-provider-info-sync berriai-litellm-provider-info-sync Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Syncs 17 existing OpenRouter rows to the live models API (https://openrouter.ai/api/v1/models). Every row below is kind api: prices come from pricing.prompt, pricing.completion, pricing.input_cache_read and pricing.overrides, context window and max output from context_length and top_provider.max_completion_tokens, and the deprecation date from expiration_date

Rows

`openrouter/deepseek/deepseek-v3.1-terminus` (api)
  input_cost_per_token: 2.7e-07 -> 3e-07
  max_output_tokens: 32768 -> 65536
  max_tokens: 32768 -> 65536
  evidence: GET /api/v1/models id "deepseek/deepseek-v3.1-terminus" pricing.completion = "0.000001" / GET /api/v1/models id "deepseek/deepseek-v3.1-terminus" pricing.input_cache_read = "0.000000135" / GET /api/v1/models id "deepseek/deepseek-v3.1-terminus" pricing.prompt = "0.0000003"

`openrouter/deepseek/deepseek-v4-flash` (api)
  cache_read_input_token_cost: 2.8e-08 -> 1.5708e-08
  input_cost_per_token: 1.4e-07 -> 7.854e-08
  output_cost_per_token: 2.8e-07 -> 1.5708e-07
  evidence: GET /api/v1/models id "deepseek/deepseek-v4-flash" pricing.completion = "0.00000015708" / GET /api/v1/models id "deepseek/deepseek-v4-flash" pricing.input_cache_read = "0.000000015708" / GET /api/v1/models id "deepseek/deepseek-v4-flash" pricing.prompt = "0.00000007854"

`openrouter/deepseek/deepseek-v4-flash-0731` (api)
  cache_read_input_token_cost: 1.6e-08 -> 8.9e-09
  input_cost_per_token: 1.8e-08 -> 8.9e-09
  max_input_tokens: 1310720 -> 1048576
  output_cost_per_token: 3.2e-07 -> 1.28e-06
  evidence: GET /api/v1/models id "deepseek/deepseek-v4-flash-0731" pricing.completion = "0.00000128" / GET /api/v1/models id "deepseek/deepseek-v4-flash-0731" pricing.input_cache_read = "0.0000000089" / GET /api/v1/models id "deepseek/deepseek-v4-flash-0731" pricing.prompt = "0.0000000089"

`openrouter/deepseek/deepseek-v4-pro` (api)
  cache_read_input_token_cost: 7.9025e-08 -> 6.525e-08
  input_cost_per_token: 9.483e-07 -> 7.83e-07
  output_cost_per_token: 1.8966e-06 -> 1.566e-06
  evidence: GET /api/v1/models id "deepseek/deepseek-v4-pro" pricing.completion = "0.000001566" / GET /api/v1/models id "deepseek/deepseek-v4-pro" pricing.input_cache_read = "0.00000006525" / GET /api/v1/models id "deepseek/deepseek-v4-pro" pricing.prompt = "0.000000783"

`openrouter/deepseek/deepseek-v4-pro-0813` (api)
  cache_read_input_token_cost: 1.72e-07 -> 4.4e-08
  input_cost_per_token: 2.4298e-07 -> 1.32e-06
  max_output_tokens: 943718 -> 393216
  max_tokens: 943718 -> 393216
  off_peak_pricing: None -> {"input_cost_per_token": 6.6e-07, "output_cost_per_token": 1.98e-06, "cache_read_input_token_cost": 2.2e-08, "windows": [{"hours_utc": "00:00-00:00", "weekdays": ["saturday", "sunday"]}, {"hours_utc": "00:00-01:00", "weekdays": ["monday", "tuesday", "wednesday", "thursday", "friday"]}, {"hours_utc": "04:00-06:00", "weekdays": ["monday", "tuesday", "wednesday", "thursday", "friday"]}, {"hours_utc": "10:00-00:00", "weekdays": ["monday", "tuesday", "wednesday", "thursday", "friday"]}]}
  output_cost_per_token: 4.2e-06 -> 3.96e-06
  evidence: GET /api/v1/models id "deepseek/deepseek-v4-pro-0813" pricing.overrides[2].completion = "0.00000396" / GET /api/v1/models id "deepseek/deepseek-v4-pro-0813" pricing.overrides[2].input_cache_read = "0.000000044" / GET /api/v1/models id "deepseek/deepseek-v4-pro-0813" pricing.overrides[2].prompt = "0.00000132" / GET /api/v1/models id "deepseek/deepseek-v4-pro-0813" pricing.overrides = "[{"utc_days":["saturday","sunday"],"prompt":"0.00000066","completion":"0.00000198","input_cache_read":"0.000000022"},{"utc_days":["monday","tuesday","wednesday","thursday","friday"],"utc_start":0,"utc_end":100,"prompt":"0.00000066","completion":"0.00000198","input_cache_read":"0.000000022"},{"utc_days":["monday","tuesday","wednesday","thursday","friday"],"utc_start":400,"utc_end":600,"prompt":"0.00000066","completion":"0.00000198","input_cache_read":"0.000000022"},{"utc_days":["monday","tuesday","wednesday","thursday","friday"],"utc_start":1000,"utc_end":0,"prompt":"0.00000066","completion":"0.00000198","input_cache_read":"0.000000022"}]"

`openrouter/deepseek/deepseek-v4.1-flash` (api)
  cache_read_input_token_cost: 6e-09 -> 2.91e-09
  input_cost_per_token: 3e-07 -> 1.98e-08
  output_cost_per_token: 1.2e-06 -> 3.96e-07
  evidence: GET /api/v1/models id "deepseek/deepseek-v4.1-flash" pricing.completion = "0.000000396" / GET /api/v1/models id "deepseek/deepseek-v4.1-flash" pricing.input_cache_read = "0.00000000291" / GET /api/v1/models id "deepseek/deepseek-v4.1-flash" pricing.prompt = "0.0000000198"

`openrouter/google/gemma-4-26b-a4b-it` (api)
  cache_read_input_token_cost: 3.75e-08 -> 4.25e-08
  input_cost_per_token: 6.75e-08 -> 7.65e-08
  output_cost_per_token: 2.25e-07 -> 2.55e-07
  evidence: GET /api/v1/models id "google/gemma-4-26b-a4b-it" pricing.completion = "0.000000255" / GET /api/v1/models id "google/gemma-4-26b-a4b-it" pricing.input_cache_read = "0.0000000425" / GET /api/v1/models id "google/gemma-4-26b-a4b-it" pricing.prompt = "0.0000000765"

`openrouter/meta/muse-glimmer-30b` (api)
  input_cost_per_token: 3e-07 -> 3.5e-07
  max_output_tokens: 16384 -> 117964
  max_tokens: 16384 -> 117964
  output_cost_per_token: 1.2e-06 -> 1.5e-06
  evidence: GET /api/v1/models id "meta/muse-glimmer-30b" pricing.completion = "0.0000015" / GET /api/v1/models id "meta/muse-glimmer-30b" pricing.input_cache_read = "0.00000004" / GET /api/v1/models id "meta/muse-glimmer-30b" pricing.prompt = "0.00000035"

`openrouter/minimax/minimax-m2.7` (api)
  cache_read_input_token_cost: 6e-08 -> 4.2e-08
  input_cost_per_token: 3e-07 -> 2.1e-07
  output_cost_per_token: 1.2e-06 -> 8.4e-07
  evidence: GET /api/v1/models id "minimax/minimax-m2.7" pricing.completion = "0.00000084" / GET /api/v1/models id "minimax/minimax-m2.7" pricing.input_cache_read = "0.000000042" / GET /api/v1/models id "minimax/minimax-m2.7" pricing.prompt = "0.00000021"

`openrouter/moonshotai/kimi-k2.7-code` (api)
  input_cost_per_token: 6.562e-07 -> 6.712e-07
  output_cost_per_token: 3.3e-06 -> 3.35e-06
  evidence: GET /api/v1/models id "moonshotai/kimi-k2.7-code" pricing.completion = "0.00000335" / GET /api/v1/models id "moonshotai/kimi-k2.7-code" pricing.input_cache_read = "0.00000018" / GET /api/v1/models id "moonshotai/kimi-k2.7-code" pricing.prompt = "0.0000006712"

`openrouter/moonshotai/kimi-k3` (api)
  cache_read_input_token_cost: 3e-07 -> 2.7e-07
  input_cost_per_token: 3e-06 -> 2.8e-07
  output_cost_per_token: 1.5e-05 -> 1e-05
  evidence: GET /api/v1/models id "moonshotai/kimi-k3" pricing.completion = "0.00001" / GET /api/v1/models id "moonshotai/kimi-k3" pricing.input_cache_read = "0.00000027" / GET /api/v1/models id "moonshotai/kimi-k3" pricing.prompt = "0.00000028"

`openrouter/openai/gpt-oss-120b` (api)
  cache_read_input_token_cost: 7.5e-08 -> removed (no input_cache_read published, ruling openrouter-gpt-oss-120b-no-cached-input-price)
  input_cost_per_token: 1.5e-07 -> 3.7e-08
  max_output_tokens: 65536 -> 117964
  max_tokens: 65536 -> 117964
  output_cost_per_token: 6e-07 -> 1.7e-07
  evidence: GET /api/v1/models id "openai/gpt-oss-120b" pricing.completion = "0.00000017" / GET /api/v1/models id "openai/gpt-oss-120b" pricing.prompt = "0.000000037"

`openrouter/qwen/qwen3-30b-a3b-instruct-2507` (api)
  input_cost_per_token: 1e-07 -> 4.815e-08
  max_output_tokens: 235929 -> 32000
  max_tokens: 235929 -> 32000
  output_cost_per_token: 3e-07 -> 1.9305e-07
  evidence: GET /api/v1/models id "qwen/qwen3-30b-a3b-instruct-2507" pricing.completion = "0.00000019305" / GET /api/v1/models id "qwen/qwen3-30b-a3b-instruct-2507" pricing.prompt = "0.00000004815"

`openrouter/stealth/space-bunny-alpha` (api)
  deprecation_date: 2098-12-31 -> 2026-10-05
  max_output_tokens: 128000 -> 524288
  max_tokens: 128000 -> 524288
  evidence: GET /api/v1/models id "stealth/space-bunny-alpha" pricing.prompt = "0" / GET /api/v1/models id "stealth/space-bunny-alpha" pricing.completion = "0" / GET /api/v1/models id "stealth/space-bunny-alpha" expiration_date = "2026-10-05"

`openrouter/tencent/hy4-preview` (api)
  off_peak_pricing: None -> {"input_cost_per_token": 7.506e-07, "output_cost_per_token": 2.2509e-06, "cache_read_input_token_cost": 3.78e-08, "hours_utc": "16:00-00:00"}
  evidence: GET /api/v1/models id "tencent/hy4-preview" pricing.overrides[0].completion = "0.000002501" / GET /api/v1/models id "tencent/hy4-preview" pricing.overrides[0].input_cache_read = "0.000000042" / GET /api/v1/models id "tencent/hy4-preview" pricing.overrides[0].prompt = "0.000000834" / GET /api/v1/models id "tencent/hy4-preview" pricing.overrides = "[{"utc_start":1600,"utc_end":0,"prompt":"0.0000007506","completion":"0.0000022509","input_cache_read":"0.0000000378"}]"

`openrouter/z-ai/glm-5.1` (api)
  cache_read_input_token_cost: 2.6e-07 -> 1.7914e-07
  input_cost_per_token: 1.4e-06 -> 9.646e-07
  output_cost_per_token: 4.4e-06 -> 3.0316e-06
  evidence: GET /api/v1/models id "z-ai/glm-5.1" pricing.completion = "0.0000030316" / GET /api/v1/models id "z-ai/glm-5.1" pricing.input_cache_read = "0.00000017914" / GET /api/v1/models id "z-ai/glm-5.1" pricing.prompt = "0.0000009646"

`openrouter/z-ai/glm-5.2` (api)
  cache_read_input_token_cost: 1.2064e-07 -> 2.6e-07
  input_cost_per_token: 6.496e-07 -> 3.249e-07
  max_output_tokens: 131072 -> 943718
  max_tokens: 131072 -> 943718
  output_cost_per_token: 2.0416e-06 -> 3.99e-06
  evidence: GET /api/v1/models id "z-ai/glm-5.2" pricing.completion = "0.00000399" / GET /api/v1/models id "z-ai/glm-5.2" pricing.input_cache_read = "0.00000026" / GET /api/v1/models id "z-ai/glm-5.2" pricing.prompt = "0.0000003249"

Notes

openrouter/openai/gpt-oss-120b leaves cache_read_input_token_cost unset per ruling openrouter-gpt-oss-120b-no-cached-input-price: the API no longer returns pricing.input_cache_read for this id, so no cached-input price is published. This PR is sent in replace mode with full rows so that field is dropped

openrouter/deepseek/deepseek-v4-pro-0813 base prices are the peak rates from pricing.overrides and off_peak_pricing carries the discounted windows, the same weekday window shape openrouter/~deepseek/deepseek-pro-latest already uses

openrouter/tencent/hy3 off_peak_pricing already matches the API (the diff only saw a string versus object encoding), so it is unchanged

Rulings followed: openrouter-gpt-oss-120b-no-cached-input-price, openrouter-merge-despite-baseline-ci-failures, openrouter-deepseek-v4-pro-drop-legacy-cache-hit and openrouter-deepseek-v4-pro-0813-drop-legacy-cache-hit; neither row carries input_cost_per_token_cache_hit any longer, and this PR does not reintroduce it

The changed https://openrouter.ai/announcements page lists blog posts only, with no model launch, price change or deprecation, so it needs no rows

Delisted by the provider

None


Note

Medium Risk
Incorrect token costs or limits directly affect LiteLLM billing and budget estimates; changes are API-sourced but some rows shift dramatically (pricing and max tokens).

Overview
Updates 17 OpenRouter model entries in model_prices_and_context_window.json (and its backup) to match the live OpenRouter models API: per-token input/output/cache-read rates, max_input_tokens / max_output_tokens, and capability flags where the sync reordered or reaffirmed them.

Notable non-price changes: openrouter/deepseek/deepseek-v4-pro-0813 and openrouter/tencent/hy4-preview gain off_peak_pricing (time-window discounts used by cost calculation); openrouter/openai/gpt-oss-120b drops cache_read_input_token_cost because the API no longer publishes cached-input pricing; openrouter/stealth/space-bunny-alpha gets a real deprecation_date (2026-10-05) and higher output limits. Several models see large pricing corrections (e.g. kimi-k3 input cost) or tighter output caps (e.g. qwen3-30b-a3b-instruct-2507).

Reviewed by Cursor Bugbot for commit 2737f77. Bugbot is set up for automated code reviews on this repo. Configure here.

@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

@greptile-apps

greptile-apps Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

[Medium risk] Updates pricing and token limits for external API models.

No new actionable issue in the current PR diff was identified.

Summary

Synchronizes OpenRouter pricing, token limits, deprecation metadata, and off-peak pricing across the primary and backup cost maps. The previously reported cached-batch pricing concern was reintroduced by the removal of the cache-read price, but its thread is resolved, so it is not reposted.

Reviews (3) · Last reviewed commit: "chore(cost-map): sync openrouter prices ..."

Comment thread model_prices_and_context_window.json

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@berriai-litellm-provider-info-sync
berriai-litellm-provider-info-sync Bot force-pushed the litellm-providers/price-sync-openrouter branch from cde0e62 to 7b9838f Compare October 1, 2026 00:18
@codecov

codecov Bot commented Oct 1, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@codspeed

codspeed Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm-providers/price-sync-openrouter (2737f77) with main (38b0762)

Open in CodSpeed

@berriai-litellm-provider-info-sync
berriai-litellm-provider-info-sync Bot force-pushed the litellm-providers/price-sync-openrouter branch from 7b9838f to 2737f77 Compare October 1, 2026 00:40

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 2737f77. Configure here.

@kerry-berri
kerry-berri merged commit 8a1f356 into main Oct 1, 2026
92 of 98 checks passed
jan-sauer-reef added a commit to jan-sauer-reef/litellm that referenced this pull request Oct 1, 2026
…ject_key_prefix

* upstream/main: (62 commits)
  fix(guardrails): scan Responses API input in Azure Prompt Shield (BerriAI#43786)
  feat(lens): investigate sampled traces and retain batch results (BerriAI#43942)
  fix(proxy): restore pre-config-wins handling of pass-through endpoints (BerriAI#43962)
  fix(cost-map): raise baseten DeepSeek-V4.1-Flash max output to 262144 (BerriAI#43916)
  chore(cost-map): add deprecation date for anthropic claude-sonnet-4-5 (BerriAI#43898)
  chore(cost-map): add fireworks inkling priority prices from the prices api (BerriAI#43949)
  feat(guardrails): honor litellm_params.timeout in every HTTP guardrail (BerriAI#43134)
  test(e2e): typed per-test metadata for the e2e suite (BerriAI#42044)
  fix(caching): write the response-cache SET to Redis at once instead of on the post-call batch (BerriAI#43973)
  feat(ui): filter tags by name and description on the Tag Management page (BerriAI#42949)
  feat(providers): add Cortecs as an OpenAI-compatible provider (BerriAI#43872)
  feat(e2e): record each e2e test's steps, starting with ProxyClient (BerriAI#42393)
  test(ci): repair stale tests and move retired OpenAI text-completion fixtures (BerriAI#43958)
  feat(proxy): record in spend logs whether a request used a client-forwarded Anthropic OAuth token (BerriAI#43063)
  fix(azure_storage): keep the DataLakeServiceClient alive until its TTL elapses (BerriAI#43082)
  chore(deps): bump gitpython and tornado, extend diskcache osv ignore to Nov 1 (BerriAI#43961)
  fix(guardrails): treat an unknown straiker api_version as unset instead of skipping the guardrail (BerriAI#43956)
  fix(azure_storage): name Data Lake objects without base64 padding or slashes (BerriAI#43914)
  fix(grayswan): send request conversation and tool calls to post-call monitor (BerriAI#43770)
  chore(cost-map): sync openrouter prices from the models API (BerriAI#43950)
  ...
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant