Skip to content

fix(cost-map): raise baseten DeepSeek-V4.1-Flash max output to 262144 - #43916

Merged
kerry-berri merged 1 commit into
mainfrom
litellm-providers/price-sync-baseten
Oct 1, 2026
Merged

kerry-berri merged 1 commit into
mainfrom
litellm-providers/price-sync-baseten

Conversation

@berriai-litellm-provider-info-sync

@berriai-litellm-provider-info-sync berriai-litellm-provider-info-sync Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Raises the max output limit on baseten/deepseek-ai/DeepSeek-V4.1-Flash to match the Baseten models API, which now lists 262144 output tokens instead of 32768

Rows

baseten/deepseek-ai/DeepSeek-V4.1-Flash (kind api)

Limits from GET https://inference.baseten.co/v1/models via the providers list API: contextWindow 1048576 (max_input_tokens, unchanged), maxOutput 262144 (max_output_tokens and max_tokens, was 32768)

Prices from the same endpoint via the providers price API already match the catalog row and are resent unchanged:

  • input_cost_per_token 3e-07 from pricing.prompt = 0.0000003
  • output_cost_per_token 1.2e-06 from pricing.completion = 0.0000012
  • cache_read_input_token_cost 7e-09 from pricing.input_cache_read = 0.000000007

Changed page

https://inference.baseten.co/v1/models changed; every other Baseten id it lists already matches its catalog row in price, context window and max output

Delisted by the provider

baseten/deepseek-ai/DeepSeek-V4.1-Flash-Fast is no longer returned by the models API and no deprecation date is stated, so the key is kept for a human to decide. It was added in #43735 under ruling baseten-deepseek-v4-1-flash-fast-merge-despite-main-test-failure


Note

Low Risk
Metadata-only limit correction in the model price map; no runtime logic changes.

Overview
Updates the LiteLLM cost/context catalog for baseten/deepseek-ai/DeepSeek-V4.1-Flash so max_output_tokens and max_tokens go from 32768 to 262144, matching Baseten’s models API. The same change is applied in model_prices_and_context_window.json and litellm/model_prices_and_context_window_backup.json; pricing and max_input_tokens (1M) are unchanged.

Reviewed by Cursor Bugbot for commit 4a21007. Bugbot is set up for automated code reviews on this repo. Configure here.

@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

@greptile-apps

greptile-apps Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 4/5

[Medium risk] Updates model capability limits for a third-party LLM.

The PR should not merge until the larger catalog limit no longer causes output-token-limited deployments to reject otherwise valid requests.

Findings

  1. P1 Requests rejected by token limit ▶

Summary

Raises the cataloged output limit for Baseten DeepSeek-V4.1-Flash from 32768 to 262144 in the primary and backup price maps, while leaving its input limit and prices unchanged.

  • The larger limit also increases pre-call output-token reservations for requests without an explicit output limit.

Reviews (1) · Last reviewed commit: "fix(cost-map): raise baseten DeepSeek-V4..."

Comment on lines +77204 to +77205
"max_output_tokens": 262144,
"max_tokens": 262144,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Requests rejected by token limit

If a deployment has an output-token-per-minute limit below 262144, a request without an explicit max_tokens now reserves the full 262144 tokens before it runs. The pre-call check rejects the request even if it would produce a short response. For example, a deployment limited to 100000 output tokens per minute previously admitted such requests but now rejects them all. The backup catalog mirrors this change.

Knowledge Base Used: Provider adapters and capabilities

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That’s fair. This row should reflect Baseten’s current maxOutput, and the router’s OTPM reservation behavior is separate from the catalog correction. I’m withdrawing the finding for this PR; any change to avoid reserving the full catalog maximum belongs in the router’s rate-limit logic, with its own tests, rather than in this metadata update.

@codecov

codecov Bot commented Sep 30, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@codspeed

codspeed Bot commented Sep 30, 2026

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm-providers/price-sync-baseten (4a21007) with main (50f5cc9)1

Open in CodSpeed

Footnotes

  1. No successful run was found on main (405ed41) during the generation of this report, so 50f5cc9 was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩

@kerry-berri
kerry-berri merged commit 2b19ddb into main Oct 1, 2026
98 of 100 checks passed
jan-sauer-reef added a commit to jan-sauer-reef/litellm that referenced this pull request Oct 1, 2026
…ject_key_prefix

* upstream/main: (62 commits)
  fix(guardrails): scan Responses API input in Azure Prompt Shield (BerriAI#43786)
  feat(lens): investigate sampled traces and retain batch results (BerriAI#43942)
  fix(proxy): restore pre-config-wins handling of pass-through endpoints (BerriAI#43962)
  fix(cost-map): raise baseten DeepSeek-V4.1-Flash max output to 262144 (BerriAI#43916)
  chore(cost-map): add deprecation date for anthropic claude-sonnet-4-5 (BerriAI#43898)
  chore(cost-map): add fireworks inkling priority prices from the prices api (BerriAI#43949)
  feat(guardrails): honor litellm_params.timeout in every HTTP guardrail (BerriAI#43134)
  test(e2e): typed per-test metadata for the e2e suite (BerriAI#42044)
  fix(caching): write the response-cache SET to Redis at once instead of on the post-call batch (BerriAI#43973)
  feat(ui): filter tags by name and description on the Tag Management page (BerriAI#42949)
  feat(providers): add Cortecs as an OpenAI-compatible provider (BerriAI#43872)
  feat(e2e): record each e2e test's steps, starting with ProxyClient (BerriAI#42393)
  test(ci): repair stale tests and move retired OpenAI text-completion fixtures (BerriAI#43958)
  feat(proxy): record in spend logs whether a request used a client-forwarded Anthropic OAuth token (BerriAI#43063)
  fix(azure_storage): keep the DataLakeServiceClient alive until its TTL elapses (BerriAI#43082)
  chore(deps): bump gitpython and tornado, extend diskcache osv ignore to Nov 1 (BerriAI#43961)
  fix(guardrails): treat an unknown straiker api_version as unset instead of skipping the guardrail (BerriAI#43956)
  fix(azure_storage): name Data Lake objects without base64 padding or slashes (BerriAI#43914)
  fix(grayswan): send request conversation and tool calls to post-call monitor (BerriAI#43770)
  chore(cost-map): sync openrouter prices from the models API (BerriAI#43950)
  ...
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant