Repository navigation
fix(cost-map): raise baseten DeepSeek-V4.1-Flash max output to 262144 - #43916
Conversation
Price-Sync: litellm-providers
|
|
|
| "max_output_tokens": 262144, | ||
| "max_tokens": 262144, |
There was a problem hiding this comment.
Requests rejected by token limit
If a deployment has an output-token-per-minute limit below 262144, a request without an explicit max_tokens now reserves the full 262144 tokens before it runs. The pre-call check rejects the request even if it would produce a short response. For example, a deployment limited to 100000 output tokens per minute previously admitted such requests but now rejects them all. The backup catalog mirrors this change.
Knowledge Base Used: Provider adapters and capabilities
There was a problem hiding this comment.
That’s fair. This row should reflect Baseten’s current maxOutput, and the router’s OTPM reservation behavior is separate from the catalog correction. I’m withdrawing the finding for this PR; any change to avoid reserving the full catalog maximum belongs in the router’s rate-limit logic, with its own tests, rather than in this metadata update.
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
…ject_key_prefix * upstream/main: (62 commits) fix(guardrails): scan Responses API input in Azure Prompt Shield (BerriAI#43786) feat(lens): investigate sampled traces and retain batch results (BerriAI#43942) fix(proxy): restore pre-config-wins handling of pass-through endpoints (BerriAI#43962) fix(cost-map): raise baseten DeepSeek-V4.1-Flash max output to 262144 (BerriAI#43916) chore(cost-map): add deprecation date for anthropic claude-sonnet-4-5 (BerriAI#43898) chore(cost-map): add fireworks inkling priority prices from the prices api (BerriAI#43949) feat(guardrails): honor litellm_params.timeout in every HTTP guardrail (BerriAI#43134) test(e2e): typed per-test metadata for the e2e suite (BerriAI#42044) fix(caching): write the response-cache SET to Redis at once instead of on the post-call batch (BerriAI#43973) feat(ui): filter tags by name and description on the Tag Management page (BerriAI#42949) feat(providers): add Cortecs as an OpenAI-compatible provider (BerriAI#43872) feat(e2e): record each e2e test's steps, starting with ProxyClient (BerriAI#42393) test(ci): repair stale tests and move retired OpenAI text-completion fixtures (BerriAI#43958) feat(proxy): record in spend logs whether a request used a client-forwarded Anthropic OAuth token (BerriAI#43063) fix(azure_storage): keep the DataLakeServiceClient alive until its TTL elapses (BerriAI#43082) chore(deps): bump gitpython and tornado, extend diskcache osv ignore to Nov 1 (BerriAI#43961) fix(guardrails): treat an unknown straiker api_version as unset instead of skipping the guardrail (BerriAI#43956) fix(azure_storage): name Data Lake objects without base64 padding or slashes (BerriAI#43914) fix(grayswan): send request conversation and tool calls to post-call monitor (BerriAI#43770) chore(cost-map): sync openrouter prices from the models API (BerriAI#43950) ...
Raises the max output limit on
baseten/deepseek-ai/DeepSeek-V4.1-Flashto match the Baseten models API, which now lists 262144 output tokens instead of 32768Rows
baseten/deepseek-ai/DeepSeek-V4.1-Flash(kind api)Limits from
GET https://inference.baseten.co/v1/modelsvia the providers list API:contextWindow1048576 (max_input_tokens, unchanged),maxOutput262144 (max_output_tokens and max_tokens, was 32768)Prices from the same endpoint via the providers price API already match the catalog row and are resent unchanged:
input_cost_per_token3e-07 frompricing.prompt = 0.0000003output_cost_per_token1.2e-06 frompricing.completion = 0.0000012cache_read_input_token_cost7e-09 frompricing.input_cache_read = 0.000000007Changed page
https://inference.baseten.co/v1/modelschanged; every other Baseten id it lists already matches its catalog row in price, context window and max outputDelisted by the provider
baseten/deepseek-ai/DeepSeek-V4.1-Flash-Fastis no longer returned by the models API and no deprecation date is stated, so the key is kept for a human to decide. It was added in #43735 under rulingbaseten-deepseek-v4-1-flash-fast-merge-despite-main-test-failureNote
Low Risk
Metadata-only limit correction in the model price map; no runtime logic changes.
Overview
Updates the LiteLLM cost/context catalog for
baseten/deepseek-ai/DeepSeek-V4.1-Flashsomax_output_tokensandmax_tokensgo from 32768 to 262144, matching Baseten’s models API. The same change is applied inmodel_prices_and_context_window.jsonandlitellm/model_prices_and_context_window_backup.json; pricing andmax_input_tokens(1M) are unchanged.Reviewed by Cursor Bugbot for commit 4a21007. Bugbot is set up for automated code reviews on this repo. Configure here.