Repository navigation
chore(cost-map): add fireworks inkling priority prices from the prices api - #43949
Conversation
…s api Price-Sync: litellm-providers
|
|
|
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit f533413. Configure here.
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
…ject_key_prefix * upstream/main: (62 commits) fix(guardrails): scan Responses API input in Azure Prompt Shield (BerriAI#43786) feat(lens): investigate sampled traces and retain batch results (BerriAI#43942) fix(proxy): restore pre-config-wins handling of pass-through endpoints (BerriAI#43962) fix(cost-map): raise baseten DeepSeek-V4.1-Flash max output to 262144 (BerriAI#43916) chore(cost-map): add deprecation date for anthropic claude-sonnet-4-5 (BerriAI#43898) chore(cost-map): add fireworks inkling priority prices from the prices api (BerriAI#43949) feat(guardrails): honor litellm_params.timeout in every HTTP guardrail (BerriAI#43134) test(e2e): typed per-test metadata for the e2e suite (BerriAI#42044) fix(caching): write the response-cache SET to Redis at once instead of on the post-call batch (BerriAI#43973) feat(ui): filter tags by name and description on the Tag Management page (BerriAI#42949) feat(providers): add Cortecs as an OpenAI-compatible provider (BerriAI#43872) feat(e2e): record each e2e test's steps, starting with ProxyClient (BerriAI#42393) test(ci): repair stale tests and move retired OpenAI text-completion fixtures (BerriAI#43958) feat(proxy): record in spend logs whether a request used a client-forwarded Anthropic OAuth token (BerriAI#43063) fix(azure_storage): keep the DataLakeServiceClient alive until its TTL elapses (BerriAI#43082) chore(deps): bump gitpython and tornado, extend diskcache osv ignore to Nov 1 (BerriAI#43961) fix(guardrails): treat an unknown straiker api_version as unset instead of skipping the guardrail (BerriAI#43956) fix(azure_storage): name Data Lake objects without base64 padding or slashes (BerriAI#43914) fix(grayswan): send request conversation and tool calls to post-call monitor (BerriAI#43770) chore(cost-map): sync openrouter prices from the models API (BerriAI#43950) ...
TLDR
Problem this solves:
How it solves it:
*_priorityprices tofireworks_ai/accounts/fireworks/models/inklingfrom the prices APIUser Flow
Before: a cost map reader looking up priority rates for Inkling on Fireworks finds none
fireworks_ai/accounts/fireworks/models/inklingAfter: the same lookup returns the published priority rates, like the other Fireworks priority rows
Rows
fireworks-ai sync under ruling
fireworks-ai-priority-prices-add-api-first, which says to add priority prices and take them from the prices API when it states themfireworks_ai/accounts/fireworks/models/inkling(kind: api), source https://api.fireworks.ai/v1/serverless/models?format=nestedserverlessModes/prioritysku "LLM input tokens (uncached)" = 1 per 1M tokens, setsinput_cost_per_token_priority1e-06serverlessModes/prioritysku "LLM input tokens (cached)" = 0.17 per 1M tokens, setscache_read_input_token_cost_priority1.7e-07serverlessModes/prioritysku "LLM output tokens" = 4.05 per 1M tokens, setsoutput_cost_per_token_priority4.05e-06defaultmode skus are 1 / 0.17 / 4.05 per 1M, matching the standard prices already on the row, sosourcemoves to the prices API url. The pricing page table does not list Inkling. Context window 1048576 matches the list rowEvery other prices API record matches the catalog. The 2026-10-01 changelog entry says "On October 1, 2026 at 00:00 UTC, serverless pricing for DeepSeek V4.1 Flash changes" to $0.30 / $0.006 / $1.20 (priority $0.375 / $0.0075 / $1.50) and the pricing page shows the same, but the prices API still states $0.22 / $0.007 / $0.66, so
deepseek-v4p1-flashstays at the API value per rulingfireworks-ai-deepseek-v4p1-flash-base-use-api-priceThe changed blog page adds "Inside Fireworks Multi-region Deployments", which states no per-token prices, and "Introducing FireRouter with Opus", left out under ruling
fireworks-ai-firerouter-opus-kerry-separate-prDelisted by the provider
glm-5p2-fastandglm-5p2-fast-usalready carrydeprecation_date2026-09-25 andkimi-k2p6-fastandkimi-k2p7-code-fastalready carry 2026-08-27 (rulingfireworks-ai-changelog-deprecation-dates).glm-5p1-fastis kept under rulingfireworks-ai-delisted-no-date-keepNaming or coverage gaps rather than delistings:
fireworks_ai/glm-5p3-fastandfireworks_ai/kimi-k3-fastare listed by the API under theiraccounts/fireworks/routers/form;fireworks_ai/glm-5p3-us,fireworks_ai/glm-5p3-flash-us,fireworks_ai/kimi-k3-usandfireworks_ai/deepseek-v4p1-flash-usare US-only endpoints the model list does not expose but the pricing page still prices; the size-tier keys (fireworks-ai-up-to-4b,fireworks-ai-4.1b-to-16b,fireworks-ai-above-16b,fireworks-ai-moe-up-to-56b,fireworks-ai-56b-to-176b,fireworks-ai-default,fireworks-ai-embedding-up-to-150m,fireworks-ai-embedding-150m-to-350m) are never on the model list by designPre-Submission checklist
Note
Low Risk
Cost-map metadata only; no runtime logic, auth, or billing code changes.
Overview
Adds Fireworks Inkling priority-tier cost fields (
input_cost_per_token_priority,cache_read_input_token_cost_priority,output_cost_per_token_priority) so lookups align with other Fireworks serverless models that expose priority pricing.The new priority rates match the existing standard input, cache-read, and output values from the Fireworks prices API. The row’s
sourceis updated tohttps://api.fireworks.ai/v1/serverless/models?format=nested. The same edit is applied in the main cost map andlitellm/model_prices_and_context_window_backup.json.Reviewed by Cursor Bugbot for commit f533413. Bugbot is set up for automated code reviews on this repo. Configure here.