Skip to content

chore(cost-map): add fireworks inkling priority prices from the prices api - #43949

Merged
kerry-berri merged 1 commit into
mainfrom
litellm-providers/price-sync-fireworks-ai
Oct 1, 2026
Merged

kerry-berri merged 1 commit into
mainfrom
litellm-providers/price-sync-fireworks-ai

Conversation

@berriai-litellm-provider-info-sync

@berriai-litellm-provider-info-sync berriai-litellm-provider-info-sync Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Fireworks Inkling lacked the priority tier prices the Fireworks prices API now publishes

How it solves it:

  • Adds the three *_priority prices to fireworks_ai/accounts/fireworks/models/inkling from the prices API
  • Same edit in the main cost map and its backup copy

User Flow

Before: a cost map reader looking up priority rates for Inkling on Fireworks finds none

  1. They GET https://litellm-domain/model/info for fireworks_ai/accounts/fireworks/models/inkling
  2. The response lists only the standard $1.00 in, $0.17 cached and $4.05 out per million

After: the same lookup returns the published priority rates, like the other Fireworks priority rows

  1. They GET the same https://litellm-domain/model/info for the same model
  2. The response also lists priority input at $1.00, cached at $0.17 and output at $4.05 per million

Rows

fireworks-ai sync under ruling fireworks-ai-priority-prices-add-api-first, which says to add priority prices and take them from the prices API when it states them

  • fireworks_ai/accounts/fireworks/models/inkling (kind: api), source https://api.fireworks.ai/v1/serverless/models?format=nested
    • serverlessModes/priority sku "LLM input tokens (uncached)" = 1 per 1M tokens, sets input_cost_per_token_priority 1e-06
    • serverlessModes/priority sku "LLM input tokens (cached)" = 0.17 per 1M tokens, sets cache_read_input_token_cost_priority 1.7e-07
    • serverlessModes/priority sku "LLM output tokens" = 4.05 per 1M tokens, sets output_cost_per_token_priority 4.05e-06
    • The default mode skus are 1 / 0.17 / 4.05 per 1M, matching the standard prices already on the row, so source moves to the prices API url. The pricing page table does not list Inkling. Context window 1048576 matches the list row

Every other prices API record matches the catalog. The 2026-10-01 changelog entry says "On October 1, 2026 at 00:00 UTC, serverless pricing for DeepSeek V4.1 Flash changes" to $0.30 / $0.006 / $1.20 (priority $0.375 / $0.0075 / $1.50) and the pricing page shows the same, but the prices API still states $0.22 / $0.007 / $0.66, so deepseek-v4p1-flash stays at the API value per ruling fireworks-ai-deepseek-v4p1-flash-base-use-api-price

The changed blog page adds "Inside Fireworks Multi-region Deployments", which states no per-token prices, and "Introducing FireRouter with Opus", left out under ruling fireworks-ai-firerouter-opus-kerry-separate-pr

Delisted by the provider

glm-5p2-fast and glm-5p2-fast-us already carry deprecation_date 2026-09-25 and kimi-k2p6-fast and kimi-k2p7-code-fast already carry 2026-08-27 (ruling fireworks-ai-changelog-deprecation-dates). glm-5p1-fast is kept under ruling fireworks-ai-delisted-no-date-keep

Naming or coverage gaps rather than delistings: fireworks_ai/glm-5p3-fast and fireworks_ai/kimi-k3-fast are listed by the API under their accounts/fireworks/routers/ form; fireworks_ai/glm-5p3-us, fireworks_ai/glm-5p3-flash-us, fireworks_ai/kimi-k3-us and fireworks_ai/deepseek-v4p1-flash-us are US-only endpoints the model list does not expose but the pricing page still prices; the size-tier keys (fireworks-ai-up-to-4b, fireworks-ai-4.1b-to-16b, fireworks-ai-above-16b, fireworks-ai-moe-up-to-56b, fireworks-ai-56b-to-176b, fireworks-ai-default, fireworks-ai-embedding-up-to-150m, fireworks-ai-embedding-150m-to-350m) are never on the model list by design

Pre-Submission checklist

  • Cost map change only, no code or tests touched

Note

Low Risk
Cost-map metadata only; no runtime logic, auth, or billing code changes.

Overview
Adds Fireworks Inkling priority-tier cost fields (input_cost_per_token_priority, cache_read_input_token_cost_priority, output_cost_per_token_priority) so lookups align with other Fireworks serverless models that expose priority pricing.

The new priority rates match the existing standard input, cache-read, and output values from the Fireworks prices API. The row’s source is updated to https://api.fireworks.ai/v1/serverless/models?format=nested. The same edit is applied in the main cost map and litellm/model_prices_and_context_window_backup.json.

Reviewed by Cursor Bugbot for commit f533413. Bugbot is set up for automated code reviews on this repo. Configure here.

@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

@greptile-apps

greptile-apps Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

[Medium risk] Updates pricing data for a single model.

The PR appears safe to merge.

Summary

Adds Fireworks Inkling priority input, cached-input, and output prices to both cost maps, and updates the row’s pricing source to the Fireworks serverless-models API.

Reviews (1) · Last reviewed commit: "chore(cost-map): add fireworks inkling p..."

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit f533413. Configure here.

@codecov

codecov Bot commented Oct 1, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@codspeed

codspeed Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm-providers/price-sync-fireworks-ai (f533413) with main (ed4caeb)

Open in CodSpeed

@kerry-berri
kerry-berri merged commit 9b8ddb0 into main Oct 1, 2026
94 of 100 checks passed
jan-sauer-reef added a commit to jan-sauer-reef/litellm that referenced this pull request Oct 1, 2026
…ject_key_prefix

* upstream/main: (62 commits)
  fix(guardrails): scan Responses API input in Azure Prompt Shield (BerriAI#43786)
  feat(lens): investigate sampled traces and retain batch results (BerriAI#43942)
  fix(proxy): restore pre-config-wins handling of pass-through endpoints (BerriAI#43962)
  fix(cost-map): raise baseten DeepSeek-V4.1-Flash max output to 262144 (BerriAI#43916)
  chore(cost-map): add deprecation date for anthropic claude-sonnet-4-5 (BerriAI#43898)
  chore(cost-map): add fireworks inkling priority prices from the prices api (BerriAI#43949)
  feat(guardrails): honor litellm_params.timeout in every HTTP guardrail (BerriAI#43134)
  test(e2e): typed per-test metadata for the e2e suite (BerriAI#42044)
  fix(caching): write the response-cache SET to Redis at once instead of on the post-call batch (BerriAI#43973)
  feat(ui): filter tags by name and description on the Tag Management page (BerriAI#42949)
  feat(providers): add Cortecs as an OpenAI-compatible provider (BerriAI#43872)
  feat(e2e): record each e2e test's steps, starting with ProxyClient (BerriAI#42393)
  test(ci): repair stale tests and move retired OpenAI text-completion fixtures (BerriAI#43958)
  feat(proxy): record in spend logs whether a request used a client-forwarded Anthropic OAuth token (BerriAI#43063)
  fix(azure_storage): keep the DataLakeServiceClient alive until its TTL elapses (BerriAI#43082)
  chore(deps): bump gitpython and tornado, extend diskcache osv ignore to Nov 1 (BerriAI#43961)
  fix(guardrails): treat an unknown straiker api_version as unset instead of skipping the guardrail (BerriAI#43956)
  fix(azure_storage): name Data Lake objects without base64 padding or slashes (BerriAI#43914)
  fix(grayswan): send request conversation and tool calls to post-call monitor (BerriAI#43770)
  chore(cost-map): sync openrouter prices from the models API (BerriAI#43950)
  ...
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant