Skip to content

feat(pricing): add gemini-live-2.5-flash-native-audio to model cost map - #21

Closed
mubashir1osmani wants to merge 583 commits into
litellm_internal_stagingfrom
litellm_gemini_live_native_audio_pricing
Closed

mubashir1osmani wants to merge 583 commits into
litellm_internal_stagingfrom
litellm_gemini_live_native_audio_pricing

Conversation

@mubashir1osmani

Copy link
Copy Markdown
Owner

TLDR

Problem this solves:

  • vertex_ai/gemini-live-2.5-flash-native-audio logged zero usage and $0 spend on every request
  • The model was absent from the pricing JSON, which caused a cascade: TEXT modality was never coerced to AUDIO in the Vertex setup message, so the session hung and never produced a response.done

How it solves it:

  • Adds gemini-live-2.5-flash-native-audio (vertex_ai-language-models provider) and gemini/gemini-live-2.5-flash-native-audio (gemini provider) with GA pricing and gemini_native_audio: true
  • Existing _model_cost_entry fallback (strip vertex_ai/ prefix, look up bare key) now resolves the model correctly

User Flow

Before: a proxy operator using vertex_ai/gemini-live-2.5-flash-native-audio sees no usage logged

  1. Client connects to wss://proxy/v1/realtime?model=vertex_ai/gemini-live-2.5-flash-native-audio
  2. Client sends session.update with modalities: ["text"]
  3. Proxy forwards responseModalities: ["TEXT"] to Vertex AI (native-audio model only supports AUDIO), session never completes a turn
  4. response.done never arrives, usage field is null, spend logged as $0

After: the same model produces a completed response with real usage

  1. Client connects to wss://proxy/v1/realtime?model=vertex_ai/gemini-live-2.5-flash-native-audio
  2. Client sends session.update with modalities: ["text"]
  3. Proxy detects the model as native-audio-only (via the new pricing entry) and coerces to responseModalities: ["AUDIO"]
  4. response.done arrives with non-zero audio token counts; spend is logged at the published rate ($3.00/M audio in, $12.00/M audio out)

Relevant issues

Linear ticket

Pre-Submission checklist

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Screenshots / Proof of Fix

Proof requires a live proxy with this model deployed; connect below once it's available on the gateway.

Type

🐛 Bug Fix

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

mateo-berri and others added 30 commits August 26, 2026 10:56
…idden_alias_explicit_lookup

fix(router): resolve hidden aliases for explicit lookup
…ublic_name_focus

fix(ui): keep focus in the add model public name input while typing
The badges flagged UI that shipped a while ago, so they no longer tell
anyone anything. Dropped all four render sites: the Settings and Admin
Settings items in the left nav, the UI Settings tab in the admin panel,
and the Submitted MCPs tab.

The NewBadge component stays so the next genuinely new surface can use
it again. BetaBadge and the "hide new badges" account toggle are
untouched, since that toggle still gates BetaBadge.
…_cache_pricing

fix(model_prices): price 1-hour cache writes on claude-3-haiku and claude-3-opus at 2x input
…ustom_auth_google_token

fix(proxy): keep the caller's Google token on credential-less Vertex passthrough under custom auth
…cp_bridge_provider_token_lifetime

fix(mcp): preserve provider access token lifetime
…ew-badges-f08c0f

chore(ui): remove stale "New" badges from the dashboard
…ricing_edges

test(cost-estimate): pin the prices and period totals /cost/estimate returns
…gy-audit-c39e33

fix(ci): let the mutation workflow find covered lines so it generates mutants
…tle_gpt55_gpt54_context_window

fix(model_prices): raise bedrock_mantle gpt-5.5 and gpt-5.4 max_input_tokens to Mantle's enforced 1050000
…entra_id_auth

fix(azure/realtime): authenticate realtime websocket with Azure AD token when no api-key
Creating a search tool through the UI only wrote the row; the router was updated
solely by the add_deployment job, so the tool was unusable for up to
PROXY_CONFIG_RELOAD_INTERVAL_SECONDS (30s by default) even on the worker that
served the write. Tools declared in config.yaml load straight into the router at
startup, which is why they never showed the delay.

The create, update and delete endpoints now refresh the router inline, matching
what the MCP server endpoints already do. The refresh is best-effort: the row is
already committed, so a failure must not surface as a 500 and push the caller
into a retry that creates duplicates.

Two related gaps go with it. _init_search_tools_in_db skipped the router update
whenever the merged list came back empty, so deleting the last search tool left
it live in memory forever. And in store_model_in_db-off deployments the
add_deployment job is never scheduled, so DB-backed search tools never reached
the router at all; that branch now loads them at startup and keeps them fresh on
its own interval, the same way MCP servers already do.
…gpt5_reasoning_effort

fix(bedrock): map reasoning_effort to reasoning.effort for OpenAI GPT-5.x on Converse
…#38380)

* test(prometheus): cover caller-identity config failure cases

* test(prometheus): narrow pytest.raises with match to satisfy PT011
…_credential_provider

fix(redis): support credential providers across clients
mateo-berri and others added 28 commits August 27, 2026 10:01
…ype_unbound_local

fix(exception_mapping_utils): map unmapped exceptions when model and provider are unset
…live

Adds a Gemini audio transcription config that maps /v1/audio/transcriptions
onto the Interactions API (speaker attribution and word timestamps land on
the OpenAI verbose_json shape), registers both models with published pricing,
routes text-only Live sessions to TEXT responseModalities so
gemini-3.5-transcribe-live sessions survive, and makes the token-priced
transcription cost path provider-aware instead of hardcoding OpenAI.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…laim (BerriAI#36728)

* fix(ui_sso): resolve highest privilege Entra app role, not first in claim

A user assigned more than one Entra app role — commonly by belonging to
several assigned groups — arrives at the Microsoft SSO callback with every
role in the id_token `roles` claim. LiteLLM stores a single role per user,
and get_microsoft_callback_response collapsed the list by taking the first
value that resolved to a LitellmUserRoles and breaking.

Entra does not guarantee the ordering of the `roles` claim, so which role
won was effectively arbitrary: a user in one group mapped to internal_user
and another mapped to proxy_admin_viewer could be silently demoted to
internal_user, and proxy_admin could lose to either.

The generic/Okta path already resolves this correctly via
determine_role_from_groups, which walks a documented privilege hierarchy.
Hoist that hierarchy into LITELLM_USER_ROLE_HIERARCHY and reuse it, so
app-role logins and group-mapping logins agree.

Extract the selection into MicrosoftSSOHandler.get_user_role_from_app_roles
so it is directly testable — the existing tests re-implemented the loop
inline, which is why the ordering bug was not caught.

Behaviour is unchanged for single-role claims, unrecognised values, and
empty claims. Roles the hierarchy does not rank (org_admin, team, customer)
are resolved deterministically rather than by claim order.

* refactor(ui_sso): trim role selection prose and use immutable annotations

Addresses review feedback on the app role selection helper.

Drop the explanatory comments and the Args/Returns docstring boilerplate that
restated the control flow, keeping only the part a reader cannot infer from the
code: that Entra does not guarantee claim ordering, and how unranked roles
resolve.

Type the parameter as Sequence[str] rather than list[str] and build the resolved
set as a frozenset, so the helper stops adding an LIT001 mutable-collection
annotation. Make LITELLM_USER_ROLE_HIERARCHY a tuple for the same reason.

No behaviour change: the ordering regression tests still fail against the
previous first-match-wins logic and pass here.
…_openapi_snapshot

fix(proxy): regenerate lazy OpenAPI snapshot and guard it in CI
* feat(ui): add cache hit/miss filter to Request Logs

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: guard cache_hit_filter validation for direct handler calls

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(ui): drop redundant cache filter comment

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…sigv4_bearer_fix

fix(bedrock): sign rerank requests with the shared header-filtered SigV4 helper (internal copy of BerriAI#36462)
Gemini Live sends no usageMetadata and no turnComplete for
gemini-3.5-transcribe-live sessions, so realtime spend logged as 0.0.
Attach estimated usage to the input_audio_transcription.completed event
using Google's published billing estimate (25 audio tokens/sec of input,
175 text tokens/min of output) derived from the streamed pcm16 audio
duration, gated to audio_transcription-mode models so conversational
Live models keep billing through usageMetadata. Also capture that usage
in the provider_config backend path so realtime cost calculation sees it.
…nds on page one (BerriAI#38545)

/v2/model/info returns llm_router.model_list, which carries no defined order: the DB
read has no order_by and an edited deployment is popped and re-appended. The Auto
routers table rendered that order verbatim behind a ten-row first page, so on a proxy
with more than ten auto routers a router created moments ago was drawn wherever the
API happened to return it, in practice last, and read as never created

Adopt the ordering the rest of the dashboard already uses, with the two cases this
table has and its siblings do not. created_at is enterprise-gated and config.yaml
routers never carry one, so seeding created_at desc alone leaves every comparison
tied on a non-premium proxy and the fix a no-op. The column now declares
sortUndefined last, which table-core applies before the desc flip so undated rows
stay last in both directions, and the row emits undefined rather than null so that
branch is reachable at all. Name is the secondary key, giving the undated block a
defined order too

Page size is deliberately unchanged: it exposes the missing order rather than
causing it
…at high thinking (BerriAI#38490)

The preset put Fable 5 in REASONING, sitting above Opus in a Haiku to Sonnet to
Opus ladder even though Fable is the lighter model. Run Opus 5 there instead, at
high thinking, so the tier above COMPLEX is the same model thinking harder rather
than a different and lighter one.

This is the first bundled preset to carry tier_model_configs. The round trip was
already built and unit tested, but nothing between the bundled JSON and the
create payload asserted on it, so add that coverage here.
…tions (BerriAI#38532)

* feat(otel): support per-team/per-key service.name for OTel v2 destinations

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(otel): pin key-level otel_service_name_override surviving team metadata merge

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): key-level otel_service_name outranks team's after metadata merge

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…lback_cache

test(e2e): de-flake the cost-header cache read and the router fallback control
…cribe

feat(gemini): day-0 support for gemini-3.5-transcribe and transcribe-live
…ks and health-check routing (BerriAI#38539)

* feat(health): opt-in model-group allowlist for background health checks and health-check routing

* fix(health): merge shared health states per writer scope instead of replacing

* refactor(health): drop restating comment and parameterize test scope annotations

* chore: remove stray generated prisma migration file

* fix(health): merge health states against the Redis snapshot, not the pod-local copy

* fix(health): fall back to the pod-local snapshot when the Redis read returns nothing
…cts one on tools/call (BerriAI#38555)

* fix(mcp): keep upstream OAuth Authorization when jwt signer hook injects one on tools/call

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): only treat server credential as occupying Authorization when it maps to that header

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…entries without custom pricing (BerriAI#38542)

* fix: suppress misleading register_model unresolved-cost warnings for entries without custom pricing

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: do not warn about zero cache costs for tiered pricing entries

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ndow (BerriAI#38514)

* feat(proxy): opt-in budget rollover carrying overage into the next window

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): zero under-cap rows before decrementing over-cap rows in cascade resets

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Missing entry caused a cascade: _model_cost_entry returned {} for
vertex_ai/gemini-live-2.5-flash-native-audio, so _is_audio_only_live_model
returned False, _coerce_response_modalities never coerced TEXT -> AUDIO in
the Vertex setup message, the session hung, and no usage was ever logged.

Add bare key (vertex_ai-language-models provider) and gemini/ prefixed key
(gemini provider) with GA pricing: $0.50/M text in, $3.00/M audio in,
$2.00/M text out, $12.00/M audio out, gemini_native_audio=true.

Also extend the existing parametrized tests to assert the vertex_ai/ prefix
resolves correctly through _model_cost_entry's stripped-prefix fallback.
… sentinel (BerriAI#38471)

* fix(auth): skip guaranteed-miss team lookup for the litellm-dashboard sentinel

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style: ruff format

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: assert builder result instead of swallowing exceptions; drop redundant comment

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…rants the key already holds (BerriAI#38463)

* fix(key_management): allow /key/update to keep or shrink MCP server grants the key already holds

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(key_management): reuse key row's included object_permission instead of a second lookup

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ex_ai/chirp-3, vertex_ai/veo-3.1-lite-generate-001
Resolved conflicts in both pricing JSON files (kept our /v1/realtime
endpoint addition for vertex_ai entry, plus adopted upstream's
supports_prompt_caching and search_context_cost_per_query additions).

Also adopted upstream's transformation.py refactor (which dropped
_is_native_audio_model), updated tests accordingly: removed the stale
test for the now-absent method, added a catalog regression test that
reads the JSON file directly instead of relying on the mocked fixture.
@mubashir1osmani

Copy link
Copy Markdown
Owner Author

Reopened upstream as BerriAI#38573; the fork's staging base was 578 commits stale and broke the lint gate.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.