Repository navigation
Conversation
A bulk item that carried only tags reached the DB with max_budget, team_id, and budget_id as explicit nulls, wiping the key's budget and detaching it from its team. The per-key update is now built from the fields the item actually set, so a field left out keeps its value and an explicit null still clears it, the same as /key/update. Items carrying a field the bulk path cannot apply (object_permission and the like) are rejected with 422 instead of being silently dropped.
…dget_alert_wording fix(alerting): clarify budget threshold messages
…when a spend row is placeholdered With store_prompts_in_spend_logs on, the persisted request body kept the client's model string even when the row's model, model_group, and error text had been replaced by the unknown-model placeholder. The body's model now takes the same placeholder on those rows. Also annotates the new test locals with Final and wraps the four test lines that ran past 120 characters.
…parametrized hang test covers both invocations
The cost tracking callback f-stringed chosen_metadata, litellm_metadata, and old_metadata into the failed_tracking_spend alert on every failure, at every log level, so one 250-byte request produced a 23 KB alert carrying the client's metadata, headers, and key-auth reprs four times over. The alert now carries the exception, the traceback, the model, and the call type; the metadata keys are logged once at debug level through lazy formatting, so nothing is built at warning level
…6887 fix(cost): carry image and video input tokens through the Responses usage bridge (internal copy of BerriAI#36887)
… drop a test helper docstring
…test_timeout ci(unit): fail a hung test in 120s with a traceback instead of idling the shard to its step timeout
…its id A registered S3 Vectors store usually carries only its "bucket:index" id, and the previous commit stopped forwarding the caller's bucket and index for a managed store, so ingesting into one raised KeyError 'vector_bucket_name'. The ingestion now derives both from vector_store_id with the rule the search side already uses, explicit keys still winning. The caller's litellm_credential_name is dropped for a managed store too, since it expands into api_key and api_base, and max_embedding_requests_per_min joins the per-upload options a caller may still set.
refactor(types): replace Any with proven types in 6 files
fix(proxy): register transcribe as a known provider for model grants
…ches feat(batches): support Mistral files/batches and per-page OCR batch cost tracking (internal copy of BerriAI#40484)
A "bucket:" or ":index" id split into an empty name, so ingestion silently generated a fresh index and search sent the empty name to AWS. Both sides now raise the existing format error through the shared helper.
…ons bridge Hosted Responses API tools with no Chat Completions equivalent were forwarded verbatim, so Codex 0.140+ got a 400 from the provider on every turn. The bridge now drops tool_search and local_shell the same way it drops computer_use, image_generation, and shell, and also drops parallel_tool_calls when no chat tools remain, since chat completions only accepts it alongside tools
…s stored request body
…s without self-recursion
…oice_400 fix(utils): reject an untranslatable tool_choice with a 400 instead of a 500
…_vectors The ingest-side bucket and index precedence now sits next to the shared store id split instead of under litellm/rag/, where provider-specific parsing does not belong.
… orphaned uploads
…ounded_error_msg fix(proxy): keep request metadata out of the cost tracking failure alert
…th_fail_closed fix(masker): memoize shared nodes and fail closed past the depth cap
…nner # Conflicts: # tests/test_litellm/proxy/batches_endpoints/test_endpoints.py
…dding model The S3 Vectors ingestion embedded every chunk with the request's embedding.model or the default, never the embedding_model the store was registered with, while search on the same store embeds with the registered model. A registered store uploaded to by id alone therefore embedded with the wrong model and AWS rejected the vectors on the dimension mismatch. The store's embedding model now wins for S3 Vectors ingestion through a helper next to the one search already uses
…s outside peak hours DeepSeek charges half the listed rate outside 01:00-04:00 and 06:00-10:00 UTC Monday to Friday, so every deepseek-flash, deepseek-v4-flash, deepseek-v4-flash-vision-exp, and deepseek-v4-pro entry now carries an off_peak_pricing block with those windows and the halved input, output, and cache-hit rates. The generated cost map schema picks up the block, and the regression tests pin the peak and off-peak cost of one call at fixed moments.
…l_search fix(responses): drop tool_search and local_shell in the chat completions bridge
…log golden The spend-log payload now carries metadata.autorouter_savings_estimate and metadata.autorouter_baseline_observation, but the golden this suite compares against does not list them, so the exact-match check reports them as extra keys and logging_testing goes red on the release candidate. This backports the golden verbatim from main, where the same change landed in 09e14a4 after rc/1.103.0 was cut. tests/logging_callback_tests/test_gcs_pub_sub.py::test_async_gcs_pub_sub_v1: 1 failed before, 1 passed after.
…_test_helpers test(mcp): migrate the mcp test helpers to the mcp 2.x MCPServer API
…autorouter_golden test(logging): add autorouter estimate keys to the GCS pub/sub spend-log golden
…are provisioned The test resolves four AWS_GOVCLOUD_* values through os.environ/ references on the deployment, and none of them is injected into the e2e stack, so the upload returns 500 "S3 bucket_name is required" and the test fails on configuration rather than on the GovCloud partition behavior it guards. The runner podSpec mounts 29 individual secretKeyRef entries from litellm-provider-keys with no envFrom, and the string "govcloud" appears nowhere in project-releaser, so nothing supplies these values today. The commercial-partition Bedrock batch tests resolve their own os.environ/ bucket and pass, which rules out the reference syntax and leaves the absent secrets as the only cause. Skipped rather than weakened, matching the LIT-5027 and LIT-4820 markers in this file: the assertions stay as the correct contract and the marker names the condition for removing it. Reports 1 skipped instead of 1 failed.
…vcloud_batch_e2e test(batches): skip the Bedrock GovCloud batch e2e until its secrets are provisioned
…ding them in a 1024-byte chunker (BerriAI#42607) * fix(bedrock): stream /v1/messages Invoke bytes through instead of holding them in a 1024-byte chunker Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * style(bedrock): apply ruff format to invoke messages stream passthrough Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * style(bedrock): drop drive-by reformat of existing invoke messages tests Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(bedrock): collect streamed chunks into a tuple in passthrough regression test Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(bedrock): give the passthrough regression test a 10s first-chunk budget * test(bedrock): type the eventstream frame helper's payload as Mapping[str, object] * test(bedrock): take the gated byte stream's chunks as an immutable Sequence --------- Co-authored-by: mateo <mateo@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com> (cherry picked from commit 975bd28)
…rc_1_103_0 fix(bedrock): backport the /v1/messages Invoke streaming pass-through to rc/1.103.0 (BerriAI#42607)
…pend rollup code from rc/1.103.0 (BerriAI#43326) Reverts the code from BerriAI#41324 (merge 87694c2) and BerriAI#41293 (merge 2e46b10). The LiteLLM_DailyGlobalSpend model and migration stay so databases that applied rc.1 keep a consistent table and rollup marker when the feature returns.
… poison the pool connection (BerriAI#43029) (BerriAI#43323) Prisma types a raw array parameter from the first batch a connection sees. After a flush in which every member cost was a whole number (a free model), the connection's cached statement expected int8[] and every later fractional batch on it failed with "improper binary format in array element", so member spend silently stopped landing while team spend kept rising. The rows now travel as one JSON document unpacked by jsonb_to_recordset with the column types declared in SQL, so Postgres types the numbers and the batch shape no longer matters. (cherry picked from commit 5e4b1b9) Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
* fix: apply configured cache_control_injection_points beside client cache_control marks (BerriAI#41956) (cherry picked from commit d8267d5) * fix(token_counter): count replayed redacted_thinking blocks so prompt_caching keeps pinning (BerriAI#42069) (cherry picked from commit 0e7cf51) * fix(mcp): return camelCase tool keys from /v1/mcp/tools after the SDK 2 upgrade (BerriAI#42352) (cherry picked from commit 5eb4e30) * fix(rust_bridge): keep the Messages route on Python until the Rust path is ready (BerriAI#42517) rc/1.103.0 has no token counter or tokenizer routes in the catalog, so only the Messages rule moves to PYTHON_ONLY (cherry picked from commit 85ed18e) * fix(proxy): stop leaking periodic tasks on every DB config reload (BerriAI#42784) (cherry picked from commit fc87a06) --------- Co-authored-by: Mateo Wang <277851410+mateo-berri@users.noreply.github.com> Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…election to rc/1.103.0 (BerriAI#43343) * feat(otel): emit gen_ai.conversation.id from the caller's session id on v2 LLM spans (BerriAI#42486) * feat(otel): emit gen_ai.conversation.id from the caller's session id on v2 LLM spans Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(otel): keep the caller's header session under missing_session_id: generate and read replayed payload session ids Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(otel): drop only the proxy-minted session id so a caller id on the other metadata key survives Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(otel): keep a replayed session id hidden when it only echoes the payload trace id Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(otel): keep a replayed session id even when the payload trace id fell back to it Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(otel): stop reading the replayed payload's session id, the generated marker does not survive replay Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): audit gen_ai.conversation.id on otel v2 spans through a real proxy, sink and postgres Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): keep otel conversation rigs alive for the whole session so shuffled shards do not reboot the proxy per test Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(otel): stop the audit rig proxies from probing sibling test peers for model info Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(otel): record accepted OTLP batches in the sink instead of mutating the collector Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(otel): guard the accepted batch deque so snapshots cannot race sink appends Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: mrinal <mrinal@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: yucheng <yucheng@berri.ai> (cherry picked from commit a319690) * fix(jwt): accept a team alias in x-litellm-team-id (BerriAI#42445) * fix(jwt): accept a team alias in x-litellm-team-id The header only matched canonical team ids, so a JWT caller selecting one of their teams by its alias got a 403 even though they belonged to it. The header value is now resolved through the existing alias lookup before the JWT allowed-team check and the DB membership fallback, while a value that is already a team id never costs an alias lookup and denials keep naming the value the caller sent Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(jwt): only alias a header team id the database provably lacks Under fallback_to_db_teams a header value whose team row read fails for any reason other than TeamNotFoundError now keeps the membership denial instead of falling through to the alias lookup, so a degraded read cannot select a different team that carries the value as an alias. Drops the HeaderTeam docstring that only restated its fields Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: ryan <ryan@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> (cherry picked from commit 071cb49) * fix(jwt): say x-litellm-team-id matched no team id or alias in the 403 (BerriAI#42495) * fix(jwt): say x-litellm-team-id matched no team id or alias in the 403 Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(jwt): tell the caller when x-litellm-team-id names an alias shared by several teams Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(jwt): deny a shared x-litellm-team-id alias exactly like an unknown value A distinct 403 for an alias several teams share was raised before the allowed-teams check, so any JWT could probe which aliases exist. The alias lookup now treats the duplicate as a miss, and both denials say the value does not resolve to a team id or a unique team alias, which is true for unknown, unauthorized and duplicate values alike Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: ryan <ryan@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> (cherry picked from commit 08639fc) * fix(jwt): let x-litellm-team-id select DB membership teams when the token also carries a team claim (BerriAI#43206) * fix(jwt): let x-litellm-team-id select DB membership teams when the token also carries a team claim Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * docs(jwt): describe header team selection under fallback_to_db_teams Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: yassin <yassin@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> (cherry picked from commit 7b4fd47) --------- Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: mrinal <mrinal@berri.ai> Co-authored-by: yucheng <yucheng@berri.ai> Co-authored-by: ryan <ryan@berri.ai> Co-authored-by: yassin <yassin@berri.ai>
… direct client mark (BerriAI#43342) * fix(caching): stand default cache points down when extra_body hides a direct client mark On native /v1/messages the extra_body envelope is dropped, so a client tool mark or root cache_control reaches Anthropic even when extra_body overrides it. The stand-down check only counted the envelope-merged view and injected two default marks on top of the client's. (cherry picked from commit 01ef7c4) * fix(caching): keep chat completions on the envelope-merged mark count for the default stand-down Chat completions merge extra_body over the request, so a direct tool mark that extra_body replaces never reaches the provider there. Only /v1/messages, where the native transforms drop the envelope, needs to count marks on both sides. (cherry picked from commit a9c5fa7)
…s again (BerriAI#43382) Adds the OwnedProxy harness the BerriAI#42486 backport's OTel test imports, registers the BerriAI#43029 backport's spend flush test in contracts.json with its covers marker, and aligns the MCP tool failure assertion with main.
…ers (BerriAI#43388) Backports the batch cleanup half of BerriAI#43321 so an Azure cancel still in progress after 120s leaves a warning instead of a teardown error, and points the MCP Tools UI test at DeepWiki's renamed ask_wiki_question tool.
…erriAI#43382 and BerriAI#43388 (BerriAI#43391) The batch cleanup leftover warning used a class defined in a test-directory module. The xdist controller cannot import it, so an uncaught leftover warning crashed the whole e2e run. It is now a plain UserWarning. The integration harness backport missed the workers option and Gateway.request headers that the OTel conversation id tests use.
…lake tolerance to rc/1.103.0 (BerriAI#43400) * test(e2e): tolerate provider-side flakes on five full-suite cells (BerriAI#42628) * test(e2e): tolerate provider-side flakes on five full-suite cells Mistral OCR retries a provider-relayed 429 with backoff, the Vertex vision probe turns reasoning off so its 32 tokens go to the answer, the Vertex cache cell spaces eight never-seen prefixes 15s apart around Google's nondeterministic minimum-token rejection and prices the cached tokens instead of prompt_tokens, and the Azure content-policy cell resends the jailbreak prompt while Azure skips its filter * test(e2e): shorten the new helper docstrings * test(e2e): accept a relayed provider 429 on the rust OCR cells The gateway already retries a provider 429 three times per call and the Mistral key is shared across pipelines, so a throttle can hold across all four attempts of the OCR cell. After the bounded retries the cell now accepts the gateway's faithful relay of the provider's 429 (throttling_error, code 429) as its second expected outcome; the gateway's own 429 and any other error still fail the cell at once. * test(e2e): drop the harness unit tests, the live cells cover the helpers --------- Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com> (cherry picked from commit 41ca465) * fix(streaming): keep litellm Usage on text-completion usage chunks (BerriAI#43047) * fix(streaming): keep litellm Usage on text-completion usage chunks * fix(streaming): convert provider usage to litellm Usage instead of dropping it (cherry picked from commit be35b22) --------- Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
…ig (BerriAI#43429) * fix(proxy): unregister logging callbacks removed from the stored config POST /config/callback/delete saved the config and resynced, but the resync only ever added callbacks, so a deleted callback kept exporting and kept showing in /get/config/callbacks as read-only on every worker. ProxyConfig now tracks which callback list entries each DB config sync registered and unregisters them once the stored config stops listing them. Callbacks it did not register (YAML, code) are never touched, and a failed config load skips the sync instead of treating the config as empty. * refactor(proxy): keep callback sync comprehensions to one for clause
metadata.requester_metadata embeds per-request headers/traceparent, so dumping it on gen_ai metrics grows in-process SDK aggregations without bound. Lift the two bounded grouping fields first so exclude_list can drop the blob without losing the dimensions Prometheus already exposes. Co-authored-by: Cursor <cursoragent@cursor.com> (cherry picked from commit 9af833c) (cherry picked from commit 14f97d3)
Re-apply of e2beef8 (Noma fork PR #12) onto the 1.102 base: upstream refactored bedrock_converse_supports_strict_tools with Final annotations, so the original patch did not cherry-pick cleanly. Same intent — pin the no-strict payload by returning False, leaving _get_bedrock_converse_strict_tools_flag wired for a one-line revert once we ship a release with BerriAI#35688. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 1000cb8)
Automates the previously-manual `docker buildx --push` of the -custom image. Triggers on push to v*-custom branches + workflow_dispatch; builds docker/Dockerfile.non_root multi-arch and pushes to ECR tagged with the branch name (e.g. v1.102.0-custom) + a -<sha> variant. Build-only; deploy stays a values bump + canary (no ArgoCD sync). Prereqs for DevOps before first run: fork access to arc-runners-prod; OIDC trust on role/github-actions-region-deploy for repo Noma-Security/litellm; ecr:PutImage on the litellm repo for that role (or a dedicated build role). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 90c1873)
The arc-runners-prod scale set isn't scoped to this fork, so the job queued forever with no runner. Switch to ubuntu-latest (no ARC scoping) and add setup-qemu for the linux/arm64 leg. Still needs the OIDC trust (repo:Noma-Security/litellm:*) + ecr:PutImage on role/github-actions-region-deploy for the push step. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit d77daae)
The spend counter DualCache used the generic InMemoryCache default of 200 entries, so deployments with more than 200 active budget scopes evicted still-valid counters and reseeded them from lagging DB spend on the next request. A key the proxy had just rejected for being over budget could be admitted again until the batch writer flushed. Fixes BerriAI#40221 Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> (cherry picked from commit 1de5a82) (cherry picked from commit 3ce17a5)
Bedrock serves its openai.* models (gpt-oss, gpt-5.6-sol/terra/luna) over the plain Invoke route, and they speak OpenAI Chat Completions in both directions. `transform_request` has delegated those to AmazonBedrockOpenAIConfig since they landed, but the response side never grew the matching branch, so a valid body came back through Titan's `results[0].outputText` and blew up as `BedrockException - ... 'NoneType' object is not subscriptable`. Streaming had the mirrored gap: no shape the base chunk parser sniffs for matches an OpenAI chunk, so every delta decoded to empty text -- a finished stream with no tokens. The trigger was deploy-new-model validating bedrock/us.openai.gpt-5.6-luna: Bedrock answered correctly and we reported its answer as a Bedrock error. Rather than special-case openai, mirror the request dispatch. `transform_request` enumerates providers explicitly and raises 404 on anything unknown; the response side used a catch-all `else` that assumed Titan, so every provider it did not name inherited Titan's parse. Making it explicit shows the fallthrough was already mis-parsing moonshot, qwen2, qwen3 and stability the same way (each has its own config today, so nothing reached it -- but the trap was armed for the next provider added to the request side). Note: BedrockError now propagates instead of being re-wrapped as a blanket 422. That is required for the unknown-provider 404 to survive, and it stops delegated configs' status codes (429, 400) from being flattened into 422. Upstream BerriAI#37132 reports all of this; it was closed with no fix and main is still affected, so this stays a fork patch. Co-authored-by: Cursor <cursoragent@cursor.com> (cherry picked from commit ad3790c)
…forwarded The Noma fork pins bedrock_converse_supports_strict_tools to False, so toolSpec.strict is never sent to Converse. The upstream tests asserting strict is kept for Sonnet 4.5/4.6 and Opus 4.5/4.6 contradict that pin. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every change on main is either upstream in v1.103.0 (Redis breaker fixes BerriAI#40764/BerriAI#40624/BerriAI#40817) or re-applied on this branch (otel lift, bedrock strict pin, spend_counter_cache, custom-image CI), so the 1.103.0 tree is kept as is. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
There was a problem hiding this comment.
CodeQL found more than 20 potential problems in the proposed changes. Check the Files changed tab for more details.
EvictedClientCloser held queued clients through weakrefs. An SDK client is a
reference cycle that a generational collection reclaims within minutes, well
before the 900s grace window, so the AsyncAzureOpenAI client evicted from the
LLM client cache every hour was never closed: its aiohttp session was
finalized by the garbage collector instead ("Unclosed client session").
In prod every liveness-probe kill of the proxy since the v1.102 rollout
started at one of those finalizations (27 of 28 kills; ~2 expected by chance).
- Hold queued clients strongly until they are closed on their own loop.
- Drop entries whose event loop has closed; nothing can close them anymore.
- A transport on a shared aiohttp session reports leases for every client on
that session, so don't treat them as this client's in-flight requests.
Otherwise clients on the proxy's shared session look busy forever, are never
closed and accumulate once they are held strongly.
(cherry picked from commit 143e720)
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, have a team admin enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit 6f1b150. Configure here.
| def _drop_buckets_of_closed_loops_locked(self) -> None: | ||
| """A loop's bucket is never drained once that loop is closed, so release it.""" | ||
| for key in [k for k, b in self._buckets.items() if b and _loop_is_gone(b[0])]: | ||
| self._pending_count -= len(self._buckets.pop(key)) |
There was a problem hiding this comment.
Closed loops drop closable clients
Medium Severity
loop_ref is stored for every eviction, including sync clients that do not need a loop. _drop_buckets_of_closed_loops_locked then discards an entire bucket when the first entry's loop is gone, so _CLOSABLE_ANYWHERE clients are dropped without close(). Sync handlers queued from a worker loop that later ends keep their pools until GC, which is the leak this change is meant to stop.
Additional Locations (2)
Reviewed by Cursor Bugbot for commit 6f1b150. Configure here.


Why
Move the Noma custom line from
v1.102.0-customto upstreamv1.103.0.v1.102.0-custom, which prod, dev and lab run, is upstreammainatv1.102.0-dev.2, not the 1.102.0 release. This jump covers about 3,300 upstream commits.What
Branch = upstream tag
v1.103.0(==stable/1.103.x,c991f4b01f) + these commits:Fork patches carried from
main/v1.102.0-custom:fix(otel)keepcomponent/platformas metric attributes beforerequester_metadatais excluded (clean pick of 14f97d3)fix(bedrock)never sendtoolSpec.strictto Converse (clean pick of 1000cb8). Still needed: the 1.103.0 cost map still sendsstrictfor Haiku 4.5, Opus 4.6 and Sonnet 4.6, which we serve in prod.fix(proxy)givespend_counter_cacheits own 10000-entry in-memory cache (pick of 3ce17a5, import-line conflict resolved). Still needed: upstream closed fix(proxy): give spend_counter_cache its own 10000-entry in-memory cache BerriAI/litellm#42414 without merging it, and [Bug]: Spend counter cache evicts active budget counters after 200 entries BerriAI/litellm#40221 is still open.ciNoma custom-image build workflow (ci: add workflow to build the Noma litellm custom image #18, ci: build custom image on github-hosted runner (unblock queued build) #19)Re-added from the closed PR #15:
fix(bedrock)parse OpenAI-format Invoke responses instead of Titan's (pick of ad3790c). 1.103.0 still falls through to Titan'sresults[0]foropenai.*on the Invoke route. That happens for newopenai.*models not yet in the cost map, and for explicitbedrock/invoke/...models.Dropped because 1.103.0 already has them: the Redis breaker fixes BerriAI#40764, BerriAI#40624 and BerriAI#40817 (their upstream originals are ancestors of
v1.103.0).Tests: removed the upstream tests that expect
toolSpec.strictto be kept for Sonnet 4.5/4.6 and Opus 4.5/4.6. They contradict the strict pin.Merge of
main: a-s oursmerge, so this PR merges cleanly intomain. Every change onmainis either upstream in 1.103.0 or re-applied above, and the tree is byte-identical to the pre-merge branch.Testing
Full unit suite (
tests/test_litellm, 62.2k tests): this branch vs pristinev1.103.0on the same machine and depstest_otel_metrics_cardinality.py,test_bedrock_openai_invoke_response.py,test_spend_counters.py, and the spend_counter test intest_proxy_server.py).Live proxy smoke test, each patch checked on this branch vs pristine
v1.103.0The proxy config mirrors prod: prometheus + otel callbacks, the otel
exclude_listwithmetadata.requester_metadata, Noma v1 guardrails pre_call + post_call withanonymize_input, and Bedrock Converse + Invoke against a local fake upstream./health/liveliness,/health/readinessRequest blocked by Noma guardrailstrict: truetooltoolSpec.strictandadditionalPropertiesabsentstrict: truesentopenai.gpt-5.6-luna(forcedinvoke/route)'NoneType' object is not subscriptable(the #15 bug)gen_ai.*metric points withcomponent/platformwhilerequester_metadatais excludedmetadata_component/metadata_platformlabelsImage + DB upgrade test (built
docker/Dockerfile.non_rootlocally for linux/arm64 and linux/amd64)Postgres 16 was seeded with the
v1.102.0-dev.2migrations (the current prod schema) plus 500kLiteLLM_SpendLogsrows (575 MB). Each image then ran withDATABASE_URL,store_model_in_db,USE_V2_MIGRATION_RESOLVER=trueand prod's OTel env:litellm_call_idcolumn, both SpendLogs indexes,LiteLLM_DailyGlobalSpendandLiteLLM_ManagedFileContentTableare present./key/generate→ virtual key → mock, Bedrock Converse (luna, Sonnet 4.6 with astricttool) all 200. Noma block 400.toolSpec.strictnot sent.litellm_call_id, key spend flushed (amd64 run). OTelgen_ai.*points carrycomponent, and the Prometheus request series is present. The only ERROR log line is the expected Noma block.Not tested locally: real Redis/ElastiCache paths (no server).
Upstream CI workflows in this fork failed on #17 for environment reasons (secrets, runners), so the local baseline comparison above is the real signal.
Rollout notes (upstream changes to handle before prod)
20260823000000builds an index onLiteLLM_SpendLogs(api_key, startTime)withoutCONCURRENTLY. That blocks inserts while it runs, and prod stores prompts in spend logs. Before the rollout, runCREATE INDEX CONCURRENTLY IF NOT EXISTS "LiteLLM_SpendLogs_api_key_startTime_idx" ON "LiteLLM_SpendLogs"("api_key","startTime");so the migration does nothing. The other 13 migrations only add tables and columns.store_model_in_db: true, any key set in YAML ignores the DB value, and UI or API writes to those keys get a 400. CheckLiteLLM_Configon prod and lab first.max_budget(fix!: re-check budget on router fallback targets BerriAI/litellm#41379). Opt out withenforce_fallback_budget: false.rpmget a matching per-pod in-flight cap (feat(router): reject with 429 when a deployment's max_parallel_requests slots are all in use BerriAI/litellm#41555). A full cap returns 429 immediately.llm_exceptionsalert also fires on proxy 5xx (fix(alerting): send llm_exceptions Slack alert for 5xx HTTPException and ProxyException BerriAI/litellm#41125).session_id(fix(proxy): default litellm_trace_id to the OTel server span trace id BerriAI/litellm#41386).mcpgoes from 1.x to 2.2.0. Dockerfile and entrypoint are unchanged.Not included, still in progress elsewhere:
gateway_name(#20 / upstream BerriAI#43678),on_flagged_action, the Noma scan retry on 5xx, and the Anthropic safeguards backport (upstream BerriAI#43662).Rollout: this PR's push builds
v1.103.0-customin ECR. Bump argo-cd values and canary lab → dev → prod.🤖 Generated with Claude Code
Note
Medium Risk
Adds blocking DB index migration and shifts critical integration coverage to new CircleCI runners with network isolation; removes main-branch guard workflow and changes issue/PR automation behavior.
Overview
CircleCI gains a dedicated integration workflow that matrix-runs contract suites (management, accounting, database, providers, extensions, SDK, cost, browser) via new
run_integration.sh, with iptables egress isolation, owned Postgres/Redis, and helper scripts for readiness, browser verification, and process cleanup. A separateprovider_replay_harnessjob runs capture/replay e2e tests with path-based skipping (provider-harness). Path classification expands for MCP dependencies and cost-map-only changes; integration test deps use a new uv cache key.GitHub automation replaces the old Greptile/Agent Shin PR closer and simple duplicate checker with Codex + LiteLLM workflows for issue classification, duplicate detection, label sync from
issue-labels.json, and “issue fixed” comments. Bug/feature templates use domain-based dropdowns. Agent Shin scripts, triage requirements, daily branch workflow,guard-main-branch, andcheck_duplicate_issuesare removed.E2E stack adds Keycloak for JWT/OIDC management tests; JWT e2e files join the changed-test canary set.
assert_ci_coverageenforces that canonical integration/browser contracts run only on CircleCI.Other notable changes: Friendli model pricing in the cost-map auto-update script; MCP 2.2 in codspeed and a new multi-Python MCP dependency resolution workflow; unit tests get pytest-timeout and optional legacy MCP peer venv; Redis chaos e2e workflow; Prisma migrations for SpendLogs
(api_key, startTime)index and policy attachment priority; Grype config for Wolfi zlib CVE; OCR job drops rust bridge test from glob; docker e2e seeds routing strategy before tests.Reviewed by Cursor Bugbot for commit 6f1b150. Bugbot is set up for automated code reviews on this repo. Configure here.