Skip to content

feat: move the Noma custom line to upstream v1.103.0 - #21

Open
giladops wants to merge 3357 commits into
mainfrom
v1.103.0-custom
Open

giladops wants to merge 3357 commits into
mainfrom
v1.103.0-custom

Conversation

@giladops

@giladops giladops commented Sep 29, 2026 •

Copy link
Copy Markdown

Why

Move the Noma custom line from v1.102.0-custom to upstream v1.103.0. v1.102.0-custom, which prod, dev and lab run, is upstream main at v1.102.0-dev.2, not the 1.102.0 release. This jump covers about 3,300 upstream commits.

What

Branch = upstream tag v1.103.0 (== stable/1.103.x, c991f4b01f) + these commits:

Fork patches carried from main / v1.102.0-custom:

Re-added from the closed PR #15:

  • fix(bedrock) parse OpenAI-format Invoke responses instead of Titan's (pick of ad3790c). 1.103.0 still falls through to Titan's results[0] for openai.* on the Invoke route. That happens for new openai.* models not yet in the cost map, and for explicit bedrock/invoke/... models.

Dropped because 1.103.0 already has them: the Redis breaker fixes BerriAI#40764, BerriAI#40624 and BerriAI#40817 (their upstream originals are ancestors of v1.103.0).

Tests: removed the upstream tests that expect toolSpec.strict to be kept for Sonnet 4.5/4.6 and Opus 4.5/4.6. They contradict the strict pin.

Merge of main: a -s ours merge, so this PR merges cleanly into main. Every change on main is either upstream in 1.103.0 or re-applied above, and the tree is byte-identical to the pre-merge branch.

Testing

Full unit suite (tests/test_litellm, 62.2k tests): this branch vs pristine v1.103.0 on the same machine and deps

  • Branch: 81 failing. Pristine 1.103.0: 88 failing. All of these are environment failures shared by both (network, local bind, timing).
  • 11 tests failed only on the branch. Rerun in isolation 3 times on both trees they behave the same (1 intermittent failure on both), so there are no regressions.
  • Tests for the carried patches pass (test_otel_metrics_cardinality.py, test_bedrock_openai_invoke_response.py, test_spend_counters.py, and the spend_counter test in test_proxy_server.py).

Live proxy smoke test, each patch checked on this branch vs pristine v1.103.0
The proxy config mirrors prod: prometheus + otel callbacks, the otel exclude_list with metadata.requester_metadata, Noma v1 guardrails pre_call + post_call with anonymize_input, and Bedrock Converse + Invoke against a local fake upstream.

Check branch pristine 1.103.0
/health/liveliness, /health/readiness 200 200
mock chat completion (with pre + post Noma scans) 200 200
Noma scan verdict = block 400 Request blocked by Noma guardrail 400
Bedrock Converse Claude Sonnet 4.6 with a strict: true tool toolSpec.strict and additionalProperties absent strict: true sent
Bedrock Invoke openai.gpt-5.6-luna (forced invoke/ route) 200, correct content 400 'NoneType' object is not subscriptable (the #15 bug)
OTel gen_ai.* metric points with component/platform while requester_metadata is excluded 5 0
Prometheus metadata_component/metadata_platform labels present present

Image + DB upgrade test (built docker/Dockerfile.non_root locally for linux/arm64 and linux/amd64)
Postgres 16 was seeded with the v1.102.0-dev.2 migrations (the current prod schema) plus 500k LiteLLM_SpendLogs rows (575 MB). Each image then ran with DATABASE_URL, store_model_in_db, USE_V2_MIGRATION_RESOLVER=true and prod's OTel env:

  • Migrations: 170 → 184 applied, 0 failed, 0 invalid indexes. Startup was 10s on arm64 and 24s on amd64 (emulated), including the non-concurrent SpendLogs index build over 500k rows. New litellm_call_id column, both SpendLogs indexes, LiteLLM_DailyGlobalSpend and LiteLLM_ManagedFileContentTable are present.
  • /key/generate → virtual key → mock, Bedrock Converse (luna, Sonnet 4.6 with a strict tool) all 200. Noma block 400. toolSpec.strict not sent.
  • Spend logs written with litellm_call_id, key spend flushed (amd64 run). OTel gen_ai.* points carry component, and the Prometheus request series is present. The only ERROR log line is the expected Noma block.

Not tested locally: real Redis/ElastiCache paths (no server).

Upstream CI workflows in this fork failed on #17 for environment reasons (secrets, runners), so the local baseline comparison above is the real signal.

Rollout notes (upstream changes to handle before prod)

  1. Migration 20260823000000 builds an index on LiteLLM_SpendLogs(api_key, startTime) without CONCURRENTLY. That blocks inserts while it runs, and prod stores prompts in spend logs. Before the rollout, run CREATE INDEX CONCURRENTLY IF NOT EXISTS "LiteLLM_SpendLogs_api_key_startTime_idx" ON "LiteLLM_SpendLogs"("api_key","startTime"); so the migration does nothing. The other 13 migrations only add tables and columns.
  2. The config file now overrides the DB (refactor(proxy): make the config file win over the database BerriAI/litellm#41779, breaking). With store_model_in_db: true, any key set in YAML ignores the DB value, and UI or API writes to those keys get a 400. Check LiteLLM_Config on prod and lab first.
  3. Before each fallback, the router re-checks the key's max_budget (fix!: re-check budget on router fallback targets BerriAI/litellm#41379). Opt out with enforce_fallback_budget: false.
  4. Deployments with rpm get a matching per-pod in-flight cap (feat(router): reject with 429 when a deployment's max_parallel_requests slots are all in use BerriAI/litellm#41555). A full cap returns 429 immediately.
  5. Failure accounting changes: a Noma block on a stream now counts as a failure (fix(proxy): log blocked streaming guardrail responses as failures, not success BerriAI/litellm#40191), 401s count as failures (fix(prometheus): count 401 auth failures in litellm_proxy_failed_requests_metric BerriAI/litellm#41170), and the Slack llm_exceptions alert also fires on proxy 5xx (fix(alerting): send llm_exceptions Slack alert for 5xx HTTPException and ProxyException BerriAI/litellm#41125).
  6. The OTel trace id now becomes the spend-log session_id (fix(proxy): default litellm_trace_id to the OTel server span trace id BerriAI/litellm#41386).
  7. mcp goes from 1.x to 2.2.0. Dockerfile and entrypoint are unchanged.

Not included, still in progress elsewhere: gateway_name (#20 / upstream BerriAI#43678), on_flagged_action, the Noma scan retry on 5xx, and the Anthropic safeguards backport (upstream BerriAI#43662).

Rollout: this PR's push builds v1.103.0-custom in ECR. Bump argo-cd values and canary lab → dev → prod.

🤖 Generated with Claude Code


Note

Medium Risk
Adds blocking DB index migration and shifts critical integration coverage to new CircleCI runners with network isolation; removes main-branch guard workflow and changes issue/PR automation behavior.

Overview
CircleCI gains a dedicated integration workflow that matrix-runs contract suites (management, accounting, database, providers, extensions, SDK, cost, browser) via new run_integration.sh, with iptables egress isolation, owned Postgres/Redis, and helper scripts for readiness, browser verification, and process cleanup. A separate provider_replay_harness job runs capture/replay e2e tests with path-based skipping (provider-harness). Path classification expands for MCP dependencies and cost-map-only changes; integration test deps use a new uv cache key.

GitHub automation replaces the old Greptile/Agent Shin PR closer and simple duplicate checker with Codex + LiteLLM workflows for issue classification, duplicate detection, label sync from issue-labels.json, and “issue fixed” comments. Bug/feature templates use domain-based dropdowns. Agent Shin scripts, triage requirements, daily branch workflow, guard-main-branch, and check_duplicate_issues are removed.

E2E stack adds Keycloak for JWT/OIDC management tests; JWT e2e files join the changed-test canary set. assert_ci_coverage enforces that canonical integration/browser contracts run only on CircleCI.

Other notable changes: Friendli model pricing in the cost-map auto-update script; MCP 2.2 in codspeed and a new multi-Python MCP dependency resolution workflow; unit tests get pytest-timeout and optional legacy MCP peer venv; Redis chaos e2e workflow; Prisma migrations for SpendLogs (api_key, startTime) index and policy attachment priority; Grype config for Wolfi zlib CVE; OCR job drops rust bridge test from glob; docker e2e seeds routing strategy before tests.

Reviewed by Cursor Bugbot for commit 6f1b150. Bugbot is set up for automated code reviews on this repo. Configure here.

A bulk item that carried only tags reached the DB with max_budget, team_id,
and budget_id as explicit nulls, wiping the key's budget and detaching it
from its team. The per-key update is now built from the fields the item
actually set, so a field left out keeps its value and an explicit null still
clears it, the same as /key/update. Items carrying a field the bulk path
cannot apply (object_permission and the like) are rejected with 422 instead
of being silently dropped.
…dget_alert_wording

fix(alerting): clarify budget threshold messages
…when a spend row is placeholdered

With store_prompts_in_spend_logs on, the persisted request body kept the client's model string even when the row's model, model_group, and error text had been replaced by the unknown-model placeholder. The body's model now takes the same placeholder on those rows. Also annotates the new test locals with Final and wraps the four test lines that ran past 120 characters.
…parametrized hang test covers both invocations
The cost tracking callback f-stringed chosen_metadata, litellm_metadata,
and old_metadata into the failed_tracking_spend alert on every failure,
at every log level, so one 250-byte request produced a 23 KB alert
carrying the client's metadata, headers, and key-auth reprs four times
over. The alert now carries the exception, the traceback, the model, and
the call type; the metadata keys are logged once at debug level through
lazy formatting, so nothing is built at warning level
…6887

fix(cost): carry image and video input tokens through the Responses usage bridge (internal copy of BerriAI#36887)
…test_timeout

ci(unit): fail a hung test in 120s with a traceback instead of idling the shard to its step timeout
…its id

A registered S3 Vectors store usually carries only its "bucket:index" id,
and the previous commit stopped forwarding the caller's bucket and index for
a managed store, so ingesting into one raised KeyError 'vector_bucket_name'.
The ingestion now derives both from vector_store_id with the rule the search
side already uses, explicit keys still winning. The caller's
litellm_credential_name is dropped for a managed store too, since it expands
into api_key and api_base, and max_embedding_requests_per_min joins the
per-upload options a caller may still set.
refactor(types): replace Any with proven types in 6 files
fix(proxy): register transcribe as a known provider for model grants
…ches

feat(batches): support Mistral files/batches and per-page OCR batch cost tracking (internal copy of BerriAI#40484)
A "bucket:" or ":index" id split into an empty name, so ingestion silently
generated a fresh index and search sent the empty name to AWS. Both sides now
raise the existing format error through the shared helper.
…ons bridge

Hosted Responses API tools with no Chat Completions equivalent were forwarded
verbatim, so Codex 0.140+ got a 400 from the provider on every turn. The bridge
now drops tool_search and local_shell the same way it drops computer_use,
image_generation, and shell, and also drops parallel_tool_calls when no chat
tools remain, since chat completions only accepts it alongside tools
…oice_400

fix(utils): reject an untranslatable tool_choice with a 400 instead of a 500
…_vectors

The ingest-side bucket and index precedence now sits next to the shared
store id split instead of under litellm/rag/, where provider-specific
parsing does not belong.
…ounded_error_msg

fix(proxy): keep request metadata out of the cost tracking failure alert
…th_fail_closed

fix(masker): memoize shared nodes and fail closed past the depth cap
…nner

# Conflicts:
#	tests/test_litellm/proxy/batches_endpoints/test_endpoints.py
…dding model

The S3 Vectors ingestion embedded every chunk with the request's
embedding.model or the default, never the embedding_model the store was
registered with, while search on the same store embeds with the
registered model. A registered store uploaded to by id alone therefore
embedded with the wrong model and AWS rejected the vectors on the
dimension mismatch. The store's embedding model now wins for S3 Vectors
ingestion through a helper next to the one search already uses
…s outside peak hours

DeepSeek charges half the listed rate outside 01:00-04:00 and 06:00-10:00 UTC
Monday to Friday, so every deepseek-flash, deepseek-v4-flash,
deepseek-v4-flash-vision-exp, and deepseek-v4-pro entry now carries an
off_peak_pricing block with those windows and the halved input, output, and
cache-hit rates. The generated cost map schema picks up the block, and the
regression tests pin the peak and off-peak cost of one call at fixed moments.
…l_search

fix(responses): drop tool_search and local_shell in the chat completions bridge
yuneng-berri and others added 26 commits September 19, 2026 19:25
…log golden

The spend-log payload now carries metadata.autorouter_savings_estimate and
metadata.autorouter_baseline_observation, but the golden this suite compares
against does not list them, so the exact-match check reports them as extra
keys and logging_testing goes red on the release candidate.

This backports the golden verbatim from main, where the same change landed in
09e14a4 after rc/1.103.0 was cut.

tests/logging_callback_tests/test_gcs_pub_sub.py::test_async_gcs_pub_sub_v1:
1 failed before, 1 passed after.
…_test_helpers

test(mcp): migrate the mcp test helpers to the mcp 2.x MCPServer API
…autorouter_golden

test(logging): add autorouter estimate keys to the GCS pub/sub spend-log golden
…are provisioned

The test resolves four AWS_GOVCLOUD_* values through os.environ/ references on
the deployment, and none of them is injected into the e2e stack, so the upload
returns 500 "S3 bucket_name is required" and the test fails on configuration
rather than on the GovCloud partition behavior it guards.

The runner podSpec mounts 29 individual secretKeyRef entries from
litellm-provider-keys with no envFrom, and the string "govcloud" appears
nowhere in project-releaser, so nothing supplies these values today. The
commercial-partition Bedrock batch tests resolve their own os.environ/ bucket
and pass, which rules out the reference syntax and leaves the absent secrets as
the only cause.

Skipped rather than weakened, matching the LIT-5027 and LIT-4820 markers in
this file: the assertions stay as the correct contract and the marker names the
condition for removing it.

Reports 1 skipped instead of 1 failed.
…vcloud_batch_e2e

test(batches): skip the Bedrock GovCloud batch e2e until its secrets are provisioned
…ding them in a 1024-byte chunker (BerriAI#42607)

* fix(bedrock): stream /v1/messages Invoke bytes through instead of holding them in a 1024-byte chunker

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(bedrock): apply ruff format to invoke messages stream passthrough

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(bedrock): drop drive-by reformat of existing invoke messages tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(bedrock): collect streamed chunks into a tuple in passthrough regression test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(bedrock): give the passthrough regression test a 10s first-chunk budget

* test(bedrock): type the eventstream frame helper's payload as Mapping[str, object]

* test(bedrock): take the gated byte stream's chunks as an immutable Sequence

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
(cherry picked from commit 975bd28)
…rc_1_103_0

fix(bedrock): backport the /v1/messages Invoke streaming pass-through to rc/1.103.0 (BerriAI#42607)
…pend rollup code from rc/1.103.0 (BerriAI#43326)

Reverts the code from BerriAI#41324 (merge 87694c2) and BerriAI#41293 (merge 2e46b10).
The LiteLLM_DailyGlobalSpend model and migration stay so databases that applied
rc.1 keep a consistent table and rollup marker when the feature returns.
… poison the pool connection (BerriAI#43029) (BerriAI#43323)

Prisma types a raw array parameter from the first batch a connection sees. After a flush in
which every member cost was a whole number (a free model), the connection's cached statement
expected int8[] and every later fractional batch on it failed with "improper binary format in
array element", so member spend silently stopped landing while team spend kept rising.

The rows now travel as one JSON document unpacked by jsonb_to_recordset with the column types
declared in SQL, so Postgres types the numbers and the batch shape no longer matters.

(cherry picked from commit 5e4b1b9)

Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
* fix: apply configured cache_control_injection_points beside client cache_control marks (BerriAI#41956)

(cherry picked from commit d8267d5)

* fix(token_counter): count replayed redacted_thinking blocks so prompt_caching keeps pinning (BerriAI#42069)

(cherry picked from commit 0e7cf51)

* fix(mcp): return camelCase tool keys from /v1/mcp/tools after the SDK 2 upgrade (BerriAI#42352)

(cherry picked from commit 5eb4e30)

* fix(rust_bridge): keep the Messages route on Python until the Rust path is ready (BerriAI#42517)

rc/1.103.0 has no token counter or tokenizer routes in the catalog, so only the Messages rule moves to PYTHON_ONLY

(cherry picked from commit 85ed18e)

* fix(proxy): stop leaking periodic tasks on every DB config reload (BerriAI#42784)

(cherry picked from commit fc87a06)

---------

Co-authored-by: Mateo Wang <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…election to rc/1.103.0 (BerriAI#43343)

* feat(otel): emit gen_ai.conversation.id from the caller's session id on v2 LLM spans (BerriAI#42486)

* feat(otel): emit gen_ai.conversation.id from the caller's session id on v2 LLM spans

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): keep the caller's header session under missing_session_id: generate and read replayed payload session ids

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): drop only the proxy-minted session id so a caller id on the other metadata key survives

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): keep a replayed session id hidden when it only echoes the payload trace id

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): keep a replayed session id even when the payload trace id fell back to it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): stop reading the replayed payload's session id, the generated marker does not survive replay

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): audit gen_ai.conversation.id on otel v2 spans through a real proxy, sink and postgres

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): keep otel conversation rigs alive for the whole session so shuffled shards do not reboot the proxy per test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(otel): stop the audit rig proxies from probing sibling test peers for model info

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(otel): record accepted OTLP batches in the sink instead of mutating the collector

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(otel): guard the accepted batch deque so snapshots cannot race sink appends

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mrinal <mrinal@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yucheng <yucheng@berri.ai>
(cherry picked from commit a319690)

* fix(jwt): accept a team alias in x-litellm-team-id (BerriAI#42445)

* fix(jwt): accept a team alias in x-litellm-team-id

The header only matched canonical team ids, so a JWT caller selecting one of their teams by its alias got a 403 even though they belonged to it. The header value is now resolved through the existing alias lookup before the JWT allowed-team check and the DB membership fallback, while a value that is already a team id never costs an alias lookup and denials keep naming the value the caller sent

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(jwt): only alias a header team id the database provably lacks

Under fallback_to_db_teams a header value whose team row read fails for any reason other than TeamNotFoundError now keeps the membership denial instead of falling through to the alias lookup, so a degraded read cannot select a different team that carries the value as an alias. Drops the HeaderTeam docstring that only restated its fields

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: ryan <ryan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit 071cb49)

* fix(jwt): say x-litellm-team-id matched no team id or alias in the 403 (BerriAI#42495)

* fix(jwt): say x-litellm-team-id matched no team id or alias in the 403

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(jwt): tell the caller when x-litellm-team-id names an alias shared by several teams

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(jwt): deny a shared x-litellm-team-id alias exactly like an unknown value

A distinct 403 for an alias several teams share was raised before the
allowed-teams check, so any JWT could probe which aliases exist. The
alias lookup now treats the duplicate as a miss, and both denials say
the value does not resolve to a team id or a unique team alias, which
is true for unknown, unauthorized and duplicate values alike

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: ryan <ryan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit 08639fc)

* fix(jwt): let x-litellm-team-id select DB membership teams when the token also carries a team claim (BerriAI#43206)

* fix(jwt): let x-litellm-team-id select DB membership teams when the token also carries a team claim

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs(jwt): describe header team selection under fallback_to_db_teams

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit 7b4fd47)

---------

Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mrinal <mrinal@berri.ai>
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: ryan <ryan@berri.ai>
Co-authored-by: yassin <yassin@berri.ai>
… direct client mark (BerriAI#43342)

* fix(caching): stand default cache points down when extra_body hides a direct client mark

On native /v1/messages the extra_body envelope is dropped, so a client tool mark or root cache_control reaches Anthropic even when extra_body overrides it. The stand-down check only counted the envelope-merged view and injected two default marks on top of the client's.

(cherry picked from commit 01ef7c4)

* fix(caching): keep chat completions on the envelope-merged mark count for the default stand-down

Chat completions merge extra_body over the request, so a direct tool mark that extra_body replaces never reaches the provider there. Only /v1/messages, where the native transforms drop the envelope, needs to count marks on both sides.

(cherry picked from commit a9c5fa7)
…s again (BerriAI#43382)

Adds the OwnedProxy harness the BerriAI#42486 backport's OTel test imports, registers the BerriAI#43029 backport's spend flush test in contracts.json with its covers marker, and aligns the MCP tool failure assertion with main.
…ers (BerriAI#43388)

Backports the batch cleanup half of BerriAI#43321 so an Azure cancel still in progress after 120s leaves a warning instead of a teardown error, and points the MCP Tools UI test at DeepWiki's renamed ask_wiki_question tool.
…erriAI#43382 and BerriAI#43388 (BerriAI#43391)

The batch cleanup leftover warning used a class defined in a test-directory module. The xdist controller cannot import it, so an uncaught leftover warning crashed the whole e2e run. It is now a plain UserWarning.

The integration harness backport missed the workers option and Gateway.request headers that the OTel conversation id tests use.
…lake tolerance to rc/1.103.0 (BerriAI#43400)

* test(e2e): tolerate provider-side flakes on five full-suite cells (BerriAI#42628)

* test(e2e): tolerate provider-side flakes on five full-suite cells

Mistral OCR retries a provider-relayed 429 with backoff, the Vertex vision
probe turns reasoning off so its 32 tokens go to the answer, the Vertex
cache cell spaces eight never-seen prefixes 15s apart around Google's
nondeterministic minimum-token rejection and prices the cached tokens
instead of prompt_tokens, and the Azure content-policy cell resends the
jailbreak prompt while Azure skips its filter

* test(e2e): shorten the new helper docstrings

* test(e2e): accept a relayed provider 429 on the rust OCR cells

The gateway already retries a provider 429 three times per call and the
Mistral key is shared across pipelines, so a throttle can hold across all
four attempts of the OCR cell. After the bounded retries the cell now
accepts the gateway's faithful relay of the provider's 429 (throttling_error,
code 429) as its second expected outcome; the gateway's own 429 and any
other error still fail the cell at once.

* test(e2e): drop the harness unit tests, the live cells cover the helpers

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
(cherry picked from commit 41ca465)

* fix(streaming): keep litellm Usage on text-completion usage chunks (BerriAI#43047)

* fix(streaming): keep litellm Usage on text-completion usage chunks

* fix(streaming): convert provider usage to litellm Usage instead of dropping it

(cherry picked from commit be35b22)

---------

Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
…ig (BerriAI#43429)

* fix(proxy): unregister logging callbacks removed from the stored config

POST /config/callback/delete saved the config and resynced, but the resync only
ever added callbacks, so a deleted callback kept exporting and kept showing in
/get/config/callbacks as read-only on every worker.

ProxyConfig now tracks which callback list entries each DB config sync
registered and unregisters them once the stored config stops listing them.
Callbacks it did not register (YAML, code) are never touched, and a failed
config load skips the sync instead of treating the config as empty.

* refactor(proxy): keep callback sync comprehensions to one for clause
metadata.requester_metadata embeds per-request headers/traceparent, so
dumping it on gen_ai metrics grows in-process SDK aggregations without
bound. Lift the two bounded grouping fields first so exclude_list can
drop the blob without losing the dimensions Prometheus already exposes.

Co-authored-by: Cursor <cursoragent@cursor.com>
(cherry picked from commit 9af833c)
(cherry picked from commit 14f97d3)
Re-apply of e2beef8 (Noma fork PR #12) onto the 1.102 base: upstream
refactored bedrock_converse_supports_strict_tools with Final annotations,
so the original patch did not cherry-pick cleanly. Same intent — pin the
no-strict payload by returning False, leaving _get_bedrock_converse_strict_tools_flag
wired for a one-line revert once we ship a release with BerriAI#35688.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit 1000cb8)
Automates the previously-manual `docker buildx --push` of the -custom image.
Triggers on push to v*-custom branches + workflow_dispatch; builds
docker/Dockerfile.non_root multi-arch and pushes to ECR tagged with the
branch name (e.g. v1.102.0-custom) + a -<sha> variant. Build-only; deploy
stays a values bump + canary (no ArgoCD sync).

Prereqs for DevOps before first run: fork access to arc-runners-prod;
OIDC trust on role/github-actions-region-deploy for repo Noma-Security/litellm;
ecr:PutImage on the litellm repo for that role (or a dedicated build role).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit 90c1873)
The arc-runners-prod scale set isn't scoped to this fork, so the job
queued forever with no runner. Switch to ubuntu-latest (no ARC scoping)
and add setup-qemu for the linux/arm64 leg. Still needs the OIDC trust
(repo:Noma-Security/litellm:*) + ecr:PutImage on role/github-actions-region-deploy
for the push step.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit d77daae)
The spend counter DualCache used the generic InMemoryCache default of 200
entries, so deployments with more than 200 active budget scopes evicted
still-valid counters and reseeded them from lagging DB spend on the next
request. A key the proxy had just rejected for being over budget could be
admitted again until the batch writer flushed.

Fixes BerriAI#40221

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit 1de5a82)
(cherry picked from commit 3ce17a5)
Bedrock serves its openai.* models (gpt-oss, gpt-5.6-sol/terra/luna) over the
plain Invoke route, and they speak OpenAI Chat Completions in both directions.
`transform_request` has delegated those to AmazonBedrockOpenAIConfig since they
landed, but the response side never grew the matching branch, so a valid body
came back through Titan's `results[0].outputText` and blew up as
`BedrockException - ... 'NoneType' object is not subscriptable`. Streaming had
the mirrored gap: no shape the base chunk parser sniffs for matches an OpenAI
chunk, so every delta decoded to empty text -- a finished stream with no tokens.

The trigger was deploy-new-model validating bedrock/us.openai.gpt-5.6-luna:
Bedrock answered correctly and we reported its answer as a Bedrock error.

Rather than special-case openai, mirror the request dispatch. `transform_request`
enumerates providers explicitly and raises 404 on anything unknown; the response
side used a catch-all `else` that assumed Titan, so every provider it did not
name inherited Titan's parse. Making it explicit shows the fallthrough was
already mis-parsing moonshot, qwen2, qwen3 and stability the same way (each has
its own config today, so nothing reached it -- but the trap was armed for the
next provider added to the request side).

Note: BedrockError now propagates instead of being re-wrapped as a blanket 422.
That is required for the unknown-provider 404 to survive, and it stops delegated
configs' status codes (429, 400) from being flattened into 422.

Upstream BerriAI#37132 reports all of this; it was closed with no fix
and main is still affected, so this stays a fork patch.

Co-authored-by: Cursor <cursoragent@cursor.com>
(cherry picked from commit ad3790c)
…forwarded

The Noma fork pins bedrock_converse_supports_strict_tools to False, so
toolSpec.strict is never sent to Converse. The upstream tests asserting
strict is kept for Sonnet 4.5/4.6 and Opus 4.5/4.6 contradict that pin.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every change on main is either upstream in v1.103.0 (Redis breaker fixes
BerriAI#40764/BerriAI#40624/BerriAI#40817) or re-applied on this branch (otel lift, bedrock strict
pin, spend_counter_cache, custom-image CI), so the 1.103.0 tree is kept as is.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@github-advanced-security github-advanced-security AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

CodeQL found more than 20 potential problems in the proposed changes. Check the Files changed tab for more details.

EvictedClientCloser held queued clients through weakrefs. An SDK client is a
reference cycle that a generational collection reclaims within minutes, well
before the 900s grace window, so the AsyncAzureOpenAI client evicted from the
LLM client cache every hour was never closed: its aiohttp session was
finalized by the garbage collector instead ("Unclosed client session").

In prod every liveness-probe kill of the proxy since the v1.102 rollout
started at one of those finalizations (27 of 28 kills; ~2 expected by chance).

- Hold queued clients strongly until they are closed on their own loop.
- Drop entries whose event loop has closed; nothing can close them anymore.
- A transport on a shared aiohttp session reports leases for every client on
  that session, so don't treat them as this client's in-flight requests.
  Otherwise clients on the proxy's shared session look busy forever, are never
  closed and accumulate once they are held strongly.

(cherry picked from commit 143e720)

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, have a team admin enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 6f1b150. Configure here.

def _drop_buckets_of_closed_loops_locked(self) -> None:
"""A loop's bucket is never drained once that loop is closed, so release it."""
for key in [k for k, b in self._buckets.items() if b and _loop_is_gone(b[0])]:
self._pending_count -= len(self._buckets.pop(key))

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Closed loops drop closable clients

Medium Severity

loop_ref is stored for every eviction, including sync clients that do not need a loop. _drop_buckets_of_closed_loops_locked then discards an entire bucket when the first entry's loop is gone, so _CLOSABLE_ANYWHERE clients are dropped without close(). Sync handlers queued from a worker loop that later ends keep their pools until GC, which is the leak this change is meant to stop.

Additional Locations (2)
Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit 6f1b150. Configure here.

This branch had an error being deployed

1 failed deployment
prod — 6f1b1502 Deployed Oct 6, 2026 by omerbd21-noma via build #4
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.