Repository navigation
fix(websearch): forward the deployment api_base to agentic follow-up calls on /v1/messages - #41918
Conversation
…calls on /v1/messages
|
I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".
|
|
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit b1b7af8. Configure here.
TLDR
Problem this solves:
litellm_web_searchtool call; Claude Code showsDid 0 searchesHow it solves it:
api_baseandapi_key/v1/messagesURLUser Flow
Before: a developer running Claude Code through LiteLLM against an Azure Foundry Claude deployment with web search interception on gets
Did 0 searchesfor every WebSearchazure_ai/claude-fable-5-1with itsapi_baseandapi_key, a tavily entry undersearch_tools, and turns onwebsearch_interceptionforazure_aiANTHROPIC_BASE_URL=https://litellm-domainandANTHROPIC_MODEL=foundry-claude-fable-5-1and asks it to use WebSearch"tools": [{"type": "web_search_20250305", "name": "web_search"}]"stop_reason": "tool_use"and atool_useblock namedlitellm_web_search, noweb_search_tool_result; Claude Code rendersDid 0 searchesand has to fetch pages by hand to answerAfter: the same WebSearch returns real results, so Claude Code shows
Did 1 searchand answers from themazure_ai/claude-fable-5-1with itsapi_baseandapi_key, a tavily entry undersearch_tools, and turns onwebsearch_interceptionforazure_aiANTHROPIC_BASE_URL=https://litellm-domainandANTHROPIC_MODEL=foundry-claude-fable-5-1and asks it to use WebSearch"tools": [{"type": "web_search_20250305", "name": "web_search"}]"stop_reason": "end_turn", aserver_tool_useblock, aweb_search_tool_resultblock holding 10 results, and the answer text; Claude Code rendersDid 1 search in 17sand cites the resultsRelevant issues
Affected release
Linear ticket
Resolves LIT-5418
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Shared setup, identical for both legs. Two proxies, each booted from its own worktree with
--num_workers 2and no database, against the real Azure Foundry, Vertex AI, and Tavily APIs:.env(mode 0600) carriesAZURE_FOUNDRY_API_BASE,AZURE_FOUNDRY_API_KEY,GOOGLE_APPLICATION_CREDENTIALS,VERTEXAI_PROJECT,VERTEXAI_LOCATION=us-east5,TAVILY_API_KEY, andLITELLM_MASTER_KEYlit5418-config.yaml:req.json, the body every curl case sends (MODELswapped per deployment):{"model":"MODEL","max_tokens":1024,"system":"You are a web search assistant. Use the web_search tool to answer, then summarize the results briefly with the source URLs.","messages":[{"role":"user","content":"Perform a web search for the query: latest LiteLLM release version on GitHub"}],"tools":[{"type":"web_search_20250305","name":"web_search","max_uses":8}]}run.sh <label> <port> <model...>, which postsreq.jsonper model and prints one summary line per response:Claude Code v2.1.277 was driven interactively in tmux against each proxy:
with the prompt
Use your WebSearch tool to find the latest LiteLLM release version on GitHub and answer in one sentence with the source URLBefore (15f63c3)
Claude Code WebSearch, foundry-claude-fable-5-1
Web Search("LiteLLM latest release GitHub BerriAI litellm releases")followed byDid 0 searches in 7s; Claude Code then falls back toFetch(https://github.com/BerriAI/litellm/releases/latest)to find an answer at allcurl /v1/messages, three Foundry Claude deployments
./run.sh BASE 41678 foundry-claude-fable-5-1 foundry-claude-opus-4-8 foundry-claude-sonnet-4-6litellm_web_searchtool call and carries no search results and no answer:curl /v1/messages streaming, foundry-claude-fable-5-1
curl -s -N http://127.0.0.1:41678/v1/messages -H "x-api-key: $LITELLM_MASTER_KEY" -H 'anthropic-version: 2023-06-01' -H 'content-type: application/json' -d "$(jq -c '. + {stream: true, model: "foundry-claude-fable-5-1"}' req.json)" | tee out/BASE-stream.sse | grep '^data:' | sed 's/^data: //' | jq -c 'select(.type=="content_block_start" or .type=="message_delta") | if .type=="content_block_start" then {index, block: .content_block.type, name: .content_block.name} else {stop_reason: .delta.stop_reason, usage} end'web_search_tool_resultblock:grep -c web_search_result out/BASE-stream.sseprints0;grep -c litellm_web_search out/BASE-stream.sseprints1curl /v1/messages, vertex-claude-sonnet-4-5 (control)
./run.sh BASE 41678 vertex-claude-sonnet-4-5After (b1b7af8)
Claude Code WebSearch, foundry-claude-fable-5-1
Web Search("LiteLLM latest release GitHub BerriAI")followed byDid 1 search in 17s, and the answer cites the release from the results:The latest stable LiteLLM release on GitHub is v1.101.0, published Sep 15, 2026curl /v1/messages, three Foundry Claude deployments
./run.sh AFTER 46224 foundry-claude-fable-5-1 foundry-claude-opus-4-8 foundry-claude-sonnet-4-6server_tool_use, aweb_search_tool_resultblock holding 10 results, and an answer naming v1.101.0:curl /v1/messages streaming, foundry-claude-fable-5-1
curl -s -N http://127.0.0.1:46224/v1/messages -H "x-api-key: $LITELLM_MASTER_KEY" -H 'anthropic-version: 2023-06-01' -H 'content-type: application/json' -d "$(jq -c '. + {stream: true, model: "foundry-claude-fable-5-1"}' req.json)" | tee out/AFTER-stream.sse | grep '^data:' | sed 's/^data: //' | jq -c 'select(.type=="content_block_start" or .type=="message_delta") | if .type=="content_block_start" then {index, block: .content_block.type, name: .content_block.name} else {stop_reason: .delta.stop_reason, usage} end'grep -c web_search_result out/AFTER-stream.sseprints1;grep -c litellm_web_search out/AFTER-stream.sseprints0curl /v1/messages, vertex-claude-sonnet-4-5 (control)
./run.sh AFTER 46224 vertex-claude-sonnet-4-5Observations from the run:
Did 0 searches; pre-existingMerged-tree check (origin/main 078a604 + this tip, local merge only)
mainmoved past the merge base, including #41905 in the same handler, so the same scenarios were re-run on a two-worker proxy booted from a local merge of the tip into origin/main 078a604 (PYTHONPATHasserted on the listener):Type
🐛 Bug Fix
Caveats (if any)
Low
litellm_web_searchtool call: LIT-8165UserByModelByMinuteUncachedInputTokens) made Claude Code showDid 0 searchesmainand untouched by this diff (classified, not fixed here, per the fleet-wide-red rule; CircleCI is not a required check)litellm_utils_testing/test_models_by_provider:transcribecost-map entry from feat(proxy): add Amazon Transcribe pass-through with completion-time job pricing #41515, in the merge base; fails identically at origin/main locally; fix PR fix(proxy): register transcribe as a known provider for model grants #41926proxy_e2e_anthropic_messages_tests/test_bedrock_invoke_messages_with_all_beta_headers[...]x2: Bedrockinvalid beta flagsince fix(anthropic): register thinking-binding-controls-2026-08-01 in beta headers config #41203; red on main's own pipeline 89818 (job 2188120); LIT-8149integration-providers/test_retry_counts_and_public_errors_match_actual_provider_attempts:'server_error' != 'internal_server_error'since fix(responses): emit typed streaming failure events #40243 echoes the upstream 500 body's type; green on branch pipeline 89821 (main at 00:30Z), red on every tree that also has the fix(proxy): refuse config-owned keys on POST /config/update #41868 fixhandles_web_search_natively()still defaults to True; flipping it is a separate change per the ticketFinal Attestation
The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR
b1b7af8 passes /live-pr-risk