Skip to content

fix(websearch): forward the deployment api_base to agentic follow-up calls on /v1/messages - #41918

Merged
mateo-berri merged 1 commit into
mainfrom
litellm_websearch_followup_api_base
Sep 19, 2026
Merged

mateo-berri merged 1 commit into
mainfrom
litellm_websearch_followup_api_base

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Azure Foundry Claude with web search interception never returns search results
  • Clients get a dangling litellm_web_search tool call; Claude Code shows Did 0 searches
  • Vertex Claude already works, so native handling stays on for both providers

How it solves it:

  • The follow-up search call now carries the deployment's api_base and api_key
  • Only set values are forwarded, so Vertex and Anthropic calls are unchanged
  • Regression test: the follow-up must hit the Foundry /v1/messages URL

User Flow

Before: a developer running Claude Code through LiteLLM against an Azure Foundry Claude deployment with web search interception on gets Did 0 searches for every WebSearch

  1. The proxy admin adds azure_ai/claude-fable-5-1 with its api_base and api_key, a tavily entry under search_tools, and turns on websearch_interception for azure_ai
  2. The developer starts Claude Code with ANTHROPIC_BASE_URL=https://litellm-domain and ANTHROPIC_MODEL=foundry-claude-fable-5-1 and asks it to use WebSearch
  3. Claude Code sends POST https://litellm-domain/v1/messages with "tools": [{"type": "web_search_20250305", "name": "web_search"}]
  4. 200 comes back with "stop_reason": "tool_use" and a tool_use block named litellm_web_search, no web_search_tool_result; Claude Code renders Did 0 searches and has to fetch pages by hand to answer

After: the same WebSearch returns real results, so Claude Code shows Did 1 search and answers from them

  1. The proxy admin adds azure_ai/claude-fable-5-1 with its api_base and api_key, a tavily entry under search_tools, and turns on websearch_interception for azure_ai
  2. The developer starts Claude Code with ANTHROPIC_BASE_URL=https://litellm-domain and ANTHROPIC_MODEL=foundry-claude-fable-5-1 and asks it to use WebSearch
  3. Claude Code sends POST https://litellm-domain/v1/messages with "tools": [{"type": "web_search_20250305", "name": "web_search"}]
  4. 200 comes back with "stop_reason": "end_turn", a server_tool_use block, a web_search_tool_result block holding 10 results, and the answer text; Claude Code renders Did 1 search in 17s and cites the results

Relevant issues

Affected release

Linear ticket

Resolves LIT-5418

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Shared setup, identical for both legs. Two proxies, each booted from its own worktree with --num_workers 2 and no database, against the real Azure Foundry, Vertex AI, and Tavily APIs:

# Before: worktree at the merge base 15f63c33bf, port 41678
# After:  worktree at the PR tip b1b7af884a, port 46224
PYTHONPATH=<worktree> python litellm/proxy/proxy_cli.py --config lit5418-config.yaml --port <port> --num_workers 2 --detailed_debug

.env (mode 0600) carries AZURE_FOUNDRY_API_BASE, AZURE_FOUNDRY_API_KEY, GOOGLE_APPLICATION_CREDENTIALS, VERTEXAI_PROJECT, VERTEXAI_LOCATION=us-east5, TAVILY_API_KEY, and LITELLM_MASTER_KEY

lit5418-config.yaml:

model_list:
  - model_name: vertex-claude-sonnet-4-5
    litellm_params:
      model: vertex_ai/claude-sonnet-4-5@20250929
      vertex_project: os.environ/VERTEXAI_PROJECT
      vertex_location: os.environ/VERTEXAI_LOCATION
  - model_name: foundry-claude-fable-5-1
    litellm_params:
      model: azure_ai/claude-fable-5-1
      api_base: os.environ/AZURE_FOUNDRY_API_BASE
      api_key: os.environ/AZURE_FOUNDRY_API_KEY
  - model_name: foundry-claude-opus-4-8
    litellm_params:
      model: azure_ai/claude-opus-4-8
      api_base: os.environ/AZURE_FOUNDRY_API_BASE
      api_key: os.environ/AZURE_FOUNDRY_API_KEY
  - model_name: foundry-claude-sonnet-4-6
    litellm_params:
      model: azure_ai/claude-sonnet-4-6
      api_base: os.environ/AZURE_FOUNDRY_API_BASE
      api_key: os.environ/AZURE_FOUNDRY_API_KEY

search_tools:
  - search_tool_name: tavily-search
    litellm_params:
      search_provider: tavily
      api_key: os.environ/TAVILY_API_KEY

general_settings:
  master_key: os.environ/LITELLM_MASTER_KEY

litellm_settings:
  callbacks: ["websearch_interception"]
  websearch_interception_params:
    enabled_providers: ["vertex_ai", "azure_ai"]
    search_tool_name: tavily-search

req.json, the body every curl case sends (MODEL swapped per deployment):

{"model":"MODEL","max_tokens":1024,"system":"You are a web search assistant. Use the web_search tool to answer, then summarize the results briefly with the source URLs.","messages":[{"role":"user","content":"Perform a web search for the query: latest LiteLLM release version on GitHub"}],"tools":[{"type":"web_search_20250305","name":"web_search","max_uses":8}]}

run.sh <label> <port> <model...>, which posts req.json per model and prints one summary line per response:

label=$1; port=$2; shift 2
for m in "$@"; do
  body=$(sed "s/MODEL/${m}/" req.json)
  out="out/${label}-${m}.json"
  code=$(curl -s -m 300 -o "$out" -w '%{http_code}' "http://127.0.0.1:${port}/v1/messages" \
    -H "x-api-key: ${LITELLM_MASTER_KEY}" -H 'anthropic-version: 2023-06-01' -H 'content-type: application/json' -d "$body")
  python3 - "$out" "$label" "$m" "$code" <<'PY'
import json, sys
p, label, m, code = sys.argv[1:5]
d = json.load(open(p))
if d.get("type") == "error":
    print(f"== {label} {m}: HTTP {code} ERROR {json.dumps(d)[:400]}"); sys.exit()
content = d.get("content", [])
blocks = [b.get("type") for b in content]
names = [b.get("name") for b in content if b.get("type") in ("server_tool_use", "tool_use")]
n_results = [len(b.get("content", [])) for b in content if b.get("type") == "web_search_tool_result"]
text = " ".join(b.get("text", "") for b in content if b.get("type") == "text")[:220].replace("\n", " ")
u = d.get("usage", {})
print(f"== {label} {m}: HTTP {code} id={d.get('id')} stop={d.get('stop_reason')} blocks={blocks} tool_names={names} search_results={n_results} usage=in{u.get('input_tokens')}/out{u.get('output_tokens')}")
print(f"   text: {text}")
for b in content:
    if b.get("type") == "tool_use":
        print(f"   tool_use input: {json.dumps(b.get('input'))}")
PY
done

Claude Code v2.1.277 was driven interactively in tmux against each proxy:

env ANTHROPIC_BASE_URL=http://127.0.0.1:<port> ANTHROPIC_API_KEY=$LITELLM_MASTER_KEY ANTHROPIC_MODEL=foundry-claude-fable-5-1 \
  CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1 CLAUDE_CODE_ENABLE_GATEWAY_MODEL_DISCOVERY=1 claude --model foundry-claude-fable-5-1

with the prompt Use your WebSearch tool to find the latest LiteLLM release version on GitHub and answer in one sentence with the source URL

Before (15f63c3)

Claude Code WebSearch, foundry-claude-fable-5-1

  1. Start Claude Code against port 41678 and send the prompt
  2. The pane shows Web Search("LiteLLM latest release GitHub BerriAI litellm releases") followed by Did 0 searches in 7s; Claude Code then falls back to Fetch(https://github.com/BerriAI/litellm/releases/latest) to find an answer at all

pr41918-b1b7af884a-lit5418-cc2-before.png

curl /v1/messages, three Foundry Claude deployments

  1. ./run.sh BASE 41678 foundry-claude-fable-5-1 foundry-claude-opus-4-8 foundry-claude-sonnet-4-6
  2. Every response stops on a dangling litellm_web_search tool call and carries no search results and no answer:
== BASE foundry-claude-fable-5-1: HTTP 200 id=msg_011CfBwHr4QMtpHTPPmHRXZ3 stop=tool_use blocks=['thinking', 'tool_use'] tool_names=['litellm_web_search'] search_results=[] usage=in463/out86
   text:
   tool_use input: {"query": "latest LiteLLM release version on GitHub"}
== BASE foundry-claude-opus-4-8: HTTP 200 id=msg_011CfBwJGoSMPcezmQWT9db5 stop=tool_use blocks=['text', 'tool_use'] tool_names=['litellm_web_search'] search_results=[] usage=in465/out88
   text: I'll search for the latest LiteLLM release version on GitHub.
   tool_use input: {"query": "latest LiteLLM release version on GitHub"}
== BASE foundry-claude-sonnet-4-6: HTTP 200 id=msg_011CfBwJdgYp4RB3xH5rJHaw stop=tool_use blocks=['tool_use'] tool_names=['litellm_web_search'] search_results=[] usage=in638/out67
   text:
   tool_use input: {"query": "latest LiteLLM release version on GitHub"}

curl /v1/messages streaming, foundry-claude-fable-5-1

  1. curl -s -N http://127.0.0.1:41678/v1/messages -H "x-api-key: $LITELLM_MASTER_KEY" -H 'anthropic-version: 2023-06-01' -H 'content-type: application/json' -d "$(jq -c '. + {stream: true, model: "foundry-claude-fable-5-1"}' req.json)" | tee out/BASE-stream.sse | grep '^data:' | sed 's/^data: //' | jq -c 'select(.type=="content_block_start" or .type=="message_delta") | if .type=="content_block_start" then {index, block: .content_block.type, name: .content_block.name} else {stop_reason: .delta.stop_reason, usage} end'
  2. The stream ends on the same dangling tool call, with no web_search_tool_result block:
{"index":0,"block":"thinking","name":null}
{"index":1,"block":"tool_use","name":"litellm_web_search"}
{"stop_reason":"tool_use","usage":{"output_tokens":101,"input_tokens":463,"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
  1. grep -c web_search_result out/BASE-stream.sse prints 0; grep -c litellm_web_search out/BASE-stream.sse prints 1

curl /v1/messages, vertex-claude-sonnet-4-5 (control)

  1. ./run.sh BASE 41678 vertex-claude-sonnet-4-5
  2. Vertex already returns the native blocks at the merge base:
== BASE vertex-claude-sonnet-4-5: HTTP 200 id=msg_vrtx_011CfBwKBF6WP3jePAAA6G6u stop=end_turn blocks=['server_tool_use', 'web_search_tool_result', 'text'] tool_names=['web_search'] search_results=[10] usage=in4991/out270
   text: Based on the search results, here's a summary of the latest LiteLLM release information:  ## Latest LiteLLM Release Version  **Latest Stable Release: v1.101.0** (released September 15)  According to the GitHub releases p

After (b1b7af8)

Claude Code WebSearch, foundry-claude-fable-5-1

  1. Start Claude Code against port 46224 and send the prompt
  2. The pane shows Web Search("LiteLLM latest release GitHub BerriAI") followed by Did 1 search in 17s, and the answer cites the release from the results: The latest stable LiteLLM release on GitHub is v1.101.0, published Sep 15, 2026

pr41918-b1b7af884a-lit5418-cc2-after.png

curl /v1/messages, three Foundry Claude deployments

  1. ./run.sh AFTER 46224 foundry-claude-fable-5-1 foundry-claude-opus-4-8 foundry-claude-sonnet-4-6
  2. Every response now ends the turn with server_tool_use, a web_search_tool_result block holding 10 results, and an answer naming v1.101.0:
== AFTER foundry-claude-fable-5-1: HTTP 200 id=msg_011CfBwLfr5SFs6va7UbqWPE stop=end_turn blocks=['server_tool_use', 'web_search_tool_result', 'text'] tool_names=['web_search'] search_results=[10] usage=in6238/out600
   text: Here's a summary of what I found:  **Latest LiteLLM release on GitHub: v1.101.0** - Marked as the "Latest" release on the BerriAI/litellm Releases page, published Sept 15 by yuneng-berri (commit `18243cd`, GPG-verified).
== AFTER foundry-claude-opus-4-8: HTTP 200 id=msg_011CfBwMwqmz8WcoHMkGo9wV stop=end_turn blocks=['server_tool_use', 'web_search_tool_result', 'text'] tool_names=['web_search'] search_results=[10] usage=in6240/out472
   text: Based on my web search, here's what I found:  ## Latest LiteLLM Release on GitHub  The **latest stable release is `v1.101.0`**, released on **September 15, 2026** by BerriAI on GitHub.  ### Additional details: - There ar
== AFTER foundry-claude-sonnet-4-6: HTTP 200 id=msg_011CfBwNgeFiutYWzyULJpBd stop=end_turn blocks=['server_tool_use', 'web_search_tool_result', 'text'] tool_names=['web_search'] search_results=[10] usage=in5099/out321
   text: ## Latest LiteLLM Release on GitHub  Here's a summary of the latest LiteLLM release information found:  | Release Type | Version | Date | |---|---|---| | **Latest Stable Release** | **v1.101.0** | September 15, 2026 | |

curl /v1/messages streaming, foundry-claude-fable-5-1

  1. curl -s -N http://127.0.0.1:46224/v1/messages -H "x-api-key: $LITELLM_MASTER_KEY" -H 'anthropic-version: 2023-06-01' -H 'content-type: application/json' -d "$(jq -c '. + {stream: true, model: "foundry-claude-fable-5-1"}' req.json)" | tee out/AFTER-stream.sse | grep '^data:' | sed 's/^data: //' | jq -c 'select(.type=="content_block_start" or .type=="message_delta") | if .type=="content_block_start" then {index, block: .content_block.type, name: .content_block.name} else {stop_reason: .delta.stop_reason, usage} end'
  2. The stream now carries the native search blocks and ends the turn:
{"index":0,"block":"server_tool_use","name":"web_search"}
{"index":1,"block":"web_search_tool_result","name":null}
{"index":2,"block":"text","name":null}
{"stop_reason":"end_turn","usage":{"output_tokens":469,"input_tokens":6308,"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
  1. grep -c web_search_result out/AFTER-stream.sse prints 1; grep -c litellm_web_search out/AFTER-stream.sse prints 0

curl /v1/messages, vertex-claude-sonnet-4-5 (control)

  1. ./run.sh AFTER 46224 vertex-claude-sonnet-4-5
  2. Vertex behaves exactly as before:
== AFTER vertex-claude-sonnet-4-5: HTTP 200 id=msg_vrtx_011CfBwPUAwXNcoqN4hi9Yn7 stop=end_turn blocks=['server_tool_use', 'web_search_tool_result', 'text'] tool_names=['web_search'] search_results=[10] usage=in5117/out298
   text: Based on the search results, here's what I found about the latest LiteLLM release version on GitHub:  ## Latest LiteLLM Release Version  **Latest Stable Release: v1.101.0** - Released on September 15, 2026 - This is the

Observations from the run:

  • Both legs ran on two-worker proxies, no DB, real Foundry, Vertex, Tavily
  • A Foundry 429 on the follow-up still renders Did 0 searches; pre-existing
  • Two Claude Code sessions in parallel tripped Foundry's tokens-per-minute limit
  • Vertex Claude native search was already right at the merge base

Merged-tree check (origin/main 078a604 + this tip, local merge only)

main moved past the merge base, including #41905 in the same handler, so the same scenarios were re-run on a two-worker proxy booted from a local merge of the tip into origin/main 078a604 (PYTHONPATH asserted on the listener):

== MERGED foundry-claude-fable-5-1: HTTP 200 id=msg_011CfCB3hJnmPAoPWWLAt8Lc stop=end_turn blocks=['server_tool_use', 'web_search_tool_result', 'text'] tool_names=['web_search'] search_results=[10]
== MERGED vertex-claude-sonnet-4-5: HTTP 200 id=msg_vrtx_011CfCB4bqfBxhtrCrL4hszF stop=end_turn blocks=['server_tool_use', 'web_search_tool_result', 'text'] tool_names=['web_search'] search_results=[10]
foundry-claude-fable-5-1 stream=true: HTTP 200 events=10 stop=end_turn blocks=['server_tool_use', 'web_search_tool_result', 'text']
S1 plain, S2 plain stream, S3 client tool (get_weather -> stop=tool_use) on both providers: unchanged
S5 vertex websearch stream: HTTP 200 stop=end_turn blocks=['server_tool_use', 'web_search_tool_result', 'text']

Type

🐛 Bug Fix

Caveats (if any)

Low

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

  • b1b7af8 passes /live-pr-risk

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@codspeed

codspeed Bot commented Sep 19, 2026

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_websearch_followup_api_base (b1b7af8) with main (15f63c3)1

Open in CodSpeed

Footnotes

  1. No successful run was found on main (5c66582) during the generation of this report, so 15f63c3 was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩

@greptile-apps

greptile-apps Bot commented Sep 19, 2026

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

The PR appears safe to merge, with both response paths preserving the deployment parameters needed by agentic follow-up calls

Summary

This PR preserves the selected deployment endpoint and credential when Anthropic Messages agentic hooks issue follow-up calls.

  • Builds shared agentic-hook kwargs containing the named api_key and api_base parameters
  • Uses those kwargs for both streaming and non-streaming response paths
  • Adds regression coverage for both paths using a mocked Azure AI deployment

Reviews (1) · Last reviewed commit: "fix(websearch): forward the deployment a..."

@codecov

codecov Bot commented Sep 19, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit b1b7af8. Configure here.

@mateo-berri mateo-berri left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@mateo-berri
mateo-berri merged commit faed57f into main Sep 19, 2026
143 of 147 checks passed
@mateo-berri
mateo-berri deleted the litellm_websearch_followup_api_base branch September 19, 2026 04:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant