Skip to content

fix(responses): drop tool_search and local_shell in the chat completions bridge - #41953

Merged
mateo-berri merged 1 commit into
mainfrom
litellm_bridge_drop_tool_search
Sep 19, 2026
Merged

mateo-berri merged 1 commit into
mainfrom
litellm_bridge_drop_tool_search

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Codex 0.140+ sends a hosted tool_search tool on every turn
  • The Responses to Chat Completions bridge forwarded it verbatim
  • Azure and OpenAI chat completions answer 400, Codex never gets a reply
  • Codex's parallel title request sends tools: [] with parallel_tool_calls, a second 400

How it solves it:

  • The bridge drops tool_search and local_shell with the existing warning log
  • Same treatment computer_use, image_generation, and shell already get
  • parallel_tool_calls is dropped when no chat tools remain
  • Anthropic and Vertex server tools keep passing through untouched

User Flow

Before: a developer running Codex CLI with wire_api = "responses" against a proxy deployment that has use_chat_completions_api: true gets a 400 on their first message and never an answer

  1. They point Codex at the proxy in ~/.codex/config.toml: base_url = "http://<proxy>/v1", wire_api = "responses", model = "gpt-5.4-mini"
  2. They open Codex and type Reply with just the word pong.
  3. Codex sends POST http:///v1/responses with 10 tools (7 function tools, the custom apply_patch tool, one {"type": "tool_search", "execution": "client", ...}, and one web_search) plus parallel_tool_calls: true
  4. The proxy answers 400 with Invalid value: 'tool_search'. Supported values are: 'function' and 'custom'. and param: tools[8].type
  5. Codex prints that error in the chat pane and the prompt box goes back to idle with no reply
  6. Codex's parallel session-title request, POST http:///v1/responses with tools: [] and parallel_tool_calls: true, also gets 400 'parallel_tool_calls' is only allowed when 'tools' are specified., so the session never gets a title

After: the same Codex turn gets an answer and a session title

  1. They point Codex at the proxy in ~/.codex/config.toml: base_url = "http://<proxy>/v1", wire_api = "responses", model = "gpt-5.4-mini"
  2. They open Codex and type Reply with just the word pong.
  3. Codex sends POST http:///v1/responses with the same 10 tools plus parallel_tool_calls: true
  4. The proxy answers 200 with a streamed response whose output_text is pong
  5. Codex prints pong in the chat pane and marks the turn done
  6. The parallel session-title request, POST http:///v1/responses with tools: [] and parallel_tool_calls: true, answers 200 and the status line shows the generated title

Relevant issues

Relates to #27655 (its local_shell repro and the tool_search pass-through it lists are fixed here; the code_interpreter conversion and per-provider computer_use mapping it also asks for are out of scope)

Relates to #33779 (a Codex namespace tool reaching a chat completions provider, a sibling shape the bridge already converts)

Affected release

Linear ticket

Resolves LIT-7902

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Both sides boot the proxy from a clean worktree at the named commit with two uvicorn workers (--num_workers 2) on a random free port, no database, and the real Azure OpenAI gpt-5.4-mini deployment behind use_chat_completions_api: true, so every request below cost real Azure tokens. The end-user client is Codex CLI 0.155.1 driven interactively in its TUI under tmux. Before ran on port 21753 and After on port 37053; $PORT below stands for the side's port

Shared setup

proxy_config.yaml:

model_list:
  - model_name: gpt-5.4-mini
    litellm_params:
      model: azure/gpt-5.4-mini
      api_base: os.environ/AZURE_FOUNDRY_API_BASE
      api_key: os.environ/AZURE_FOUNDRY_API_KEY
      use_chat_completions_api: true
general_settings:
  master_key: sk-1234

Proxy boot, from the worktree checked out at the side's commit:

python litellm/proxy/proxy_cli.py --config proxy_config.yaml --port $PORT --num_workers 2
curl -s http://127.0.0.1:$PORT/health/liveliness

Codex config.toml in a scratch CODEX_HOME (the last two blocks are what Codex 0.155.1 writes itself after its folder-trust prompt and its model-migration notice, kept so the run lands straight on the prompt with gpt-5.4-mini):

model = "gpt-5.4-mini"
model_reasoning_effort = "medium"
model_provider = "litellm"

[model_providers.litellm]
name = "litellm"
base_url = "http://127.0.0.1:$PORT/v1"
env_key = "LITELLM_QA_KEY"
wire_api = "responses"

[projects."/path/to/codex_project"]
trust_level = "trusted"

[notice.model_migrations]
"gpt-5.4-mini" = "gpt-5.6-luna"

Codex launch and drive:

tmux new-session -d -s codexqa -x 220 -y 60 -c /path/to/codex_project "env CODEX_HOME=/path/to/codex_home LITELLM_QA_KEY=sk-1234 codex"
tmux send-keys -t codexqa -l 'Reply with just the word pong.'; sleep 1; tmux send-keys -t codexqa Enter
sleep 6; tmux capture-pane -p -t codexqa -S -200

codex_tools_body.json is the /v1/responses body Codex 0.155.1 sends for that turn: its 10 tools in its order (exec_command, write_stdin, request_user_input, the custom apply_patch, view_image, get_goal, create_goal, update_goal, tool_search, web_search) plus parallel_tool_calls: true. tools[8] is the hosted tool this PR drops:

{"type": "tool_search", "execution": "client", "description": "# Tool discovery\n\nSearches over deferred tool metadata with BM25 and exposes matching tools for the next model call.\n\nYou have access to tools from the following sources:\n- Multi-agent tools: Spawn and manage sub-agents.\nSome of the tools may not have been provided to you upfront, and you should use this tool (`tool_search`) to search for the required tools. For MCP tool discovery, always use `tool_search` instead of `list_mcp_resources` or `list_mcp_resource_templates`.", "parameters": {"type": "object", "properties": {"limit": {"type": "number", "description": "Maximum number of tools to return. Defaults to 8."}, "query": {"type": "string", "description": "Search query for deferred tools."}}, "required": ["query"], "additionalProperties": false}}

Before (825e287)

Codex CLI turn

  1. Launch Codex with the config above, type Reply with just the word pong., press Enter
  2. The pane prints the provider's 400 where the reply should be, the prompt goes back to idle, and the status line keeps the directory path with no session title:
› Reply with just the word pong.
■ {"error":{"message":"litellm.BadRequestError: AzureException BadRequestError - Invalid value: 'tool_search'. Supported values are: 'function' and 'custom'.. Received Model Group=gpt-5.4-mini\nAvailable Model Group Fallbacks=None","type":"invalid_request_error","param":"tools[8].type","code":"400"}}
› Ask Codex to do anything
  gpt-5.4-mini medium · /private/tmp/claude-501/-Users-mateo-Development-litellm/13337e26-dc6e-4365-a37c-57d6c7feb0a9/scratchpad/lit7902/qa/before/codex_project

pr41953-5477dbe74c-lit7902-before-codex-pane.png

curl, the exact Codex request

  1. curl -sS -i http://127.0.0.1:$PORT/v1/responses -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d @codex_tools_body.json, run twice so each worker serves one
  2. Both runs:
HTTP/1.1 400 Bad Request
{"error":{"message":"litellm.BadRequestError: AzureException BadRequestError - Invalid value: 'tool_search'. Supported values are: 'function' and 'custom'.. Received Model Group=gpt-5.4-mini\nAvailable Model Group Fallbacks=None","type":"invalid_request_error","param":"tools[8].type","code":"400"}}

curl, Codex's session-title request

  1. curl -sS -i http://127.0.0.1:$PORT/v1/responses -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{"model":"gpt-5.4-mini","input":"Reply with just the word pong.","tools":[],"parallel_tool_calls":true}', run twice
  2. Both runs:
HTTP/1.1 400 Bad Request
{"error":{"message":"litellm.BadRequestError: AzureException BadRequestError - Invalid value for 'parallel_tool_calls': 'parallel_tool_calls' is only allowed when 'tools' are specified.. Received Model Group=gpt-5.4-mini\nAvailable Model Group Fallbacks=None","type":"invalid_request_error","param":"parallel_tool_calls","code":"400"}}

curl, one function tool with parallel_tool_calls

  1. curl -sS -i http://127.0.0.1:$PORT/v1/responses -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{"model":"gpt-5.4-mini","input":"Reply with just the word pong.","parallel_tool_calls":true,"tools":[{"type":"function","name":"get_weather","description":"Get weather","parameters":{"type":"object","properties":{"city":{"type":"string"}},"required":["city"]},"strict":false}]}'
  2. HTTP/1.1 200 OK, "status":"completed", output_text pong, usage.total_tokens 132

curl, no tools

  1. curl -sS -i http://127.0.0.1:$PORT/v1/responses -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{"model":"gpt-5.4-mini","input":"Reply with just the word pong."}'
  2. HTTP/1.1 200 OK, "status":"completed", output_text pong, usage.total_tokens 17

After (5477dbe)

Codex CLI turn

  1. Launch Codex with the config above, type Reply with just the word pong., press Enter
  2. The pane shows • pong with the turn marked done, and the status line now ends in the generated session title Reply pong, so Codex's parallel title request went through as well:
› Reply with just the word pong.
• pong
  done 4:03 AM
› Ask Codex to do anything
  gpt-5.4-mini medium · /private/tmp/claude-501/-Users-mateo-Development-litellm/13337e26-dc6e-4365-a37c-57d6c7feb0a9/scratchpad/lit7902/qa/after/codex_project · Reply pong

pr41953-5477dbe74c-lit7902-after-codex-pane.png

curl, the exact Codex request

  1. curl -sS -i http://127.0.0.1:$PORT/v1/responses -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d @codex_tools_body.json, run twice so each worker serves one
  2. Both runs HTTP/1.1 200 OK, "status":"completed", output_text pong, usage.total_tokens 1551 against 17 with no tools, so the 8 function and custom tools still reach the model

curl, Codex's session-title request

  1. curl -sS -i http://127.0.0.1:$PORT/v1/responses -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{"model":"gpt-5.4-mini","input":"Reply with just the word pong.","tools":[],"parallel_tool_calls":true}', run twice
  2. Both runs HTTP/1.1 200 OK, "status":"completed", output_text pong, usage.total_tokens 17

curl, one function tool with parallel_tool_calls

  1. curl -sS -i http://127.0.0.1:$PORT/v1/responses -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{"model":"gpt-5.4-mini","input":"Reply with just the word pong.","parallel_tool_calls":true,"tools":[{"type":"function","name":"get_weather","description":"Get weather","parameters":{"type":"object","properties":{"city":{"type":"string"}},"required":["city"]},"strict":false}]}'
  2. HTTP/1.1 200 OK, "status":"completed", output_text pong, usage.total_tokens 132, unchanged from Before

curl, no tools

  1. curl -sS -i http://127.0.0.1:$PORT/v1/responses -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{"model":"gpt-5.4-mini","input":"Reply with just the word pong."}'
  2. HTTP/1.1 200 OK, "status":"completed", output_text pong, usage.total_tokens 17, unchanged from Before

Observations from the run:

  • Bridge 200s echo tools: [] and parallel_tool_calls: false; pre-existing, unchanged here
  • Both workers covered by repeat runs; access log lacks pid

Type

🐛 Bug Fix

Caveats (if any)

Medium

  • tool_search is dropped, so Codex's deferred-tool discovery is unavailable through the bridge, the same treatment the bridge gives computer_use. The alternative considered was converting Codex's execution: client entry into a function tool named tool_search so the model could still call it; not done here because whether Codex accepts a plain function_call back for that tool is unverified and would need its own QA, while the drop matches how the bridge already treats every hosted tool without a Chat Completions equivalent

Low

  • Codex 0.155.1 migrates gpt-5.4-mini to gpt-5.6-luna on first launch, and that model sends tools as an additional_tools input item instead, a shape this PR does not touch
  • Only Azure was observed rejecting parallel_tool_calls without tools; OpenAI documents the same rule but was not re-tested here, and dropping the flag is harmless either way since it only means something alongside tools
  • Codex's web_search entry was already turned into web_search_options and dropped for providers that reject it (Azure included); this PR leaves that path as it was
  • code_interpreter from Responses API → Chat Completions transformation silently passes through unsupported built-in tool types #27655 and file_search still pass through verbatim (mcp passes through on purpose since Chat Completions accepts it); out of scope here

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Note

Low Risk
Scoped to the Responses-to-Chat-Completions request transform; reduces invalid upstream payloads without changing auth or core routing.

Overview
Fixes 400 errors when Codex (and similar clients) hit LiteLLM’s Responses → Chat Completions bridge with hosted tools or empty tool lists.

The bridge now treats tool_search and local_shell like other Responses-only tools (computer_use, shell, etc.): they are dropped with a warning instead of being forwarded to chat providers that only accept function / custom. parallel_tool_calls is removed from the outgoing completion request whenever no chat tools remain after that filtering (e.g. tools: [] or only dropped hosted tools), while it is still passed through when at least one convertible function/custom tool survives.

Tests cover empty/hosted-only vs function-tool cases and Codex-style mixed tool lists.

Reviewed by Cursor Bugbot for commit 5477dbe. Bugbot is set up for automated code reviews on this repo. Configure here.

…ons bridge

Hosted Responses API tools with no Chat Completions equivalent were forwarded
verbatim, so Codex 0.140+ got a 400 from the provider on every turn. The bridge
now drops tool_search and local_shell the same way it drops computer_use,
image_generation, and shell, and also drops parallel_tool_calls when no chat
tools remain, since chat completions only accepts it alongside tools
@devin-ai-integration

devin-ai-integration Bot commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@greptile-apps

greptile-apps Bot commented Sep 19, 2026

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

The PR appears safe to merge because unsupported hosted tools are removed while supported tool conversions and related options remain intact

Summary

This PR fixes the Responses-to-Chat Completions bridge by dropping unsupported tool_search and local_shell tools and removing parallel_tool_calls when no callable tools remain

  • Preserves supported function, custom, and web-search transformations
  • Adds regression coverage for hosted-tool filtering and parameter cleanup

Reviews (1) · Last reviewed commit: "fix(responses): drop tool_search and loc..."

@codspeed

codspeed Bot commented Sep 19, 2026

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_bridge_drop_tool_search (5477dbe) with main (db04e79)

Open in CodSpeed

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@codecov

codecov Bot commented Sep 19, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 5477dbe. Configure here.

@mateo-berri mateo-berri left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@mateo-berri
mateo-berri merged commit 40f51df into main Sep 19, 2026
91 of 92 checks passed
@mateo-berri
mateo-berri deleted the litellm_bridge_drop_tool_search branch September 19, 2026 11:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant