Skip to content

fix(responses): lift additional_tools input items into tools on the chat bridge - #38388

Closed
leonardofreitass wants to merge 1 commit into
BerriAI:litellm_internal_stagingfrom
leonardofreitass:fix/responses-hoist-codex-additional-tools
Closed

leonardofreitass wants to merge 1 commit into
BerriAI:litellm_internal_stagingfrom
leonardofreitass:fix/responses-hoist-codex-additional-tools

Conversation

@leonardofreitass

@leonardofreitass leonardofreitass commented Aug 26, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Tools nested in additional_tools input items never reach the provider
  • The Responses → Chat Completions conversion drops the item silently
  • Codex CLI sessions end up with an assistant that cannot call anything

How it solves it:

  • Lift the nested tools into the tool list before conversion
  • Let allowlist and guardrail extraction see them, so nesting is not a bypass
  • Scoped to the chat bridge, so native Responses paths are untouched

User Flow

Before: a developer running Codex CLI through a LiteLLM gateway on a Bedrock-backed model gets an assistant that cannot use a single tool, and nothing says why.

  1. They start Codex against POST https://litellm-domain/v1/responses with a Bedrock GPT-5.6 deployment
  2. They ask it to do something that needs a tool
  3. It replies in prose only — "I can't access terminal or file-reading tools in this session" — and never calls anything
  4. The request returns HTTP 200 with no error, and https://litellm-domain/ui/?page=logs shows an ordinary successful call

After: the same session uses its tools normally.

  1. They start Codex against the same POST https://litellm-domain/v1/responses with the same deployment
  2. They ask it to do something that needs a tool
  3. It issues a tool call and the session proceeds
  4. https://litellm-domain/ui/?page=logs shows the same request, now with the tool call recorded

For a key restricted by metadata.allowed_tools, a tool nested in input is now refused with tool_access_denied instead of silently reaching the model, and configured guardrails now see those tools instead of receiving none.

Relevant issues

Related: #29818, #36182

Adjacent but deliberately out of scope: #27276 (tool-name and tool-type handling inside the same conversion).

Linear ticket

Pre-Submission checklist

  • I have added meaningful tests
  • The handful of test files covering my change pass locally: uv run pytest tests/test_litellm/llms/openai/responses/test_openai_responses_guardrail_handler.py tests/test_litellm/proxy/test_tools_allowlist_enforcement.py tests/test_litellm/responses/litellm_completion_transformation/test_additional_tools_lifting.py -v (130 pass)
  • My PR passes all required CI/CD checks — make lint sub-checks verified locally: format-check-changed, ruff (litellm + tests config), ruff-strict ratchet, type-discipline ratchet, test-quality ratchet, circular-imports, import-safety
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5

Screenshots / Proof of Fix

End-to-end against real Amazon Bedrock through a local proxy. No mocks.

config.yaml — resolves to the converse route via the bundled price map (litellm_provider: bedrock_converse):

model_list:
  - model_name: global.gpt-5.6-sol
    litellm_params:
      model: bedrock/global.openai.gpt-5.6-sol
      aws_region_name: us-east-1
litellm_settings:
  drop_params: true

Request body: a Codex CLI 0.149 "responses lite" payload — top-level tools: [], with 9 tool definitions nested in an additional_tools input item across two namespace containers (functions, collaboration), parallel_tool_calls: false, reasoning.context: all_turns.

Two measurements are reported. Whether the model chooses to call a tool is not deterministic, so the authoritative signal is the tool count that reaches the outbound chat request.

Before (2e73400)

  1. uv run litellm --config config.yaml --port 4000
  2. POST http://127.0.0.1:4000/v1/responses with the body above → HTTP 200
  3. Observed tools on the outbound chat request: 0
  4. Observed output item types: ['message'] — no function_call
  5. Observed assistant text: "I can't access terminal or file-reading tools in this session, so I'm unable to read ./README.md."

After (29bb4e3)

  1. uv run litellm --config config.yaml --port 4000
  2. POST http://127.0.0.1:4000/v1/responses with the same body → HTTP 200
  3. Observed tools on the outbound chat request: 8, lifted from the 2 namespace containers — functions__wait, functions__request_user_input, collaboration__followup_task, collaboration__interrupt_agent, collaboration__list_agents, collaboration__send_message, collaboration__spawn_agent, collaboration__wait_agent
  4. Observed output item types: ['function_call', 'message']
  5. Observed tool call: collaboration__spawn_agent, and 3 function_call events across the turn

Type

🐛 Bug Fix

Caveats (if any)

Medium

  • Guardrails can now block a tool nested in additional_tools, but not redact one. Nested tools are sent for inspection and excluded from the write-back by identity, so a guardrail's edits to them are discarded. Strictly better than the previous behaviour, where tool extraction was skipped entirely whenever top-level tools was empty, but not complete.
  • The two surfaces disagree on tool name form: the allowlist sees spawn_agent, guardrails see collaboration__spawn_agent (the namespace prefix). The bare inner name was chosen for the allowlist because that is what an operator writes in allowed_tools. Worth a maintainer's opinion.
  • Native Responses providers that reject the item are unchanged; bedrock_mantle already lifts it in its own transformation, others (xai, perplexity, openrouter) would need a separate change.
  • bedrock_mantle still carries its own copy of the parse. Migrating it onto the shared helper is a follow-up: its tests couple to internals including a debug-log assertion, and it is not needed here.

Low

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

🤖 Generated with Claude Code

@codecov

codecov Bot commented Aug 26, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 98.14815% with 1 line in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
.../openai/responses/guardrail_translation/handler.py 92.85% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

@codspeed

codspeed Bot commented Aug 26, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing leonardofreitass:fix/responses-hoist-codex-additional-tools (29bb4e3) with litellm_internal_staging (dd01abc)1

Open in CodSpeed

Footnotes

  1. No successful run was found on litellm_internal_staging (4ad4db2) during the generation of this report, so dd01abc was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩

@leonardofreitass
leonardofreitass force-pushed the fix/responses-hoist-codex-additional-tools branch from f0a87cd to 54bd965 Compare August 26, 2026 21:31
@leonardofreitass
leonardofreitass marked this pull request as ready for review August 26, 2026 21:56
Comment thread litellm/responses/litellm_completion_transformation/transformation.py Outdated
@veria-ai

veria-ai Bot commented Aug 26, 2026 •

Copy link
Copy Markdown
Contributor

PR overview

This pull request updates the chat bridge to lift additional_tools input items into the request’s top-level tools collection.

One issue has been addressed, but nested tool filtering remains ineffective when a guardrail removes a tool. An attacker could preserve a rejected tool through the bridge and make it available for use, bypassing the intended guardrail with a limited, request-level blast radius.

Open issues (1)

Fixed/addressed: 1 · PR risk: 6/10

@greptile-apps

greptile-apps Bot commented Aug 26, 2026

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR updates the Responses-to-Chat bridge to extract tools from Codex-style additional_tools input items before message conversion.

  • Preserves ordinary input items while merging nested tools after top-level tools.
  • Adds regression coverage for lifting, ordering, malformed items, empty tools, and message preservation.

Confidence Score: 4/5

The PR appears safe to merge functionally, with a non-blocking repository-convention issue around duplicated parsing logic and test placement.

The changed bridge preserves surviving messages and sends lifted tools through the established converter; the remaining accepted concern is maintainability drift between parallel additional_tools implementations.

Files Needing Attention: litellm/responses/litellm_completion_transformation/transformation.py; tests/test_litellm/responses/litellm_completion_transformation/test_additional_tools_lifting.py

Important Files Changed

Filename Overview
litellm/responses/litellm_completion_transformation/transformation.py Correctly routes nested tools into the existing converter, but duplicates an existing wire-format parser rather than sharing it.
tests/test_litellm/responses/litellm_completion_transformation/test_additional_tools_lifting.py Provides meaningful bridge regression coverage, though repository guidance calls for adding bug-fix cases to the existing mapped test module.

Reviews (1): Last reviewed commit: "fix(responses): lift additional_tools in..." | Re-trigger Greptile

Comment thread litellm/responses/litellm_completion_transformation/transformation.py Outdated
@leonardofreitass
leonardofreitass force-pushed the fix/responses-hoist-codex-additional-tools branch from 54bd965 to aa44c80 Compare August 27, 2026 06:02
@leonardofreitass
leonardofreitass force-pushed the fix/responses-hoist-codex-additional-tools branch 2 times, most recently from 9b9e670 to 5d3dd54 Compare September 2, 2026 16:30
…hat bridge

An additional_tools input item carries tool definitions but no content, so the
Responses -> Chat Completions conversion dropped it silently and the model was
offered no tools at all. Lift the nested tools into the converted tool list.

Codex CLI emits this shape, so a Codex session against any provider without a
native Responses config lost its entire toolset with no error and no log line.

Lifting makes those tools live, so the /v1/responses allowlist and guardrail
extractor has to see them too; otherwise a nested tool reaches the model without
passing tool authorization. The bridge and the extractor now share one parser so
they cannot disagree about what the effective tool list is.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@leonardofreitass
leonardofreitass force-pushed the fix/responses-hoist-codex-additional-tools branch from 5d3dd54 to 29bb4e3 Compare September 4, 2026 19:34
merge_guardrailed_tools(
original_tools,
flattened_tool_groups,
tuple(t for t in guardrailed_tools if _guardrail_tool_identity(t) not in excluded),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Medium: Nested tool filtering is discarded

If a guardrail removes a nested tool from guardrailed_tools, this filter changes nothing in input; the original additional_tools item remains and the chat bridge later lifts that rejected tool. An attacker can therefore retain a tool that a filtering guardrail removed. Rebuild the nested items from the guardrail result, or remove them from input and hoist the guardrailed versions exactly once.

@mateo-berri

Copy link
Copy Markdown
Contributor

Closing as superseded by #40989, which carries the same additional_tools hoist on main plus the streaming and guardrail merge fixes. Thanks for the original fix

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants