Skip to content

fix(bedrock_mantle): hoist Codex additional_tools input items to top-level tools - #33228

Merged
mateo-berri merged 2 commits into
BerriAI:litellm_internal_stagingfrom
lyb0307:litellm_bedrock_mantle_codex_additional_tools
Jul 21, 2026
Merged

mateo-berri merged 2 commits into
BerriAI:litellm_internal_stagingfrom
lyb0307:litellm_bedrock_mantle_codex_additional_tools

Conversation

@lyb0307

@lyb0307 lyb0307 commented Jul 14, 2026

Copy link
Copy Markdown

Relevant issues

Linear ticket

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • My PR passes all CI/CD checks (e.g., lint, format, unit tests)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Setup: proxy on localhost:4000 hitting the real Bedrock Mantle endpoint in us-east-2 (SigV4 from the ambient AWS credentials, real $). gpt-5.6 is not in the price map yet, so mode and wire path are declared via model_info, the same way a proxy admin sets them in the Admin UI

model_list:
  - model_name: gpt-5.6-sol
    litellm_params:
      model: bedrock_mantle/openai.gpt-5.6-sol
      aws_region_name: us-east-2
    model_info:
      mode: responses
      use_openai_responses_path: true

general_settings:
  master_key: sk-1234

Before, at the PR base 10d5804, a Responses request shaped the way Codex CLI 0.144.x sends it (tool definitions inside input as an additional_tools item) fails:

curl -sS http://localhost:4000/v1/responses \
  -H "Authorization: Bearer sk-1234" -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.6-sol",
    "input": [
      {"type": "additional_tools", "role": "developer", "tools": [
        {"type": "function", "name": "wait", "description": "Waits on a yielded exec cell.", "strict": false,
         "parameters": {"type": "object", "properties": {"cell_id": {"type": "string"}}, "required": ["cell_id"], "additionalProperties": false}}
      ]},
      {"type": "message", "role": "user", "content": [{"type": "input_text", "text": "Say hi in one word."}]}
    ]
  }' -w "\nHTTP %{http_code}\n"
{"error":{"message":"litellm.APIConnectionError: Bedrock_mantleException - {\"error\":{\"code\":\"validation_error\",\"message\":\"invalid request body: Invalid 'input': value did not match any expected variant\",\"param\":null,\"type\":\"invalid_request_error\"}}. Received Model Group=gpt-5.6-sol\nAvailable Model Group Fallbacks=None","type":null,"param":null,"code":"500"}}
HTTP 500

After, at 20398d9, the exact same curl succeeds:

HTTP 200
status: completed | error: None
assistant said: ['Hi!']
usage: 56

End to end with the real client at 20398d9: Codex CLI 0.144.4 pointed at the proxy (wire_api = "responses", base_url = "http://127.0.0.1:4000/v1", model gpt-5.6-sol) completes a full agentic task. The run made 5 /v1/responses calls (the first turn plus tool-call turns that replay reasoning, custom_tool_call and function_call_output items) and every one returned 200

$ codex exec --skip-git-repo-check --sandbox workspace-write "Create a file named proof.txt containing exactly 'mantle fix works', then run 'wc -c proof.txt' and report the byte count."
...
Created `proof.txt` exactly as requested. Byte count: **16**.
diff --git a/proof.txt b/proof.txt
new file mode 100644
--- /dev/null
+++ b/proof.txt
@@ -0,0 +1 @@
+mantle fix works

tokens used
10,003

$ cat proof.txt
mantle fix works

$ grep -o 'POST /v1/responses HTTP/1.1" [0-9]*' litellm.log | sort | uniq -c
   5 POST /v1/responses HTTP/1.1" 200

Before the fix, the same Codex CLI setup failed on the first request of every session with the HTTP 500 above and retried forever

Type

🐛 Bug Fix

Changes

Codex CLI 0.144.x uses a "responses lite" wire format for the gpt-5.5 and gpt-5.6 model families: instructions move into developer messages and the tool definitions move out of the top-level tools param into input as a {"type": "additional_tools", "role": "developer", "tools": [...]} item. api.openai.com accepts that item type; Bedrock Mantle's request parser rejects the whole request with 400 invalid request body: Invalid 'input': value did not match any expected variant, which litellm surfaces as a 500 APIConnectionError. The result is that Codex CLI cannot talk to any bedrock_mantle/ Responses model through the proxy

Testing against the live endpoint (bedrock-mantle.us-east-2.api.aws, openai.gpt-5.6-sol) shows Mantle accepts the very same tool definitions when they are sent in the top-level tools param. BedrockMantleResponsesAPIConfig.transform_responses_api_request therefore now strips additional_tools items from input and merges their tools into the top-level tools param, running them through the existing unsupported-tool-type filter. Requests that carry no such item are passed through unchanged, and the item is stripped even when none of its tools survive the filter, since Mantle would otherwise 400 on the item itself

Unit tests cover the hoist (item stripped, tools merged after any existing tools, multiple items merged in order), the filter interaction, the malformed-item and string-input edge cases, and a pass-through case built from the item types Codex replays on later turns

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

@CLAassistant

CLAassistant commented Jul 14, 2026 •

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you all sign our Contributor License Agreement before we can accept your contribution.
1 out of 2 committers have signed the CLA.

✅ mateo-berri
❌ BingBingL
You have signed the CLA already but the status is still pending? Let us recheck it.

@greptile-apps

greptile-apps Bot commented Jul 14, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR fixes a Bedrock Mantle Responses API incompatibility where Codex CLI's "responses lite" wire format places tool definitions inside input as {"type": "additional_tools", ...} items rather than the top-level tools param — a format OpenAI accepts but Mantle rejects with a 400.

  • Adds transform_responses_api_request to BedrockMantleResponsesAPIConfig that strips additional_tools items from input, merges their tools (after applying the existing unsupported-type filter) into the top-level tools param, then delegates to the parent implementation.
  • Adds _hoist_codex_additional_tools, _is_codex_additional_tools_item, and _tools_of_additional_tools_item helpers; all edge cases (string input, no matching items, all tools filtered, multiple items, malformed item) are handled and covered by unit tests.

Confidence Score: 5/5

Safe to merge — the change is additive, scoped entirely to the Bedrock Mantle responses path, and does not touch any shared or critical infrastructure.

The override intercepts only requests destined for Bedrock Mantle's Responses API and is fully isolated from all other providers. Requests that contain no additional_tools items are passed through unchanged (confirmed by the passthrough test). All edge cases — empty filter results, malformed items, string inputs, multiple items — are explicitly tested with unit tests that make no real network calls. The tool-filter logic reuses the existing _filter_unsupported_tools helper, keeping behavior consistent with the existing map_openai_params path.

No files require special attention.

Important Files Changed

Filename Overview
litellm/llms/bedrock_mantle/responses/transformation.py Adds transform_responses_api_request override that hoists additional_tools input items to the top-level tools param before forwarding to the parent implementation; logic is clean and handles all edge cases correctly
tests/test_litellm/llms/bedrock_mantle/test_bedrock_mantle_responses_transformation.py Adds 9 new unit tests covering the hoist, merge ordering, filter interaction, empty-filter strip, malformed item, string input passthrough, agentic replay passthrough, and debug logging; all tests use mocked/local calls with no real network requests

Reviews (2): Last reviewed commit: "fix(bedrock_mantle): log additional_tool..." | Re-trigger Greptile

@codecov

codecov Bot commented Jul 14, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

Comment thread litellm/llms/bedrock_mantle/responses/transformation.py
@veria-ai

veria-ai Bot commented Jul 14, 2026 •

Copy link
Copy Markdown
Contributor

PR overview

All previously flagged issues have been addressed. No open security concerns remain on this pull request.

Security review

No open security issues remain on this pull request.

Fixed/addressed: 1 · PR risk: 0/10

@lzy7071

lzy7071 commented Jul 14, 2026 •

Copy link
Copy Markdown

Thank you! I am experiencing the same issue as well. Although I am not entirely certain if this is the "best solution" because it seems to be a deliberate design choice codex has made.

For additional evidence that this is not on the bedrock mantle side - using amazon-bed rock directly, I am able to use codex with amazon-bedrock set as provider directly.

@mateo-berri
mateo-berri changed the base branch from litellm_oss_daily_2026_07_13 to litellm_internal_staging July 21, 2026 02:30
@mateo-berri
mateo-berri force-pushed the litellm_bedrock_mantle_codex_additional_tools branch from 20398d9 to c72ae12 Compare July 21, 2026 02:30
@mateo-berri

mateo-berri commented Jul 21, 2026 •

Copy link
Copy Markdown
Contributor

Thanks for the fix and the thorough live proof. To get this mergeable I retargeted the base to litellm_internal_staging (the default base for OSS contributions), rebased your branch onto its current tip, and pushed one commit on top that logs the hoist at debug level so the request rewrite is visible when diagnosing proxy traffic. Please also sign the CLA (see the cla-assistant comment above) so we can merge

One heads up: a separate change will gate the unsupported service_tier param that Codex sends when a speed tier is selected in its config, which is the other half of what breaks Codex against Mantle. Your PR stays scoped to the input item hoist

@codspeed

codspeed Bot commented Jul 21, 2026

Copy link
Copy Markdown
Contributor

Merging this PR will improve performance by 31.47%

⚠️ Different runtime environments detected

Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.

Open the report in CodSpeed to investigate

⚡ 4 improved benchmarks
✅ 27 untouched benchmarks

Performance Changes

Benchmark BASE HEAD Efficiency
⚡ test_completion_streaming 58.8 ms 34.6 ms +70%
⚡ test_response_to_model_response_object 543.5 µs 410.7 µs +32.34%
⚡ test_cost_per_token_openai 597.7 µs 517.7 µs +15.46%
⚡ test_cost_per_token_anthropic 602.6 µs 524 µs +15%

Tip

Curious why this is faster? Comment @codspeedbot explain why this is faster on this PR, or directly use the CodSpeed MCP with your agent.


Comparing lyb0307:litellm_bedrock_mantle_codex_additional_tools (8745f35) with litellm_internal_staging (214945a)1

Open in CodSpeed

Footnotes

  1. No successful run was found on litellm_internal_staging (9ad8698) during the generation of this report, so 214945a was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩

@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri mateo-berri left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM; thanks!

@mateo-berri
mateo-berri merged commit 07a355e into BerriAI:litellm_internal_staging Jul 21, 2026
77 checks passed
mateo-berri added a commit that referenced this pull request Jul 21, 2026
prakhar-prakash-juspay added a commit to juspay/litellm that referenced this pull request Aug 20, 2026
Codex code mode declares its tools as an `additional_tools` input item and
leaves the top-level `tools` array empty. That item is not part of the public
Responses schema, so OpenAI-compatible engines do not understand it:

  vLLM 0.25.x  -> 400 "cannot pickle 'pydantic_core...ValidatorIterator' object"
  SGLang       -> 400 "Unsupported Responses API input item type"
  vLLM 0.23.x/0.26.x -> 200, item silently ignored, model gets no tools at all

The last case is the dangerous one: the agent looks healthy and quietly stops
being able to call anything.

Hoist the inner tools (unwrapping `namespace` entries) into `tools`. Dropping
the item is not enough, since Codex sends `tools: []` and the model would be
left with nothing. Also rewrite `custom` (freeform) tools as functions taking a
single string argument, because engines skip any tool whose type is not
"function" - without that the `exec` tool disappears and the agent still cannot
run anything. This discards `format` (lark/regex grammar), so such tools are no
longer grammar-constrained; sgl-project/sglang#35216 accepts the same tradeoff.

Verified against real captured Codex traffic on a LiteLLM gateway:

  model          as Codex sends it   hoist only   hoist + shim
  open-fast      400 pickle crash    calls exec   calls exec
  open-large     no tool calls       no tool calls calls exec

BerriAI#33228 applies the same hoist for bedrock_mantle; the helper
here is provider-agnostic so other providers can reuse it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
prakhar-prakash-juspay added a commit to juspay/litellm that referenced this pull request Aug 20, 2026
* fix(hosted_vllm): hoist Codex additional_tools into top-level tools

Codex code mode declares its tools as an `additional_tools` input item and
leaves the top-level `tools` array empty. That item is not part of the public
Responses schema, so OpenAI-compatible engines do not understand it:

  vLLM 0.25.x  -> 400 "cannot pickle 'pydantic_core...ValidatorIterator' object"
  SGLang       -> 400 "Unsupported Responses API input item type"
  vLLM 0.23.x/0.26.x -> 200, item silently ignored, model gets no tools at all

The last case is the dangerous one: the agent looks healthy and quietly stops
being able to call anything.

Hoist the inner tools (unwrapping `namespace` entries) into `tools`. Dropping
the item is not enough, since Codex sends `tools: []` and the model would be
left with nothing. Also rewrite `custom` (freeform) tools as functions taking a
single string argument, because engines skip any tool whose type is not
"function" - without that the `exec` tool disappears and the agent still cannot
run anything. This discards `format` (lark/regex grammar), so such tools are no
longer grammar-constrained; sgl-project/sglang#35216 accepts the same tradeoff.

Verified against real captured Codex traffic on a LiteLLM gateway:

  model          as Codex sends it   hoist only   hoist + shim
  open-fast      400 pickle crash    calls exec   calls exec
  open-large     no tool calls       no tool calls calls exec

BerriAI#33228 applies the same hoist for bedrock_mantle; the helper
here is provider-agnostic so other providers can reuse it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(codex_compat): stop truncating tool descriptions, keep nameless tools

Review of the previous commit surfaced three defects in the hoisting helper:

1. Custom-tool descriptions were capped at 1024 chars. Codex ships its entire
   code-mode API surface in `exec`'s description - 14546 chars in captured
   traffic - so the cap destroyed 93% of the reference the model needs to write
   valid calls, while leaving `function` tools untouched. The cap was arbitrary;
   nothing downstream requires it. Descriptions now pass through verbatim.

2. Tools without a `name` were dropped, because the dedupe guard was the only
   path that appended. Built-ins such as {"type": "web_search"} and MCP entries
   carry no name, so hoisting silently removed those capabilities. They are now
   forwarded.

3. `_expand_namespace` unwrapped a single level. Codex nests a namespace per MCP
   server, and an inner `namespace` entry would have been appended verbatim into
   `tools` - re-introducing the exact tool type this helper exists to remove.
   Expansion is now recursive.

Name collisions across flattened namespaces still resolve first-wins, which is
lossy: namespaces exist so two providers can both expose e.g. `read`. That now
logs a warning instead of failing silently, and is covered by a test that
documents the behaviour.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants