Repository navigation
Responses API Codex compatibility: custom tools, namespace tools, agent_message, array tool output - #35216
Open
chromecast56 wants to merge 5 commits into
Open
Responses API Codex compatibility: custom tools, namespace tools, agent_message, array tool output#35216chromecast56 wants to merge 5 commits into
chromecast56 wants to merge 5 commits into
Conversation
Declaring tools=[{"type": "custom"}] was accepted by validation but the
tool then did not exist in any useful sense:
- the declaration was silently dropped by _response_tools_to_chat_tools
(only "function" tools flowed to chat), so the model never saw the tool
and could not call it;
- replaying the custom_tool_call / custom_tool_call_output items the
declaration invites failed with a hard 400 ("Unsupported Responses API
input item type"), breaking any client that keeps a transcript — e.g.
OpenAI Codex CLI with apply_patch_tool_type = freeform, which gets a
clean 200 on turn 1 and an unavoidable 400 on turn 2.
Custom tools now round-trip: declarations are exposed to the chat layer as
a function with a single string property ("input"), so JSON tool-call
parsing and constrained decoding work unchanged; parsed calls to a
declared-custom tool are emitted as custom_tool_call items with the raw
string unwrapped (non-streaming and streaming, including
response.custom_tool_call_input.delta/done events); and replayed
custom_tool_call / custom_tool_call_output input items normalize to the
equivalent assistant tool_call / tool messages. Streaming buffers the
wrapped JSON and emits the input as one delta at close, since a JSON
string cannot be unwrapped incrementally. Unknown item types still 400.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…age, array tool output - Namespace tools: flatten inner function tools to namespace.inner chat functions; emitted call items split back to name + namespace; replayed namespaced calls re-qualify for the chat template. Verified live that the model faithfully echoes dotted tool names. - agent_message items render as text with a routing header instead of 400ing (both input_text and encrypted_content parts are plaintext against non-OpenAI providers; ciphertext-looking blobs become a placeholder). Hosted multi_agent_call actions stay unsupported. - Array-shaped tool output: text-only arrays still flatten to a string, but arrays with non-text parts (input_image from Codex view_image) pass through as normalized chat content parts instead of being dropped (fixes sgl-project#33867 / sgl-project#34927). - CPU tests: test/registered/core/test_responses_codex_compat.py. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…egression, black, prose - Namespace flattening now carries the inner function's strict flag (a valid strict:true declaration silently became non-strict). - Wire-level streaming regression for custom tools: item added -> unwrapped input delta -> input done -> item done -> completed snapshot; custom tools leak no function_call argument deltas. - Soften the agent_message encryption claim: encrypted_content may be plaintext-in-disguise or real ciphertext cross-provider; plaintext is forwarded, ciphertext-looking content becomes a visible placeholder. - black 26.1.0 over touched files. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Collaborator
|
/rerun-test test/registered/core/test_responses_codex_compat.py test/registered/core/test_responses_custom_tools.py test/registered/unit/entrypoints/openai/test_serving_responses_stream.py test/registered/unit/entrypoints/openai/test_serving_responses.py test/registered/unit/entrypoints/openai/test_responses_protocol.py |
Contributor
|
Results for 🚀 |
prakhar-prakash-juspay
added a commit
to juspay/litellm
that referenced
this pull request
Aug 20, 2026
Codex code mode declares its tools as an `additional_tools` input item and leaves the top-level `tools` array empty. That item is not part of the public Responses schema, so OpenAI-compatible engines do not understand it: vLLM 0.25.x -> 400 "cannot pickle 'pydantic_core...ValidatorIterator' object" SGLang -> 400 "Unsupported Responses API input item type" vLLM 0.23.x/0.26.x -> 200, item silently ignored, model gets no tools at all The last case is the dangerous one: the agent looks healthy and quietly stops being able to call anything. Hoist the inner tools (unwrapping `namespace` entries) into `tools`. Dropping the item is not enough, since Codex sends `tools: []` and the model would be left with nothing. Also rewrite `custom` (freeform) tools as functions taking a single string argument, because engines skip any tool whose type is not "function" - without that the `exec` tool disappears and the agent still cannot run anything. This discards `format` (lark/regex grammar), so such tools are no longer grammar-constrained; sgl-project/sglang#35216 accepts the same tradeoff. Verified against real captured Codex traffic on a LiteLLM gateway: model as Codex sends it hoist only hoist + shim open-fast 400 pickle crash calls exec calls exec open-large no tool calls no tool calls calls exec BerriAI#33228 applies the same hoist for bedrock_mantle; the helper here is provider-agnostic so other providers can reuse it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
prakhar-prakash-juspay
added a commit
to juspay/litellm
that referenced
this pull request
Aug 20, 2026
* fix(hosted_vllm): hoist Codex additional_tools into top-level tools Codex code mode declares its tools as an `additional_tools` input item and leaves the top-level `tools` array empty. That item is not part of the public Responses schema, so OpenAI-compatible engines do not understand it: vLLM 0.25.x -> 400 "cannot pickle 'pydantic_core...ValidatorIterator' object" SGLang -> 400 "Unsupported Responses API input item type" vLLM 0.23.x/0.26.x -> 200, item silently ignored, model gets no tools at all The last case is the dangerous one: the agent looks healthy and quietly stops being able to call anything. Hoist the inner tools (unwrapping `namespace` entries) into `tools`. Dropping the item is not enough, since Codex sends `tools: []` and the model would be left with nothing. Also rewrite `custom` (freeform) tools as functions taking a single string argument, because engines skip any tool whose type is not "function" - without that the `exec` tool disappears and the agent still cannot run anything. This discards `format` (lark/regex grammar), so such tools are no longer grammar-constrained; sgl-project/sglang#35216 accepts the same tradeoff. Verified against real captured Codex traffic on a LiteLLM gateway: model as Codex sends it hoist only hoist + shim open-fast 400 pickle crash calls exec calls exec open-large no tool calls no tool calls calls exec BerriAI#33228 applies the same hoist for bedrock_mantle; the helper here is provider-agnostic so other providers can reuse it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(codex_compat): stop truncating tool descriptions, keep nameless tools Review of the previous commit surfaced three defects in the hoisting helper: 1. Custom-tool descriptions were capped at 1024 chars. Codex ships its entire code-mode API surface in `exec`'s description - 14546 chars in captured traffic - so the cap destroyed 93% of the reference the model needs to write valid calls, while leaving `function` tools untouched. The cap was arbitrary; nothing downstream requires it. Descriptions now pass through verbatim. 2. Tools without a `name` were dropped, because the dedupe guard was the only path that appended. Built-ins such as {"type": "web_search"} and MCP entries carry no name, so hoisting silently removed those capabilities. They are now forwarded. 3. `_expand_namespace` unwrapped a single level. Codex nests a namespace per MCP server, and an inner `namespace` entry would have been appended verbatim into `tools` - re-introducing the exact tool type this helper exists to remove. Expansion is now recursive. Name collisions across flattened namespaces still resolve first-wins, which is lossy: namespaces exist so two providers can both expose e.g. `read`. That now logs a warning instead of failing silently, and is covered by a test that documents the behaviour. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
4 of 5 tasks
hawkli-1994
added a commit
to hawkli-1994/sglang
that referenced
this pull request
Sep 12, 2026
… dropping them
A `namespace` declaration groups inner function schemas (Codex sends its
multi-agent tool families this way; ProviderCapabilities::namespace_tools
defaults to true). The protocol accepts the declaration but nothing
executes it: the inner functions never reach the chat tools the model sees,
so the model never calls them, and a replayed call carrying a `namespace`
field loses its qualification.
Members now flatten to `{namespace}.{name}` chat function tools (inner
description and strict flags preserved; the namespace description fills
members that omit one), and calls translate back on every output path:
- non-streaming and the required-JSON fallback split the qualified name
into `name` plus a `namespace` field on the item
- streaming `output_item.added/done` carry the split pair
- replayed namespaced calls re-qualify for the chat template; only
declared namespace prefixes split, so a plain function whose name
contains a dot is never reinterpreted
- the harmony developer message renders the flattened members, and
harmony output parsing splits qualified recipients the same way
Namespaced items travel through widened pydantic unions
(`ResponseNamespacedFunctionToolCall` plus matching event subclasses);
without the widened arm leading, the SDK's `ResponseFunctionToolCall` arm
serializes the subclass and silently drops `namespace`.
`tool_choice="required"` accepts namespace declarations, and a `namespace`
field on a named choice re-qualifies it to the flattened name instead of
degrading to "auto".
Custom tools (the other half of sgl-project#35216) landed on main via sgl-project#38690; this
PR builds on that work and shares its adapters module.
Co-Authored-By: Claude Code <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
The OpenAI Codex CLI now drives providers through
/v1/responses, and several item and tool types it sends are either silently dropped or hard-rejected by the Responses serving layer. Each of these kills a Codex conversation in a different way:apply_patch_tool_type = freeform), two separate failures: the declaration returns 200 but is silently dropped (tool.type != "function"skip), so the model never calls the tool; and independently, replaying acustom_tool_call/custom_tool_call_outputitem from any prior turn 400s, killing the conversation.namespacedeclaration (ProviderCapabilities::namespace_toolsdefaults totrue, so custom providers receive them un-flattened). The protocol type already carriesResponseTool.tools("Inner schemas fornamespacetools") but nothing handles it — the model never sees the inner tools, so it never calls them.agent_messageitems: any multi-agent thread carrying a sub-agent reply 400s withUnsupported Responses API input item type: 'agent_message'(same failure reported against other providers in Multi-Agent V2 sends OpenAI-specific agent_message items to external Responses providers openai/codex#33551).textfields, silently droppinginput_imageparts — live for Codex becauseview_imagereturns image parts in tool output ([Bug] /v1/responses API fails with 400 when function_call_output output is an array #33867, [Bug] [Responses API] input_image parts in function_call_output are not converted to image_url: 400 ChatCompletionRequest validation error (post-#25881 builds), silently dropped on current main #34927).Modifications
All changes are stateless HTTP-serving-layer translation in
serving_responses.py; no tokenizer, template, or engine changes.Custom (freeform) tools — declaration shim to a one-string-argument function (
{"input": ...}), so parsers and constrained decoding keep working;custom_tool_call/custom_tool_call_outputinput items normalize to chat messages; parsed calls to a declared custom tool emitcustom_tool_calloutput items carrying the raw string; streaming emitsresponse.custom_tool_call_input.delta/done. Known limitation:format(lark/regex grammars) is ignored — grammar-constrained custom tools run unconstrained.Namespace tools — each inner function tool flattens to a chat function named
f"{namespace}.{inner}"; emitted call items split the pair back (name=innerplus anamespacefield); replayed namespaced calls re-qualify to the dotted name for the chat template. Only declared namespace prefixes split — a plain function whose name contains a dot passes through. Verified live against a Kimi K3 deployment that the model faithfully echoes dotted tool names in its calls. Deliberate simplification vs. OpenAI: inner schemas are exposed up front rather than deferred behind tool search.agent_message— renders as a user message: routing header ([agent message from X to Y]) plus the message text. Codex agent messages may contain plaintextinput_textand/or opaqueencrypted_content; plaintext is forwarded, and ciphertext-looking content (one long unbroken base64 token) is replaced with a visible placeholder because SGLang cannot decrypt it. Hostedmulti_agent_callactions remain unsupported and keep 400ing — accepting them would claim orchestration work the server never performs.Array-shaped tool output — text-only arrays still flatten to a plain string (the shape every chat template accepts); arrays containing non-text parts route through the existing content-part normalizer (
input_text→text,input_image→image_url) and pass through as chat content parts, so images in tool output reach the model. Fixes #33867 / #34927.Tests
Registered in
base-a-test-cpu(not yet run by a maintainer-triggered suite):test/registered/core/test_responses_custom_tools.py— declaration shim shape, call emission, input replay, unknown-type rejection preserved.test/registered/core/test_responses_codex_compat.py— namespace flattening (description inheritance,strictpreservation, qualified names), call-item split, undeclared-dotted-name passthrough, replay re-qualification;agent_messagerendering, ciphertext placeholder, empty-message drop; text-only output flattening andinput_imagepart survival for both output item types.test/registered/unit/entrypoints/openai/test_serving_responses_stream.py— wire-level custom-tool streaming regression:output_item.added→ unwrappedcustom_tool_call_input.delta→.done→output_item.done→ thecustom_tool_callinresponse.completed.output; nofunction_call_argumentsevents leak for custom tools.CI States
Latest PR Test (Base): ❌ Missing
run-cilabel -- add it to run CI tests.Latest PR Test (Extra): ❌ Blocked --
run-ciis required first.Latest PR Test (AMD ROCm 10): ➖ No AMD PR run found for this commit.