Skip to content

Responses API Codex compatibility: custom tools, namespace tools, agent_message, array tool output - #35216

Open
chromecast56 wants to merge 5 commits into
sgl-project:mainfrom
chromecast56:responses-custom-tools
Open

chromecast56 wants to merge 5 commits into
sgl-project:mainfrom
chromecast56:responses-custom-tools

Conversation

@chromecast56

@chromecast56 chromecast56 commented Aug 17, 2026 •

Copy link
Copy Markdown
Collaborator

Motivation

The OpenAI Codex CLI now drives providers through /v1/responses, and several item and tool types it sends are either silently dropped or hard-rejected by the Responses serving layer. Each of these kills a Codex conversation in a different way:

  1. Custom (freeform) tools (apply_patch_tool_type = freeform), two separate failures: the declaration returns 200 but is silently dropped (tool.type != "function" skip), so the model never calls the tool; and independently, replaying a custom_tool_call / custom_tool_call_output item from any prior turn 400s, killing the conversation.
  2. Namespace tools: Codex groups its multi-agent tool family under a namespace declaration (ProviderCapabilities::namespace_tools defaults to true, so custom providers receive them un-flattened). The protocol type already carries ResponseTool.tools ("Inner schemas for namespace tools") but nothing handles it — the model never sees the inner tools, so it never calls them.
  3. agent_message items: any multi-agent thread carrying a sub-agent reply 400s with Unsupported Responses API input item type: 'agent_message' (same failure reported against other providers in Multi-Agent V2 sends OpenAI-specific agent_message items to external Responses providers openai/codex#33551).
  4. Array-shaped tool output: list outputs were flattened by string-joining text fields, silently dropping input_image parts — live for Codex because view_image returns image parts in tool output ([Bug] /v1/responses API fails with 400 when function_call_output output is an array #33867, [Bug] [Responses API] input_image parts in function_call_output are not converted to image_url: 400 ChatCompletionRequest validation error (post-#25881 builds), silently dropped on current main #34927).

Modifications

All changes are stateless HTTP-serving-layer translation in serving_responses.py; no tokenizer, template, or engine changes.

Custom (freeform) tools — declaration shim to a one-string-argument function ({"input": ...}), so parsers and constrained decoding keep working; custom_tool_call / custom_tool_call_output input items normalize to chat messages; parsed calls to a declared custom tool emit custom_tool_call output items carrying the raw string; streaming emits response.custom_tool_call_input.delta/done. Known limitation: format (lark/regex grammars) is ignored — grammar-constrained custom tools run unconstrained.

Namespace tools — each inner function tool flattens to a chat function named f"{namespace}.{inner}"; emitted call items split the pair back (name=inner plus a namespace field); replayed namespaced calls re-qualify to the dotted name for the chat template. Only declared namespace prefixes split — a plain function whose name contains a dot passes through. Verified live against a Kimi K3 deployment that the model faithfully echoes dotted tool names in its calls. Deliberate simplification vs. OpenAI: inner schemas are exposed up front rather than deferred behind tool search.

agent_message — renders as a user message: routing header ([agent message from X to Y]) plus the message text. Codex agent messages may contain plaintext input_text and/or opaque encrypted_content; plaintext is forwarded, and ciphertext-looking content (one long unbroken base64 token) is replaced with a visible placeholder because SGLang cannot decrypt it. Hosted multi_agent_call actions remain unsupported and keep 400ing — accepting them would claim orchestration work the server never performs.

Array-shaped tool output — text-only arrays still flatten to a plain string (the shape every chat template accepts); arrays containing non-text parts route through the existing content-part normalizer (input_text → text, input_image → image_url) and pass through as chat content parts, so images in tool output reach the model. Fixes #33867 / #34927.

Tests

Registered in base-a-test-cpu (not yet run by a maintainer-triggered suite):

  • test/registered/core/test_responses_custom_tools.py — declaration shim shape, call emission, input replay, unknown-type rejection preserved.
  • test/registered/core/test_responses_codex_compat.py — namespace flattening (description inheritance, strict preservation, qualified names), call-item split, undeclared-dotted-name passthrough, replay re-qualification; agent_message rendering, ciphertext placeholder, empty-message drop; text-only output flattening and input_image part survival for both output item types.
  • test/registered/unit/entrypoints/openai/test_serving_responses_stream.py — wire-level custom-tool streaming regression: output_item.added → unwrapped custom_tool_call_input.delta → .done → output_item.done → the custom_tool_call in response.completed.output; no function_call_arguments events leak for custom tools.

CI States

Latest PR Test (Base): ❌ Missing run-ci label -- add it to run CI tests.
Latest PR Test (Extra): ❌ Blocked -- run-ci is required first.
Latest PR Test (AMD ROCm 10): ➖ No AMD PR run found for this commit.

Declaring tools=[{"type": "custom"}] was accepted by validation but the
tool then did not exist in any useful sense:

- the declaration was silently dropped by _response_tools_to_chat_tools
  (only "function" tools flowed to chat), so the model never saw the tool
  and could not call it;
- replaying the custom_tool_call / custom_tool_call_output items the
  declaration invites failed with a hard 400 ("Unsupported Responses API
  input item type"), breaking any client that keeps a transcript — e.g.
  OpenAI Codex CLI with apply_patch_tool_type = freeform, which gets a
  clean 200 on turn 1 and an unavoidable 400 on turn 2.

Custom tools now round-trip: declarations are exposed to the chat layer as
a function with a single string property ("input"), so JSON tool-call
parsing and constrained decoding work unchanged; parsed calls to a
declared-custom tool are emitted as custom_tool_call items with the raw
string unwrapped (non-streaming and streaming, including
response.custom_tool_call_input.delta/done events); and replayed
custom_tool_call / custom_tool_call_output input items normalize to the
equivalent assistant tool_call / tool messages. Streaming buffers the
wrapped JSON and emits the input as one delta at close, since a JSON
string cannot be unwrapped incrementally. Unknown item types still 400.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…age, array tool output

- Namespace tools: flatten inner function tools to namespace.inner chat
  functions; emitted call items split back to name + namespace; replayed
  namespaced calls re-qualify for the chat template. Verified live that the
  model faithfully echoes dotted tool names.
- agent_message items render as text with a routing header instead of
  400ing (both input_text and encrypted_content parts are plaintext against
  non-OpenAI providers; ciphertext-looking blobs become a placeholder).
  Hosted multi_agent_call actions stay unsupported.
- Array-shaped tool output: text-only arrays still flatten to a string, but
  arrays with non-text parts (input_image from Codex view_image) pass
  through as normalized chat content parts instead of being dropped
  (fixes sgl-project#33867 / sgl-project#34927).
- CPU tests: test/registered/core/test_responses_codex_compat.py.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@chromecast56 chromecast56 changed the title Fix Responses API custom (freeform) tools: declaration, calls, and replay Responses API Codex compatibility: custom tools, namespace tools, agent_message, array tool output Aug 18, 2026
chromecast56 and others added 2 commits August 18, 2026 01:39
…egression, black, prose

- Namespace flattening now carries the inner function's strict flag
  (a valid strict:true declaration silently became non-strict).
- Wire-level streaming regression for custom tools: item added ->
  unwrapped input delta -> input done -> item done -> completed snapshot;
  custom tools leak no function_call argument deltas.
- Soften the agent_message encryption claim: encrypted_content may be
  plaintext-in-disguise or real ciphertext cross-provider; plaintext is
  forwarded, ciphertext-looking content becomes a visible placeholder.
- black 26.1.0 over touched files.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@gongy

gongy commented Aug 19, 2026

Copy link
Copy Markdown
Collaborator

/rerun-test test/registered/core/test_responses_codex_compat.py test/registered/core/test_responses_custom_tools.py test/registered/unit/entrypoints/openai/test_serving_responses_stream.py test/registered/unit/entrypoints/openai/test_serving_responses.py test/registered/unit/entrypoints/openai/test_responses_protocol.py

@github-actions

github-actions Bot commented Aug 19, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-test test/registered/core/test_responses_codex_compat.py test/registered/core/test_responses_custom_tools.py test/registered/unit/entrypoints/openai/test_serving_responses_stream.py test/registered/unit/entrypoints/openai/test_serving_responses.py test/registered/unit/entrypoints/openai/test_responses_protocol.py:

🚀 ubuntu-latest (5 tests): ✅ View workflow run

cd test/ && python3 registered/core/test_responses_codex_compat.py
cd test/ && python3 registered/core/test_responses_custom_tools.py
cd test/ && python3 registered/unit/entrypoints/openai/test_serving_responses_stream.py
cd test/ && python3 registered/unit/entrypoints/openai/test_serving_responses.py
cd test/ && python3 registered/unit/entrypoints/openai/test_responses_protocol.py

prakhar-prakash-juspay added a commit to juspay/litellm that referenced this pull request Aug 20, 2026
Codex code mode declares its tools as an `additional_tools` input item and
leaves the top-level `tools` array empty. That item is not part of the public
Responses schema, so OpenAI-compatible engines do not understand it:

  vLLM 0.25.x  -> 400 "cannot pickle 'pydantic_core...ValidatorIterator' object"
  SGLang       -> 400 "Unsupported Responses API input item type"
  vLLM 0.23.x/0.26.x -> 200, item silently ignored, model gets no tools at all

The last case is the dangerous one: the agent looks healthy and quietly stops
being able to call anything.

Hoist the inner tools (unwrapping `namespace` entries) into `tools`. Dropping
the item is not enough, since Codex sends `tools: []` and the model would be
left with nothing. Also rewrite `custom` (freeform) tools as functions taking a
single string argument, because engines skip any tool whose type is not
"function" - without that the `exec` tool disappears and the agent still cannot
run anything. This discards `format` (lark/regex grammar), so such tools are no
longer grammar-constrained; sgl-project/sglang#35216 accepts the same tradeoff.

Verified against real captured Codex traffic on a LiteLLM gateway:

  model          as Codex sends it   hoist only   hoist + shim
  open-fast      400 pickle crash    calls exec   calls exec
  open-large     no tool calls       no tool calls calls exec

BerriAI#33228 applies the same hoist for bedrock_mantle; the helper
here is provider-agnostic so other providers can reuse it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
prakhar-prakash-juspay added a commit to juspay/litellm that referenced this pull request Aug 20, 2026
* fix(hosted_vllm): hoist Codex additional_tools into top-level tools

Codex code mode declares its tools as an `additional_tools` input item and
leaves the top-level `tools` array empty. That item is not part of the public
Responses schema, so OpenAI-compatible engines do not understand it:

  vLLM 0.25.x  -> 400 "cannot pickle 'pydantic_core...ValidatorIterator' object"
  SGLang       -> 400 "Unsupported Responses API input item type"
  vLLM 0.23.x/0.26.x -> 200, item silently ignored, model gets no tools at all

The last case is the dangerous one: the agent looks healthy and quietly stops
being able to call anything.

Hoist the inner tools (unwrapping `namespace` entries) into `tools`. Dropping
the item is not enough, since Codex sends `tools: []` and the model would be
left with nothing. Also rewrite `custom` (freeform) tools as functions taking a
single string argument, because engines skip any tool whose type is not
"function" - without that the `exec` tool disappears and the agent still cannot
run anything. This discards `format` (lark/regex grammar), so such tools are no
longer grammar-constrained; sgl-project/sglang#35216 accepts the same tradeoff.

Verified against real captured Codex traffic on a LiteLLM gateway:

  model          as Codex sends it   hoist only   hoist + shim
  open-fast      400 pickle crash    calls exec   calls exec
  open-large     no tool calls       no tool calls calls exec

BerriAI#33228 applies the same hoist for bedrock_mantle; the helper
here is provider-agnostic so other providers can reuse it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(codex_compat): stop truncating tool descriptions, keep nameless tools

Review of the previous commit surfaced three defects in the hoisting helper:

1. Custom-tool descriptions were capped at 1024 chars. Codex ships its entire
   code-mode API surface in `exec`'s description - 14546 chars in captured
   traffic - so the cap destroyed 93% of the reference the model needs to write
   valid calls, while leaving `function` tools untouched. The cap was arbitrary;
   nothing downstream requires it. Descriptions now pass through verbatim.

2. Tools without a `name` were dropped, because the dedupe guard was the only
   path that appended. Built-ins such as {"type": "web_search"} and MCP entries
   carry no name, so hoisting silently removed those capabilities. They are now
   forwarded.

3. `_expand_namespace` unwrapped a single level. Codex nests a namespace per MCP
   server, and an inner `namespace` entry would have been appended verbatim into
   `tools` - re-introducing the exact tool type this helper exists to remove.
   Expansion is now recursive.

Name collisions across flattened namespaces still resolve first-wins, which is
lossy: namespaces exist so two providers can both expose e.g. `read`. That now
logs a warning instead of failing silently, and is covered by a test that
documents the behaviour.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
hawkli-1994 added a commit to hawkli-1994/sglang that referenced this pull request Sep 12, 2026
… dropping them

A `namespace` declaration groups inner function schemas (Codex sends its
multi-agent tool families this way; ProviderCapabilities::namespace_tools
defaults to true). The protocol accepts the declaration but nothing
executes it: the inner functions never reach the chat tools the model sees,
so the model never calls them, and a replayed call carrying a `namespace`
field loses its qualification.

Members now flatten to `{namespace}.{name}` chat function tools (inner
description and strict flags preserved; the namespace description fills
members that omit one), and calls translate back on every output path:

- non-streaming and the required-JSON fallback split the qualified name
  into `name` plus a `namespace` field on the item
- streaming `output_item.added/done` carry the split pair
- replayed namespaced calls re-qualify for the chat template; only
  declared namespace prefixes split, so a plain function whose name
  contains a dot is never reinterpreted
- the harmony developer message renders the flattened members, and
  harmony output parsing splits qualified recipients the same way

Namespaced items travel through widened pydantic unions
(`ResponseNamespacedFunctionToolCall` plus matching event subclasses);
without the widened arm leading, the SDK's `ResponseFunctionToolCall` arm
serializes the subclass and silently drops `namespace`.

`tool_choice="required"` accepts namespace declarations, and a `namespace`
field on a named choice re-qualifies it to the flattened name instead of
degrading to "auto".

Custom tools (the other half of sgl-project#35216) landed on main via sgl-project#38690; this
PR builds on that work and shares its adapters module.

Co-Authored-By: Claude Code <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug] /v1/responses API fails with 400 when function_call_output output is an array

2 participants