Skip to content

fix: route codex payloads around the SDK's GIL-holding request transform (#93650) - #93773

Merged
kshitijk4poor merged 3 commits into
NousResearch:mainfrom
kshitijk4poor:fix/codex-sdk-transform-bypass-93650-v2
Aug 24, 2026
Merged

kshitijk4poor merged 3 commits into
NousResearch:mainfrom
kshitijk4poor:fix/codex-sdk-transform-bypass-93650-v2

Conversation

@kshitijk4poor

Copy link
Copy Markdown

Fixes #93650. Salvage of #93681 (@kchernev).

Summary

Routes the bulk payload fields (input, tools) through extra_body instead of the typed SDK kwargs, bypassing the OpenAI SDK's maybe_transform walk that can wedge for hours holding the GIL on large Codex Responses API calls.

Problem

A codex_responses streaming call can freeze the entire agent for hours before a single byte leaves the process. The OpenAI SDK's responses.create re-walks the whole request body against the ResponseCreateParams union graph (openai/_utils/_transform.py::maybe_transform) client-side, holding the GIL. In the incident documented in #93650 (openai 2.24.0, CPython 3.11.16), one such walk over a ~1.4 MB / 1,093-item conversation wedged for 12.5 CPU-hours at ~99% of a core.

Core-dump forensics showed the worker in maybe_transform → isinstance(obj, Mapping) → typing.__subclasscheck__, with 9 other threads in take_gil futex waits — including the TTFB/stale watchdog poller, which never got to run any of its detectors. No in-process watchdog can rescue this: the GIL-holding walk starves them, and even when they run their remedy is socket-level (_close_request_client_once) — a pre-network hang has no socket to kill.

Fix

Hermes assembles Codex payloads from JSON round-trips — they are already wire format, so the transform walk is pure overhead plus this failure mode. The SDK merges extra_body into the JSON body after the transform (_base_client._build_request, shallow merge), a mechanism this codebase already relies on for prompt_cache_retention. This PR routes the bulk payload fields (input, tools) through extra_body, skipping maybe_transform for them entirely while producing an identical request body.

Safety rails:

  • Plain-JSON guard — a field is only moved if every node is a plain JSON type (dict/list/str/int/float/bool/None with str keys). Non-JSON types keep the typed SDK path.
  • Caller precedence — an explicit caller-provided extra_body entry wins over a moved field.
  • Escape hatch — HERMES_CODEX_SDK_TRANSFORM=1 restores the pre-fix behavior.

Applied at both live responses.create call sites: the primary stream path (codex_runtime._open_codex_stream) and the auxiliary adapter (auxiliary_client._CodexCompletionsAdapter).

Changes

  • agent/codex_runtime.py: +74 lines — _is_plain_json_data, _bypass_sdk_request_transform, wired into _open_codex_stream
  • agent/auxiliary_client.py: +7/-1 — wired into _CodexCompletionsAdapter.create
  • tests/run_agent/test_codex_sdk_transform_bypass.py: +153 lines — 10 new tests
  • tests/agent/test_auxiliary_client.py: +14/-1 — fake-client capture updated for both wire shapes
  • tests/run_agent/test_native_compaction.py: +5/-1 — fake-client capture updated
  • tests/run_agent/test_run_agent_codex_responses.py: +6/-1 — retention guard updated for new wire shape
  • contributors/emails/: AUTHOR_MAP entry for @kchernev

Validation

Before After
Targeted tests (codex + aux + compaction) — 245 passed
New bypass tests — 10 passed
E2E (real imports, _bypass_sdk_request_transform) — 6/6 passed
Ruff check — clean

Competitor research

  • LiteLLM: bypasses the SDK entirely (builds raw httpx requests), so maybe_transform never runs
  • Codex CLI (Rust): builds JSON requests directly, structurally immune
  • LangChain: uses chat.completions.create (smaller typedict graph, not the massive ResponseCreateParams union)
  • Claude Code: routes through Anthropic's native API, not the OpenAI SDK

No other major harness uses the OpenAI Python SDK's responses.create() with payloads as large as Hermes does (558K tokens, 1.4MB body). Hermes is uniquely exposed.

kchernev and others added 3 commits August 24, 2026 15:29
…orm (NousResearch#93650)

responses.create re-walks the entire request body against the
ResponseCreateParams union graph client-side while holding the GIL.
NousResearch#93650 documents that walk wedging for 12+ hours on a ~1.4 MB
conversation, starving every other thread including the TTFB/stale
watchdogs — and no socket kill can unblock a pre-network hang.

Hermes payloads are JSON round-trips and already wire format, so the
bulk fields (input, tools) are now routed through extra_body, which the
SDK merges into the JSON body after the transform. Guarded by a
plain-JSON check (anything else keeps the typed path) and a
HERMES_CODEX_SDK_TRANSFORM=1 escape hatch. Applied to both the primary
stream path and the auxiliary adapter.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@kshitijk4poor kshitijk4poor added type/bug Something isn't working comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint provider/openai OpenAI / Codex Responses API P1 High — major feature broken, no workaround labels Aug 24, 2026
@kshitijk4poor
kshitijk4poor enabled auto-merge (rebase) August 24, 2026 10:32
auto-merge was automatically disabled August 24, 2026 10:32

Rebase failed

@kshitijk4poor
kshitijk4poor merged commit 6e534df into NousResearch:main Aug 24, 2026
39 checks passed
melon-xf added a commit to melon-xf/hermes-agent that referenced this pull request Sep 3, 2026
…k-transform-bypass-93650-v2

fix: route codex payloads around the SDK's GIL-holding request transform (NousResearch#93650)
kshitijk4poor pushed a commit to kshitijk4poor/hermes-agent that referenced this pull request Sep 9, 2026
…st transform

`chat.completions.create` re-walks the whole request body against the
`CompletionCreateParams` union graph client-side, with the GIL held, before
any byte leaves the process. NousResearch#93650 documented that class of walk wedging
for 12+ hours on a ~1.4 MB conversation: no in-process watchdog can fire
while the GIL is held, and no socket kill helps a pre-network hang.
NousResearch#93773 merged the remedy — route the already-wire-format bulk fields
through `extra_body`, which the SDK merges into the JSON body after the
transform — but scoped it to `responses.create`. The default chat path,
which every OpenRouter / Nous / xAI / DeepSeek / Kimi / llama.cpp /
Ollama / LM Studio / LiteLLM request takes, still pays the full walk.

Measured against a real `openai.OpenAI` over an `httpx.MockTransport`
(canned SSE, no network), with the request body captured from the
transport on both sides:

    101 msgs /  76 KB   13.6 ms -> 1.2 ms
    401 msgs / 190 KB   48.6 ms -> 2.0 ms
   1601 msgs / 650 KB  188.7 ms -> 5.8 ms

and the bytes the server receives are IDENTICAL — literally equal, not
merely equivalent (194,894 == 194,894 at 401 messages). The cost is paid
per API call, so a tool-using turn multiplies it by its iteration count.

The three helpers move from agent/codex_runtime.py into a shared
agent/sdk_transform_bypass.py, re-exported under their original names so
agent/auxiliary_client.py and tests/run_agent/test_codex_sdk_transform_bypass.py
keep working untouched. The field tuple is now a parameter:
("input", "tools") for Responses, ("messages", "tools") for chat.

Two chat-specific details. `messages` is a @required_args parameter, so it
stays in the typed kwargs as an empty list and the extra_body copy
replaces it in the body — hence the new `required_empty` argument, which
Responses does not use. And the bypass is gated on the target actually
being the SDK's Completions: Hermes also drives chat-completions-shaped
facades that are NOT the SDK — the in-process MoA aggregator most
importantly — and those never merge extra_body, so handing them one would
silently send an empty message list. That guard is also why this needs no
edits to the 32 test files that assert on kwargs["messages"]: they mock
with stand-ins, not the SDK.

Every rail the merged PR was reviewed on is kept: the plain-JSON-only
guard so pydantic models and generators stay on the typed path, caller
`extra_body` precedence via setdefault (load-bearing here — the chat path
already populates extra_body from custom providers, reasoning config and
Nous Portal), and an env escape hatch, HERMES_CHAT_SDK_TRANSFORM=1,
mirroring HERMES_CODEX_SDK_TRANSFORM.

The summary/compression call sites at chat_completion_helpers.py:3449 and
:3514 carry the largest payloads in the process and are deliberately left
for a follow-up: they route through a lambda whose client is not in scope
at the call site, so they need a slightly different shape and a wider
test surface than this change.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint P1 High — major feature broken, no workaround provider/openai OpenAI / Codex Responses API type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Agent-wide freeze: SDK maybe_transform can wedge for hours holding the GIL before the request is sent (codex_responses)

2 participants