Repository navigation
fix: route codex payloads around the SDK's GIL-holding request transform (#93650) - #93773
Merged
kshitijk4poor merged 3 commits intoAug 24, 2026
Conversation
…orm (NousResearch#93650) responses.create re-walks the entire request body against the ResponseCreateParams union graph client-side while holding the GIL. NousResearch#93650 documents that walk wedging for 12+ hours on a ~1.4 MB conversation, starving every other thread including the TTFB/stale watchdogs — and no socket kill can unblock a pre-network hang. Hermes payloads are JSON round-trips and already wire format, so the bulk fields (input, tools) are now routed through extra_body, which the SDK merges into the JSON body after the transform. Guarded by a plain-JSON check (anything else keeps the typed path) and a HERMES_CODEX_SDK_TRANSFORM=1 escape hatch. Applied to both the primary stream path and the auxiliary adapter. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
kshitijk4poor
enabled auto-merge (rebase)
August 24, 2026 10:32
auto-merge was automatically disabled
August 24, 2026 10:32
Rebase failed
13 tasks done
melon-xf
added a commit
to melon-xf/hermes-agent
that referenced
this pull request
Sep 3, 2026
…k-transform-bypass-93650-v2 fix: route codex payloads around the SDK's GIL-holding request transform (NousResearch#93650)
kshitijk4poor
pushed a commit
to kshitijk4poor/hermes-agent
that referenced
this pull request
Sep 9, 2026
…st transform `chat.completions.create` re-walks the whole request body against the `CompletionCreateParams` union graph client-side, with the GIL held, before any byte leaves the process. NousResearch#93650 documented that class of walk wedging for 12+ hours on a ~1.4 MB conversation: no in-process watchdog can fire while the GIL is held, and no socket kill helps a pre-network hang. NousResearch#93773 merged the remedy — route the already-wire-format bulk fields through `extra_body`, which the SDK merges into the JSON body after the transform — but scoped it to `responses.create`. The default chat path, which every OpenRouter / Nous / xAI / DeepSeek / Kimi / llama.cpp / Ollama / LM Studio / LiteLLM request takes, still pays the full walk. Measured against a real `openai.OpenAI` over an `httpx.MockTransport` (canned SSE, no network), with the request body captured from the transport on both sides: 101 msgs / 76 KB 13.6 ms -> 1.2 ms 401 msgs / 190 KB 48.6 ms -> 2.0 ms 1601 msgs / 650 KB 188.7 ms -> 5.8 ms and the bytes the server receives are IDENTICAL — literally equal, not merely equivalent (194,894 == 194,894 at 401 messages). The cost is paid per API call, so a tool-using turn multiplies it by its iteration count. The three helpers move from agent/codex_runtime.py into a shared agent/sdk_transform_bypass.py, re-exported under their original names so agent/auxiliary_client.py and tests/run_agent/test_codex_sdk_transform_bypass.py keep working untouched. The field tuple is now a parameter: ("input", "tools") for Responses, ("messages", "tools") for chat. Two chat-specific details. `messages` is a @required_args parameter, so it stays in the typed kwargs as an empty list and the extra_body copy replaces it in the body — hence the new `required_empty` argument, which Responses does not use. And the bypass is gated on the target actually being the SDK's Completions: Hermes also drives chat-completions-shaped facades that are NOT the SDK — the in-process MoA aggregator most importantly — and those never merge extra_body, so handing them one would silently send an empty message list. That guard is also why this needs no edits to the 32 test files that assert on kwargs["messages"]: they mock with stand-ins, not the SDK. Every rail the merged PR was reviewed on is kept: the plain-JSON-only guard so pydantic models and generators stay on the typed path, caller `extra_body` precedence via setdefault (load-bearing here — the chat path already populates extra_body from custom providers, reasoning config and Nous Portal), and an env escape hatch, HERMES_CHAT_SDK_TRANSFORM=1, mirroring HERMES_CODEX_SDK_TRANSFORM. The summary/compression call sites at chat_completion_helpers.py:3449 and :3514 carry the largest payloads in the process and are deliberately left for a follow-up: they route through a lambda whose client is not in scope at the call site, so they need a slightly different shape and a wider test surface than this change.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #93650. Salvage of #93681 (@kchernev).
Summary
Routes the bulk payload fields (
input,tools) throughextra_bodyinstead of the typed SDK kwargs, bypassing the OpenAI SDK'smaybe_transformwalk that can wedge for hours holding the GIL on large Codex Responses API calls.Problem
A
codex_responsesstreaming call can freeze the entire agent for hours before a single byte leaves the process. The OpenAI SDK'sresponses.createre-walks the whole request body against theResponseCreateParamsunion graph (openai/_utils/_transform.py::maybe_transform) client-side, holding the GIL. In the incident documented in #93650 (openai 2.24.0, CPython 3.11.16), one such walk over a ~1.4 MB / 1,093-item conversation wedged for 12.5 CPU-hours at ~99% of a core.Core-dump forensics showed the worker in
maybe_transform → isinstance(obj, Mapping) → typing.__subclasscheck__, with 9 other threads intake_gilfutex waits — including the TTFB/stale watchdog poller, which never got to run any of its detectors. No in-process watchdog can rescue this: the GIL-holding walk starves them, and even when they run their remedy is socket-level (_close_request_client_once) — a pre-network hang has no socket to kill.Fix
Hermes assembles Codex payloads from JSON round-trips — they are already wire format, so the transform walk is pure overhead plus this failure mode. The SDK merges
extra_bodyinto the JSON body after the transform (_base_client._build_request, shallow merge), a mechanism this codebase already relies on forprompt_cache_retention. This PR routes the bulk payload fields (input,tools) throughextra_body, skippingmaybe_transformfor them entirely while producing an identical request body.Safety rails:
dict/list/str/int/float/bool/Nonewithstrkeys). Non-JSON types keep the typed SDK path.extra_bodyentry wins over a moved field.HERMES_CODEX_SDK_TRANSFORM=1restores the pre-fix behavior.Applied at both live
responses.createcall sites: the primary stream path (codex_runtime._open_codex_stream) and the auxiliary adapter (auxiliary_client._CodexCompletionsAdapter).Changes
agent/codex_runtime.py: +74 lines —_is_plain_json_data,_bypass_sdk_request_transform, wired into_open_codex_streamagent/auxiliary_client.py: +7/-1 — wired into_CodexCompletionsAdapter.createtests/run_agent/test_codex_sdk_transform_bypass.py: +153 lines — 10 new teststests/agent/test_auxiliary_client.py: +14/-1 — fake-client capture updated for both wire shapestests/run_agent/test_native_compaction.py: +5/-1 — fake-client capture updatedtests/run_agent/test_run_agent_codex_responses.py: +6/-1 — retention guard updated for new wire shapecontributors/emails/: AUTHOR_MAP entry for @kchernevValidation
Competitor research
maybe_transformnever runschat.completions.create(smaller typedict graph, not the massiveResponseCreateParamsunion)No other major harness uses the OpenAI Python SDK's
responses.create()with payloads as large as Hermes does (558K tokens, 1.4MB body). Hermes is uniquely exposed.