Repository navigation
fix(sse): preserve parallel_tool_calls for GPT-5.6 delegation under Codex Responses Lite (#7821) - #7957
Conversation
…on under Codex Responses Lite (#7821)
There was a problem hiding this comment.
Code Review
This pull request addresses an issue where parallel tool calls were incorrectly stripped for delegation-dependent GPT-5.6 models under Codex Responses Lite, which silently broke sub-agent delegation. The changes introduce a helper to identify these models and preserve parallel tool calls, accompanied by unit tests. The review feedback points out a critical gap: the implementation only checks the model suffix for the reasoning effort, whereas clients can also specify it in the request body. The reviewer provides actionable suggestions to inspect the request body for reasoning effort, update the test helper, and add regression tests.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
| function isCodexDelegationDependentModel(model: unknown): boolean { | ||
| const { baseModel, effort } = splitCodexReasoningSuffix(model); | ||
| if (effort === "ultra" && GPT_5_6_ULTRA_ALIAS_MODELS.has(baseModel)) return true; | ||
| if (effort === "max" && baseModel === "gpt-5.6-luna") return true; | ||
| return false; | ||
| } | ||
|
|
||
| function enforceCodexResponsesLiteParallelToolCalls( | ||
| bodyInput: unknown, | ||
| clientHeaders?: Record<string, string> | null | ||
| clientHeaders: Record<string, string> | null | undefined, | ||
| model: unknown | ||
| ): unknown { | ||
| if ( | ||
| !isCodexResponsesLiteRequest(bodyInput, clientHeaders) || | ||
| !bodyInput || | ||
| typeof bodyInput !== "object" || | ||
| Array.isArray(bodyInput) | ||
| Array.isArray(bodyInput) || | ||
| isCodexDelegationDependentModel(model) | ||
| ) { | ||
| return bodyInput; | ||
| } |
There was a problem hiding this comment.
The current implementation of isCodexDelegationDependentModel only checks the model suffix (e.g., gpt-5.6-sol-ultra) to determine if the request is delegation-dependent. However, clients can also specify the reasoning effort in the request body via reasoning.effort or reasoning_effort (e.g., with a base model gpt-5.6-sol and reasoning: { effort: "ultra" }). In such cases, isCodexDelegationDependentModel will return false, and enforceCodexResponsesLiteParallelToolCalls will incorrectly force parallel_tool_calls: false, silently breaking delegation.
We should update isCodexDelegationDependentModel to also inspect the request body for the reasoning effort if it is not present in the model suffix.
function isCodexDelegationDependentModel(model: unknown, bodyInput: unknown): boolean {
const { baseModel, effort: modelEffort } = splitCodexReasoningSuffix(model);
let effort = modelEffort;
if (!effort && bodyInput && typeof bodyInput === "object" && !Array.isArray(bodyInput)) {
const body = bodyInput as Record<string, unknown>;
const reasoningRecord =
body.reasoning && typeof body.reasoning === "object" && !Array.isArray(body.reasoning)
? (body.reasoning as Record<string, unknown>)
: null;
const explicitReasoning = typeof reasoningRecord?.effort === "string" ? reasoningRecord.effort.trim().toLowerCase() : undefined;
const requestReasoningEffort = typeof body.reasoning_effort === "string" ? body.reasoning_effort.trim().toLowerCase() : undefined;
effort = (explicitReasoning || requestReasoningEffort || null) as EffortLevel | null;
}
if (effort === "ultra" && GPT_5_6_ULTRA_ALIAS_MODELS.has(baseModel)) return true;
if (effort === "max" && baseModel === "gpt-5.6-luna") return true;
return false;
}
function enforceCodexResponsesLiteParallelToolCalls(
bodyInput: unknown,
clientHeaders: Record<string, string> | null | undefined,
model: unknown
): unknown {
if (
!isCodexResponsesLiteRequest(bodyInput, clientHeaders) ||
!bodyInput ||
typeof bodyInput !== "object" ||
Array.isArray(bodyInput) ||
isCodexDelegationDependentModel(model, bodyInput)
) {
return bodyInput;
}| async function runLiteRequest(model: string): Promise<Record<string, unknown>[]> { | ||
| const executor = new CodexExecutor(); | ||
| const originalFetch = globalThis.fetch; | ||
| const capturedBodies: Record<string, unknown>[] = []; | ||
|
|
||
| globalThis.fetch = async (_url, init) => { | ||
| capturedBodies.push(JSON.parse(String(init?.body || "{}"))); | ||
| return new Response(JSON.stringify({ id: "resp_lite", object: "response" }), { | ||
| status: 200, | ||
| headers: { "Content-Type": "application/json" }, | ||
| }); | ||
| }; | ||
|
|
||
| const body = { | ||
| _nativeCodexPassthrough: true, | ||
| model, | ||
| input: [], | ||
| parallel_tool_calls: true, | ||
| }; | ||
|
|
||
| try { | ||
| await executor.execute({ | ||
| model, | ||
| body, | ||
| stream: true, | ||
| credentials: { accessToken: "codex-token" }, | ||
| clientHeaders: { "X-OpenAI-Internal-Codex-Responses-Lite": "true" }, | ||
| }); | ||
| } finally { | ||
| globalThis.fetch = originalFetch; | ||
| } | ||
|
|
||
| return capturedBodies; | ||
| } |
There was a problem hiding this comment.
To support testing delegation-dependent models where the effort is specified in the request body rather than the model suffix, we should update the runLiteRequest helper to accept an optional bodyOverride parameter.
async function runLiteRequest(model: string, bodyOverride?: Record<string, unknown>): Promise<Record<string, unknown>[]> {
const executor = new CodexExecutor();
const originalFetch = globalThis.fetch;
const capturedBodies: Record<string, unknown>[] = [];
globalThis.fetch = async (_url, init) => {
capturedBodies.push(JSON.parse(String(init?.body || "{}")));
return new Response(JSON.stringify({ id: "resp_lite", object: "response" }), {
status: 200,
headers: { "Content-Type": "application/json" },
});
};
const body = {
_nativeCodexPassthrough: true,
model,
input: [],
parallel_tool_calls: true,
...bodyOverride,
};
try {
await executor.execute({
model,
body,
stream: true,
credentials: { accessToken: "codex-token" },
clientHeaders: { "X-OpenAI-Internal-Codex-Responses-Lite": "true" },
});
} finally {
globalThis.fetch = originalFetch;
}
return capturedBodies;
}| test("Responses Lite still forces parallel_tool_calls:false for GPT-5.6 sol at non-ultra effort", async () => { | ||
| const capturedBodies = await runLiteRequest("gpt-5.6-sol-high"); | ||
| assert.equal( | ||
| capturedBodies[0].parallel_tool_calls, | ||
| false, | ||
| "Non-ultra GPT-5.6 effort tiers have no delegation dependency — must stay forced off" | ||
| ); | ||
| }); |
…#2608 stripping intact (#7821) The static RESPONSES_API_ALLOWLIST addition made parallel_tool_calls survive for ALL models, breaking the #2608 non-passthrough stripping guarantee for gpt-5.5. The real #7821 fix (isCodexDelegationDependentModel gating in enforceCodexResponsesLiteParallelToolCalls) is model/effort-scoped and does not need the allowlist entry — native Codex traffic returns before the allowlist runs.
…odex Responses Lite (diegosouzapw#7821) (diegosouzapw#7957) * fix(sse): preserve parallel_tool_calls for GPT-5.6 ultra/max delegation under Codex Responses Lite (diegosouzapw#7821) * fix(codex): drop over-broad parallel_tool_calls allowlist entry — keep diegosouzapw#2608 stripping intact (diegosouzapw#7821) The static RESPONSES_API_ALLOWLIST addition made parallel_tool_calls survive for ALL models, breaking the diegosouzapw#2608 non-passthrough stripping guarantee for gpt-5.5. The real diegosouzapw#7821 fix (isCodexDelegationDependentModel gating in enforceCodexResponsesLiteParallelToolCalls) is model/effort-scoped and does not need the allowlist entry — native Codex traffic returns before the allowlist runs.
…odex Responses Lite (diegosouzapw#7821) (diegosouzapw#7957) * fix(sse): preserve parallel_tool_calls for GPT-5.6 ultra/max delegation under Codex Responses Lite (diegosouzapw#7821) * fix(codex): drop over-broad parallel_tool_calls allowlist entry — keep diegosouzapw#2608 stripping intact (diegosouzapw#7821) The static RESPONSES_API_ALLOWLIST addition made parallel_tool_calls survive for ALL models, breaking the diegosouzapw#2608 non-passthrough stripping guarantee for gpt-5.5. The real diegosouzapw#7821 fix (isCodexDelegationDependentModel gating in enforceCodexResponsesLiteParallelToolCalls) is model/effort-scoped and does not need the allowlist entry — native Codex traffic returns before the allowlist runs. (cherry picked from commit 0d4fbfe)
…odex Responses Lite (diegosouzapw#7821) (diegosouzapw#7957) * fix(sse): preserve parallel_tool_calls for GPT-5.6 ultra/max delegation under Codex Responses Lite (diegosouzapw#7821) * fix(codex): drop over-broad parallel_tool_calls allowlist entry — keep diegosouzapw#2608 stripping intact (diegosouzapw#7821) The static RESPONSES_API_ALLOWLIST addition made parallel_tool_calls survive for ALL models, breaking the diegosouzapw#2608 non-passthrough stripping guarantee for gpt-5.5. The real diegosouzapw#7821 fix (isCodexDelegationDependentModel gating in enforceCodexResponsesLiteParallelToolCalls) is model/effort-scoped and does not need the allowlist entry — native Codex traffic returns before the allowlist runs.
Closes #7821
Root cause
enforceCodexResponsesLiteParallelToolCalls()inopen-sse/executors/codex.tswas model-blind: whenever the client marked Responses Lite (the current Codex CLI/App default), it unconditionally force-setparallel_tool_calls: falsefor every model — including GPT-5.6 at theultra/maxtiers (gpt-5.6-sol/gpt-5.6-terraatultra,gpt-5.6-lunaatmax), whose sub-agent delegation depends on parallel tool calls staying enabled (see theclampEffort()comment: "Ultra coordinates delegation in Codex clients"). GPT-5.5 has no delegation tier, so losingparallel_tool_callscost it nothing — which matches the exact split users reported (GPT-5.6 unusable, GPT-5.5 fine).The existing test (
tests/unit/executor-codex-gpt56.test.ts) covered effort-clamping and lite-forcing as two separate scenarios, never combining a lite marker with an ultra/max-tier request in the same test, so the interaction was invisible to CI.Fix
enforceCodexResponsesLiteParallelToolCalls()now takes the resolved model and skips the forced override for delegation-dependent model/effort combos (isCodexDelegationDependentModel(), reusingsplitCodexReasoningSuffix()+ the existingGPT_5_6_ULTRA_ALIAS_MODELSset). All other models keep the prior lite-forcing behavior unchanged.parallel_tool_callstoRESPONSES_API_ALLOWLISTso the same field isn't stripped a second time on the Chat-Completions-translated (non-native) path if Responses Lite support is ever extended past native Codex CLI/App traffic (secondary latent bug noted in the plan-file, not what users are hitting today since native traffic returns early before that allowlist runs).Regression test
tests/unit/executor-codex-gpt56-lite-ultra.test.ts— 5 cases:gpt-5.6-sol-ultra,gpt-5.6-terra-ultra,gpt-5.6-luna-maxkeepparallel_tool_calls: trueunder Responses Lite;gpt-5.5andgpt-5.6-sol-high(non-delegation) still get it forced tofalse.RED (against
origin/release/v3.8.49, restored viagit show— no stash used):GREEN (with the fix):
Sibling suites (touched-area regression check)
All green, unchanged behavior confirmed:
tests/unit/executor-codex-gpt56.test.tstests/unit/codex-gpt56-catalog.test.tstests/unit/codex-responses-passthrough-strip-3317.test.tstests/unit/codex-responses-ws-memory.test.tstests/unit/codex-responses-ws-model-resolution.test.tstests/unit/combo-name-codex-responses-rewrite.test.tstests/unit/auto-combo-codex-responses-3509.test.tstests/unit/responses-transformer.test.tstests/unit/responses-transformer-dense-output.test.tstests/unit/t19-codex-responses-empty-content.test.tsGates run (all green)
node scripts/check/check-file-size.mjs— no violations on touched filesnode scripts/check/check-complexity.mjs— OK (2084 violations vs baseline 2130)node scripts/check/check-cognitive-complexity.mjs— OK (903 violations vs baseline 950)node scripts/check/check-mutation-test-coverage.mjs --strict— no driftnode node_modules/typescript/bin/tsc --pretty false -p tsconfig.typecheck-core.json(typecheck:core) — cleannpx eslint --suppressions-location config/quality/eslint-suppressions.json open-sse/executors/codex.ts tests/unit/executor-codex-gpt56-lite-ultra.test.ts— cleannode scripts/check/check-changelog-integrity.mjs— OKInterim mitigation
Until this ships, affected users can work around it client-side via
codex debug models+use_responses_lite: false+multi_agent_version: v1, per the issue's acceptance criteria. Documenting that workaround indocs/frameworks/CLOUD_AGENT.md/config-codex-cliis left as a follow-up (out of scope for this surgical fix).