Repository navigation
fix(sse): preserve parallel_tool_calls for GPT-5.6 delegation under Codex Responses Lite (#7821) #7957
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
fix(sse): preserve parallel_tool_calls for GPT-5.6 delegation under Codex Responses Lite (#7821) #7957
Changes from all commits
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1 @@ | ||
| - fix(sse): preserve parallel_tool_calls for GPT-5.6 ultra/max delegation under Codex Responses Lite (#7821) |
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,89 @@ | ||
| import test from "node:test"; | ||
| import assert from "node:assert/strict"; | ||
|
|
||
| import { CodexExecutor } from "../../open-sse/executors/codex.ts"; | ||
|
|
||
| // Issue #7821: enforceCodexResponsesLiteParallelToolCalls() used to have zero model | ||
| // awareness — it force-set parallel_tool_calls:false for EVERY model when Responses | ||
| // Lite was detected, including gpt-5.6-sol/-terra at "ultra" effort (and gpt-5.6-luna | ||
| // at "max"), whose delegation-to-sub-agents capability depends on parallel_tool_calls | ||
| // staying enabled. This collided with the effort-clamp comment near clampEffort() | ||
| // ("Ultra coordinates delegation in Codex clients") and is why GPT-5.6 was reported | ||
| // unusable through the stock Codex CLI/App (which enables Responses Lite by default) | ||
| // while GPT-5.5 (no delegation tier) was unaffected. | ||
| // | ||
| // Covers the interaction (lite marker + delegation-dependent model/effort together), | ||
| // not just each behavior in isolation — that interaction was the actual blind spot. | ||
|
|
||
| async function runLiteRequest(model: string): Promise<Record<string, unknown>[]> { | ||
| const executor = new CodexExecutor(); | ||
| const originalFetch = globalThis.fetch; | ||
| const capturedBodies: Record<string, unknown>[] = []; | ||
|
|
||
| globalThis.fetch = async (_url, init) => { | ||
| capturedBodies.push(JSON.parse(String(init?.body || "{}"))); | ||
| return new Response(JSON.stringify({ id: "resp_lite", object: "response" }), { | ||
| status: 200, | ||
| headers: { "Content-Type": "application/json" }, | ||
| }); | ||
| }; | ||
|
|
||
| const body = { | ||
| _nativeCodexPassthrough: true, | ||
| model, | ||
| input: [], | ||
| parallel_tool_calls: true, | ||
| }; | ||
|
|
||
| try { | ||
| await executor.execute({ | ||
| model, | ||
| body, | ||
| stream: true, | ||
| credentials: { accessToken: "codex-token" }, | ||
| clientHeaders: { "X-OpenAI-Internal-Codex-Responses-Lite": "true" }, | ||
| }); | ||
| } finally { | ||
| globalThis.fetch = originalFetch; | ||
| } | ||
|
|
||
| return capturedBodies; | ||
| } | ||
|
Comment on lines
+18
to
+51
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. To support testing delegation-dependent models where the effort is specified in the request body rather than the model suffix, we should update the async function runLiteRequest(model: string, bodyOverride?: Record<string, unknown>): Promise<Record<string, unknown>[]> {
const executor = new CodexExecutor();
const originalFetch = globalThis.fetch;
const capturedBodies: Record<string, unknown>[] = [];
globalThis.fetch = async (_url, init) => {
capturedBodies.push(JSON.parse(String(init?.body || "{}")));
return new Response(JSON.stringify({ id: "resp_lite", object: "response" }), {
status: 200,
headers: { "Content-Type": "application/json" },
});
};
const body = {
_nativeCodexPassthrough: true,
model,
input: [],
parallel_tool_calls: true,
...bodyOverride,
};
try {
await executor.execute({
model,
body,
stream: true,
credentials: { accessToken: "codex-token" },
clientHeaders: { "X-OpenAI-Internal-Codex-Responses-Lite": "true" },
});
} finally {
globalThis.fetch = originalFetch;
}
return capturedBodies;
} |
||
|
|
||
| test("Responses Lite must not strip parallel_tool_calls for GPT-5.6 sol ultra-tier delegation", async () => { | ||
| const capturedBodies = await runLiteRequest("gpt-5.6-sol-ultra"); | ||
| assert.equal( | ||
| capturedBodies[0].parallel_tool_calls, | ||
| true, | ||
| "Responses Lite stripped parallel_tool_calls for an ultra-tier GPT-5.6 delegation " + | ||
| "request — this breaks sub-agent delegation and is the root cause of #7821" | ||
| ); | ||
| }); | ||
|
|
||
| test("Responses Lite must not strip parallel_tool_calls for GPT-5.6 terra ultra-tier delegation", async () => { | ||
| const capturedBodies = await runLiteRequest("gpt-5.6-terra-ultra"); | ||
| assert.equal(capturedBodies[0].parallel_tool_calls, true); | ||
| }); | ||
|
|
||
| test("Responses Lite must not strip parallel_tool_calls for GPT-5.6 luna max-tier delegation", async () => { | ||
| const capturedBodies = await runLiteRequest("gpt-5.6-luna-max"); | ||
| assert.equal(capturedBodies[0].parallel_tool_calls, true); | ||
| }); | ||
|
|
||
| test("Responses Lite still forces parallel_tool_calls:false for non-delegation GPT-5.5", async () => { | ||
| const capturedBodies = await runLiteRequest("gpt-5.5"); | ||
| assert.equal( | ||
| capturedBodies[0].parallel_tool_calls, | ||
| false, | ||
| "GPT-5.5 has no delegation tier — Responses Lite behavior for it must be unchanged" | ||
| ); | ||
| }); | ||
|
|
||
| test("Responses Lite still forces parallel_tool_calls:false for GPT-5.6 sol at non-ultra effort", async () => { | ||
| const capturedBodies = await runLiteRequest("gpt-5.6-sol-high"); | ||
| assert.equal( | ||
| capturedBodies[0].parallel_tool_calls, | ||
| false, | ||
| "Non-ultra GPT-5.6 effort tiers have no delegation dependency — must stay forced off" | ||
| ); | ||
| }); | ||
|
Comment on lines
+82
to
+89
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. |
||
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
The current implementation of
isCodexDelegationDependentModelonly checks the model suffix (e.g.,gpt-5.6-sol-ultra) to determine if the request is delegation-dependent. However, clients can also specify the reasoning effort in the request body viareasoning.effortorreasoning_effort(e.g., with a base modelgpt-5.6-solandreasoning: { effort: "ultra" }). In such cases,isCodexDelegationDependentModelwill returnfalse, andenforceCodexResponsesLiteParallelToolCallswill incorrectly forceparallel_tool_calls: false, silently breaking delegation.We should update
isCodexDelegationDependentModelto also inspect the request body for the reasoning effort if it is not present in the model suffix.