Skip to content

fix(sse): preserve parallel_tool_calls for GPT-5.6 delegation under Codex Responses Lite (#7821) - #7957

Merged
diegosouzapw merged 2 commits into
release/v3.8.49from
fix/7821-gpt56-codex-responses-lite
Jul 21, 2026
Merged

diegosouzapw merged 2 commits into
release/v3.8.49from
fix/7821-gpt56-codex-responses-lite

Conversation

@diegosouzapw

Copy link
Copy Markdown
Owner

Closes #7821

Root cause

enforceCodexResponsesLiteParallelToolCalls() in open-sse/executors/codex.ts was model-blind: whenever the client marked Responses Lite (the current Codex CLI/App default), it unconditionally force-set parallel_tool_calls: false for every model — including GPT-5.6 at the ultra/max tiers (gpt-5.6-sol/gpt-5.6-terra at ultra, gpt-5.6-luna at max), whose sub-agent delegation depends on parallel tool calls staying enabled (see the clampEffort() comment: "Ultra coordinates delegation in Codex clients"). GPT-5.5 has no delegation tier, so losing parallel_tool_calls cost it nothing — which matches the exact split users reported (GPT-5.6 unusable, GPT-5.5 fine).

The existing test (tests/unit/executor-codex-gpt56.test.ts) covered effort-clamping and lite-forcing as two separate scenarios, never combining a lite marker with an ultra/max-tier request in the same test, so the interaction was invisible to CI.

Fix

  • enforceCodexResponsesLiteParallelToolCalls() now takes the resolved model and skips the forced override for delegation-dependent model/effort combos (isCodexDelegationDependentModel(), reusing splitCodexReasoningSuffix() + the existing GPT_5_6_ULTRA_ALIAS_MODELS set). All other models keep the prior lite-forcing behavior unchanged.
  • Added parallel_tool_calls to RESPONSES_API_ALLOWLIST so the same field isn't stripped a second time on the Chat-Completions-translated (non-native) path if Responses Lite support is ever extended past native Codex CLI/App traffic (secondary latent bug noted in the plan-file, not what users are hitting today since native traffic returns early before that allowlist runs).

Regression test

tests/unit/executor-codex-gpt56-lite-ultra.test.ts — 5 cases: gpt-5.6-sol-ultra, gpt-5.6-terra-ultra, gpt-5.6-luna-max keep parallel_tool_calls: true under Responses Lite; gpt-5.5 and gpt-5.6-sol-high (non-delegation) still get it forced to false.

RED (against origin/release/v3.8.49, restored via git show — no stash used):

✖ Responses Lite must not strip parallel_tool_calls for GPT-5.6 sol ultra-tier delegation
✖ Responses Lite must not strip parallel_tool_calls for GPT-5.6 terra ultra-tier delegation
✖ Responses Lite must not strip parallel_tool_calls for GPT-5.6 luna max-tier delegation
  AssertionError [ERR_ASSERTION]: Expected values to be strictly equal: false !== true
✔ Responses Lite still forces parallel_tool_calls:false for non-delegation GPT-5.5
✔ Responses Lite still forces parallel_tool_calls:false for GPT-5.6 sol at non-ultra effort

GREEN (with the fix):

✔ Responses Lite must not strip parallel_tool_calls for GPT-5.6 sol ultra-tier delegation
✔ Responses Lite must not strip parallel_tool_calls for GPT-5.6 terra ultra-tier delegation
✔ Responses Lite must not strip parallel_tool_calls for GPT-5.6 luna max-tier delegation
✔ Responses Lite still forces parallel_tool_calls:false for non-delegation GPT-5.5
✔ Responses Lite still forces parallel_tool_calls:false for GPT-5.6 sol at non-ultra effort
ℹ tests 5, pass 5, fail 0

Sibling suites (touched-area regression check)

All green, unchanged behavior confirmed:

  • tests/unit/executor-codex-gpt56.test.ts
  • tests/unit/codex-gpt56-catalog.test.ts
  • tests/unit/codex-responses-passthrough-strip-3317.test.ts
  • tests/unit/codex-responses-ws-memory.test.ts
  • tests/unit/codex-responses-ws-model-resolution.test.ts
  • tests/unit/combo-name-codex-responses-rewrite.test.ts
  • tests/unit/auto-combo-codex-responses-3509.test.ts
  • tests/unit/responses-transformer.test.ts
  • tests/unit/responses-transformer-dense-output.test.ts
  • tests/unit/t19-codex-responses-empty-content.test.ts

Gates run (all green)

  • node scripts/check/check-file-size.mjs — no violations on touched files
  • node scripts/check/check-complexity.mjs — OK (2084 violations vs baseline 2130)
  • node scripts/check/check-cognitive-complexity.mjs — OK (903 violations vs baseline 950)
  • node scripts/check/check-mutation-test-coverage.mjs --strict — no drift
  • node node_modules/typescript/bin/tsc --pretty false -p tsconfig.typecheck-core.json (typecheck:core) — clean
  • npx eslint --suppressions-location config/quality/eslint-suppressions.json open-sse/executors/codex.ts tests/unit/executor-codex-gpt56-lite-ultra.test.ts — clean
  • node scripts/check/check-changelog-integrity.mjs — OK

Interim mitigation

Until this ships, affected users can work around it client-side via codex debug models + use_responses_lite: false + multi_agent_version: v1, per the issue's acceptance criteria. Documenting that workaround in docs/frameworks/CLOUD_AGENT.md / config-codex-cli is left as a follow-up (out of scope for this surgical fix).

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request addresses an issue where parallel tool calls were incorrectly stripped for delegation-dependent GPT-5.6 models under Codex Responses Lite, which silently broke sub-agent delegation. The changes introduce a helper to identify these models and preserve parallel tool calls, accompanied by unit tests. The review feedback points out a critical gap: the implementation only checks the model suffix for the reasoning effort, whereas clients can also specify it in the request body. The reviewer provides actionable suggestions to inspect the request body for reasoning effort, update the test helper, and add regression tests.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment on lines +160 to 180
function isCodexDelegationDependentModel(model: unknown): boolean {
const { baseModel, effort } = splitCodexReasoningSuffix(model);
if (effort === "ultra" && GPT_5_6_ULTRA_ALIAS_MODELS.has(baseModel)) return true;
if (effort === "max" && baseModel === "gpt-5.6-luna") return true;
return false;
}

function enforceCodexResponsesLiteParallelToolCalls(
bodyInput: unknown,
clientHeaders?: Record<string, string> | null
clientHeaders: Record<string, string> | null | undefined,
model: unknown
): unknown {
if (
!isCodexResponsesLiteRequest(bodyInput, clientHeaders) ||
!bodyInput ||
typeof bodyInput !== "object" ||
Array.isArray(bodyInput)
Array.isArray(bodyInput) ||
isCodexDelegationDependentModel(model)
) {
return bodyInput;
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The current implementation of isCodexDelegationDependentModel only checks the model suffix (e.g., gpt-5.6-sol-ultra) to determine if the request is delegation-dependent. However, clients can also specify the reasoning effort in the request body via reasoning.effort or reasoning_effort (e.g., with a base model gpt-5.6-sol and reasoning: { effort: "ultra" }). In such cases, isCodexDelegationDependentModel will return false, and enforceCodexResponsesLiteParallelToolCalls will incorrectly force parallel_tool_calls: false, silently breaking delegation.

We should update isCodexDelegationDependentModel to also inspect the request body for the reasoning effort if it is not present in the model suffix.

function isCodexDelegationDependentModel(model: unknown, bodyInput: unknown): boolean {
  const { baseModel, effort: modelEffort } = splitCodexReasoningSuffix(model);
  let effort = modelEffort;
  if (!effort && bodyInput && typeof bodyInput === "object" && !Array.isArray(bodyInput)) {
    const body = bodyInput as Record<string, unknown>;
    const reasoningRecord =
      body.reasoning && typeof body.reasoning === "object" && !Array.isArray(body.reasoning)
        ? (body.reasoning as Record<string, unknown>)
        : null;
    const explicitReasoning = typeof reasoningRecord?.effort === "string" ? reasoningRecord.effort.trim().toLowerCase() : undefined;
    const requestReasoningEffort = typeof body.reasoning_effort === "string" ? body.reasoning_effort.trim().toLowerCase() : undefined;
    effort = (explicitReasoning || requestReasoningEffort || null) as EffortLevel | null;
  }
  if (effort === "ultra" && GPT_5_6_ULTRA_ALIAS_MODELS.has(baseModel)) return true;
  if (effort === "max" && baseModel === "gpt-5.6-luna") return true;
  return false;
}

function enforceCodexResponsesLiteParallelToolCalls(
  bodyInput: unknown,
  clientHeaders: Record<string, string> | null | undefined,
  model: unknown
): unknown {
  if (
    !isCodexResponsesLiteRequest(bodyInput, clientHeaders) ||
    !bodyInput ||
    typeof bodyInput !== "object" ||
    Array.isArray(bodyInput) ||
    isCodexDelegationDependentModel(model, bodyInput)
  ) {
    return bodyInput;
  }

Comment on lines +18 to +51
async function runLiteRequest(model: string): Promise<Record<string, unknown>[]> {
const executor = new CodexExecutor();
const originalFetch = globalThis.fetch;
const capturedBodies: Record<string, unknown>[] = [];

globalThis.fetch = async (_url, init) => {
capturedBodies.push(JSON.parse(String(init?.body || "{}")));
return new Response(JSON.stringify({ id: "resp_lite", object: "response" }), {
status: 200,
headers: { "Content-Type": "application/json" },
});
};

const body = {
_nativeCodexPassthrough: true,
model,
input: [],
parallel_tool_calls: true,
};

try {
await executor.execute({
model,
body,
stream: true,
credentials: { accessToken: "codex-token" },
clientHeaders: { "X-OpenAI-Internal-Codex-Responses-Lite": "true" },
});
} finally {
globalThis.fetch = originalFetch;
}

return capturedBodies;
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

To support testing delegation-dependent models where the effort is specified in the request body rather than the model suffix, we should update the runLiteRequest helper to accept an optional bodyOverride parameter.

async function runLiteRequest(model: string, bodyOverride?: Record<string, unknown>): Promise<Record<string, unknown>[]> {
  const executor = new CodexExecutor();
  const originalFetch = globalThis.fetch;
  const capturedBodies: Record<string, unknown>[] = [];

  globalThis.fetch = async (_url, init) => {
    capturedBodies.push(JSON.parse(String(init?.body || "{}")));
    return new Response(JSON.stringify({ id: "resp_lite", object: "response" }), {
      status: 200,
      headers: { "Content-Type": "application/json" },
    });
  };

  const body = {
    _nativeCodexPassthrough: true,
    model,
    input: [],
    parallel_tool_calls: true,
    ...bodyOverride,
  };

  try {
    await executor.execute({
      model,
      body,
      stream: true,
      credentials: { accessToken: "codex-token" },
      clientHeaders: { "X-OpenAI-Internal-Codex-Responses-Lite": "true" },
    });
  } finally {
    globalThis.fetch = originalFetch;
  }

  return capturedBodies;
}

Comment on lines +82 to +89
test("Responses Lite still forces parallel_tool_calls:false for GPT-5.6 sol at non-ultra effort", async () => {
const capturedBodies = await runLiteRequest("gpt-5.6-sol-high");
assert.equal(
capturedBodies[0].parallel_tool_calls,
false,
"Non-ultra GPT-5.6 effort tiers have no delegation dependency — must stay forced off"
);
});

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Add regression tests to verify that parallel_tool_calls is preserved when the delegation-dependent effort (ultra or max) is specified in the request body (either via reasoning.effort or reasoning_effort) instead of the model suffix.

test(

…#2608 stripping intact (#7821)

The static RESPONSES_API_ALLOWLIST addition made parallel_tool_calls survive for
ALL models, breaking the #2608 non-passthrough stripping guarantee for gpt-5.5.
The real #7821 fix (isCodexDelegationDependentModel gating in
enforceCodexResponsesLiteParallelToolCalls) is model/effort-scoped and does not
need the allowlist entry — native Codex traffic returns before the allowlist runs.
@diegosouzapw
diegosouzapw merged commit 0d4fbfe into release/v3.8.49 Jul 21, 2026
20 checks passed
@diegosouzapw
diegosouzapw deleted the fix/7821-gpt56-codex-responses-lite branch July 23, 2026 00:24
HouMinXi pushed a commit to HouMinXi/OmniRoute that referenced this pull request Aug 2, 2026
…odex Responses Lite (diegosouzapw#7821) (diegosouzapw#7957)

* fix(sse): preserve parallel_tool_calls for GPT-5.6 ultra/max delegation under Codex Responses Lite (diegosouzapw#7821)

* fix(codex): drop over-broad parallel_tool_calls allowlist entry — keep diegosouzapw#2608 stripping intact (diegosouzapw#7821)

The static RESPONSES_API_ALLOWLIST addition made parallel_tool_calls survive for
ALL models, breaking the diegosouzapw#2608 non-passthrough stripping guarantee for gpt-5.5.
The real diegosouzapw#7821 fix (isCodexDelegationDependentModel gating in
enforceCodexResponsesLiteParallelToolCalls) is model/effort-scoped and does not
need the allowlist entry — native Codex traffic returns before the allowlist runs.
fenix007 pushed a commit to fenix007/OmniRoute that referenced this pull request Aug 20, 2026
…odex Responses Lite (diegosouzapw#7821) (diegosouzapw#7957)

* fix(sse): preserve parallel_tool_calls for GPT-5.6 ultra/max delegation under Codex Responses Lite (diegosouzapw#7821)

* fix(codex): drop over-broad parallel_tool_calls allowlist entry — keep diegosouzapw#2608 stripping intact (diegosouzapw#7821)

The static RESPONSES_API_ALLOWLIST addition made parallel_tool_calls survive for
ALL models, breaking the diegosouzapw#2608 non-passthrough stripping guarantee for gpt-5.5.
The real diegosouzapw#7821 fix (isCodexDelegationDependentModel gating in
enforceCodexResponsesLiteParallelToolCalls) is model/effort-scoped and does not
need the allowlist entry — native Codex traffic returns before the allowlist runs.

(cherry picked from commit 0d4fbfe)
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…odex Responses Lite (diegosouzapw#7821) (diegosouzapw#7957)

* fix(sse): preserve parallel_tool_calls for GPT-5.6 ultra/max delegation under Codex Responses Lite (diegosouzapw#7821)

* fix(codex): drop over-broad parallel_tool_calls allowlist entry — keep diegosouzapw#2608 stripping intact (diegosouzapw#7821)

The static RESPONSES_API_ALLOWLIST addition made parallel_tool_calls survive for
ALL models, breaking the diegosouzapw#2608 non-passthrough stripping guarantee for gpt-5.5.
The real diegosouzapw#7821 fix (isCodexDelegationDependentModel gating in
enforceCodexResponsesLiteParallelToolCalls) is model/effort-scoped and does not
need the allowlist entry — native Codex traffic returns before the allowlist runs.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(providers): GPT-5.6 unusable via Codex CLI/App — Responses Lite forces parallel_tool_calls:false, breaking ultra delegation

1 participant