fix(cli): Codex long-session turn-pin fallback + codex-settings key resolution (#13564 #13563) - #13566
Conversation
…oped unusable (diegosouzapw#13564) Long-running Codex sessions die with 400 NATIVE_CODEX_PINNED_MODEL_UNAVAILABLE whenever the model pinned to the current turn becomes model-scoped unusable mid-session (per-model quota lockout, connection cooldown, exhausted accounts). Claude Code has no equivalent pin and already falls back to the next healthy combo model; Codex now matches. Release the turn pin when all pinned provider+model targets are model-scoped unusable and fall through to full combo routing, re-pinning to whichever model succeeds. Preserve the pin on provider-wide outages (circuit breaker OPEN, provider cooldown) and when the pinned target is still healthy. Also prunes the stale ESLint suppression entry for combo.ts that this change orphaned (createPinnedModelUnavailableResponse import dropped; pre-existing getBootstrapLatencyMs remains the sole residual unused var).
…d of 400 (diegosouzapw#13563) Applying Codex settings from /dashboard/cli-code/codex always failed with 400 "baseUrl, apiKey and model are required" when the dashboard sent an empty apiKey (cloud mode with no management key selected) — baseUrl and model are already Zod-gated, so that response could only ever fire on the empty key. The codex-settings route had diverged from the sibling CLI tools (cline/forge/ openclaw/grok-build/jcode): an inline if(!apiKey) 400 guard plus a hand-rolled getApiKeyById lookup, instead of the shared resolveApiKey(keyId, apiKey) helper which resolves by keyId, falls back to the submitted apiKey, then to sk_omniroute. This change makes codex-settings use the canonical resolver, so: - empty apiKey + valid keyId -> the real DB key is written to auth.json - empty apiKey + no keyId -> sk_omniroute default (config still applies) - explicit apiKey -> written verbatim (unchanged)
|
Solid pair of fixes, both well-isolated with real regression coverage. I ran both new test |
…turn-pin-fallback # Conflicts: # open-sse/services/combo.ts
…settings apiKey resolution
|
Changelog fragments added per review feedback — one per fix, as in the #13098 split:
I also rebased the branch onto the current Verified at the new head (1f24cc5):
PR is now mergeable (MERGEABLE, checks re-running). Thanks for the review — the circuit-breaker/cooldown scope notes were spot on. |
|
@diegosouzapw — fragments added and branch rebased onto the current base tip (now MERGEABLE). Details here: #issuecomment-5694559093. Ready for re-review when checks finish. |
|
Thanks @opensource-elearning — merging via the release merge-train. Validated in local merge-train (.claude/worktrees/merge-train-20260918-111718-suite.log) on the devbox @ train tip 7bb373fba5e241964c0ffb17bd03a804700839bb, boarded with 55 sibling PRs: typecheck:core, file-size, complexity, cognitive-complexity, changelog-integrity green; 747/747 changed-area node:test cases + 476/476 vitest green (fast parity mode — the full suite ran today on the release tip via the base-red train and runs again on the 3b train). Merged --admin per merge-gates §7. |
65263a4
into
diegosouzapw:release/v3.8.51
…esolution (diegosouzapw#13564 diegosouzapw#13563) (diegosouzapw#13566) * fix(sse): release native Codex turn pin when pinned model is model-scoped unusable (diegosouzapw#13564) Long-running Codex sessions die with 400 NATIVE_CODEX_PINNED_MODEL_UNAVAILABLE whenever the model pinned to the current turn becomes model-scoped unusable mid-session (per-model quota lockout, connection cooldown, exhausted accounts). Claude Code has no equivalent pin and already falls back to the next healthy combo model; Codex now matches. Release the turn pin when all pinned provider+model targets are model-scoped unusable and fall through to full combo routing, re-pinning to whichever model succeeds. Preserve the pin on provider-wide outages (circuit breaker OPEN, provider cooldown) and when the pinned target is still healthy. Also prunes the stale ESLint suppression entry for combo.ts that this change orphaned (createPinnedModelUnavailableResponse import dropped; pre-existing getBootstrapLatencyMs remains the sole residual unused var). * fix(api): resolve codex-settings apiKey via canonical resolver instead of 400 (diegosouzapw#13563) Applying Codex settings from /dashboard/cli-code/codex always failed with 400 "baseUrl, apiKey and model are required" when the dashboard sent an empty apiKey (cloud mode with no management key selected) — baseUrl and model are already Zod-gated, so that response could only ever fire on the empty key. The codex-settings route had diverged from the sibling CLI tools (cline/forge/ openclaw/grok-build/jcode): an inline if(!apiKey) 400 guard plus a hand-rolled getApiKeyById lookup, instead of the shared resolveApiKey(keyId, apiKey) helper which resolves by keyId, falls back to the submitted apiKey, then to sk_omniroute. This change makes codex-settings use the canonical resolver, so: - empty apiKey + valid keyId -> the real DB key is written to auth.json - empty apiKey + no keyId -> sk_omniroute default (config still applies) - explicit apiKey -> written verbatim (unchanged) * docs(changelog): add fragments for Codex turn-pin fallback and codex-settings apiKey resolution
Summary
Two independent Codex CLI fixes that were blocking long-running Codex sessions and the dashboard configure flow, shipped together because both belong to the same tool surface and both are verified by their own regression tests.
400 NATIVE_CODEX_PINNED_MODEL_UNAVAILABLEbricks a long session mid-turnapiKeythrough the canonical resolver400 baseUrl, apiKey and model are requiredon every dashboard Apply in cloud modeIssues addressed
baseUrl, apiKey and model are requiredwhen no management key is configured (cloud mode) #13563 — Codex settings still fail to configure whenapiKeyis empty./dashboard/logs/proxy). A fresh reproduction was added to fix(resilience): auto/coding collapses to the 200k-context target and fails with context_length_exceeded instead of the 1M model #12273 today. It is a separate dispatch-path change and will get its own PR.What was happening
1. Codex long sessions die to a terminal 400
Native Codex requests carry a turn pin: for the duration of one
turn_idthe combo locks provider + model. If that model becomes model-scoped unusable mid-turn (per-model quota lockout, connection cooldown, exhausted account), every continuation request for the same turn returned a non-retryable400 NATIVE_CODEX_PINNED_MODEL_UNAVAILABLE. Claude Code has no such pin, so it naturally falls back to the next healthy combo model on every request; Codex could not.2. Codex settings always 400 in cloud mode
The
codex-settingsroute diverged from every other CLI tool.baseUrlandmodelare Zod-gated (min(1)), so the400 "baseUrl, apiKey and model are required"could only fire on an emptyapiKey. The dashboard sends an emptyapiKeyexactly when cloud mode is on and no management key is selected — while the Apply button stays enabled. The route used an inlineif (!apiKey) return 400plus a hand-rolledgetApiKeyById, instead of the sharedresolveApiKey(keyId, apiKey)used bycline/forge/openclaw/grok-build/jcode-settings.Changes
Commit 1 —
fix(sse): release native Codex turn pin when pinned model is model-scoped unusable (#13564)open-sse/services/combo.ts— the pinned-turn branch now releases the pin and falls through to full combo routing when all pinned provider+model targets are model-scoped unusable. When the pinned target is not in the combo, it releases too. The pin is preserved when the provider is unhealthy at the provider level (circuit breaker OPEN, provider cooldown).open-sse/services/combo/nativeCodexTurnPin.ts— newreleaseNativeCodexTurnPin(body, comboName)helper (idempotent delete of the turn pin).tests/unit/native-codex-turn-pin-model-scoped-fallback.test.ts— rewritten regression suite: pin released + fallback to healthy sibling and re-pinned on success; retries stay pinned without flapping; pin preserved on circuit-breaker OPEN and on provider cooldown; new turns always get free combo routing.config/quality/eslint-suppressions.json— pruned the stalecombo.tsno-unused-varscount. This change orphaned thecreatePinnedModelUnavailableResponseimport; the residual unusedgetBootstrapLatencyMsis pre-existing.Commit 2 —
fix(api): resolve codex-settings apiKey via canonical resolver instead of 400 (#13563)src/app/api/cli-tools/codex-settings/route.ts— replaced the divergent guard + hand-rolled lookup with the canonicalresolveApiKey(keyId, validation.data.apiKey)from@/shared/services/apiKeyResolver, matching the sibling CLI tools.tests/unit/codex-settings-api-key-resolution.test.ts— new: emptyapiKey+ validkeyIdresolves the real DB key intoauth.json; emptyapiKey+ nokeyIdwrites thesk_omniroutedefault instead of 400; explicitapiKeyis still written verbatim.Design notes
finallyrelease inhandleComboChatInner, so there is no counter leak.keyId, then fall back to an explicit key, thensk_omniroute. Codex now behaves identically, so the dashboard Apply flow works in cloud mode without a pre-selected key.Verification
All tests below were run from the worktree root. Red-green cycles for both regression suites were observed before the fixes (the bug paths failed first).
Turn-pin suite
All phases pass (7/7). Wider pin regression sweep (39/39 across 6 other pin files) also green.
Codex-settings key resolution
3/3 pass (was 1/3 green, 2 failing, before the route change).
Adjacent regression
codex-settings-wire-api-default.test.tsstill passes (wire-api / TOML generation untouched).eslint(exit 0) and the pre-commit gate (prettier, eslint with suppressions, any-budget, tracked-artifacts) passed on both commits.Notes for reviewers
/dashboard/logs/proxystatus=errorflood is a real upstream-rejection symptom of fix(resilience): auto/coding collapses to the 200k-context target and fails with context_length_exceeded instead of the 1M model #12273, not a separate proxy-log bug. That fix is scoped as a follow-up PR.Branch synced to the current
release/v3.8.51tip (base-red #12732 since resolved); mergeable, checks re-running.