From 3be3e55ffd4d0282bfa727cd9dcb315ebc2a15d9 Mon Sep 17 00:00:00 2001 From: JUN Date: Tue, 25 Aug 2026 12:17:08 +0900 Subject: [PATCH 01/29] fix(management): validate the sidecar pair against the submitted backend (#2514) * devlog: operator visibility train roadmap unit (260825) Docs-only roadmap for three operator-visibility defects that share one shape: OpenCodex computes the truth and does not report it. - 010 (#2457): the sidecar pair check collapses a five-member union into a two-arm ternary, so a submitted gemini backend is validated against the stored openai backend and the dashboard save 400s. - 020 (#2411): collectStatus already computes routingKind and ships it in status --json, but the human renderer never prints it, so a healthy proxy reads green while nothing routes through it. - 030 (#2412): a version-manager overwrite leaves a message-less ineligible verdict, and the CLI warns only when a message exists. 002 records the plan audit, including two amendments: WP4 must plumb messages to every reachable silent ineligible return rather than only the one the reporter hit, and WP3 must pin custom-local/unknown as intentionally silent. * fix(management): validate the sidecar pair against the submitted backend Both web-search sidecar writers resolved the backend for the model pair check with a ternary that handled only openai, anthropic, and null, then fell back to the STORED backend for everything else. The accepted union has five members, so a submitted `gemini` never matched an arm and was validated against whatever was already saved. That is two defects, not one. With a stored openai backend the dashboard could not save a Gemini pair the picker itself offered (#2457, the reported 400), and symmetrically `backend=gemini` with `model=gpt-5.6-luna` was accepted with 200 because it was checked against openai. Writing the same pair directly into config.json always worked, which is what narrowed this to the write gate. Both routes now resolve the submitted backend across the whole union. The union check above each site has already rejected unknown literals, so a surviving string is a member; Array.includes does not narrow, so the cast carries that proof. The two null policies are deliberately NOT unified: on /api/sidecar-settings null unsets the backend and unset resolves to openai, while on /api/claude-code null drops the override and inherits the global backend. Collapsing them into one helper would silently change what clearing a Claude override means. The executor is untouched. resolveSidecarBackend and planWebSearch already handled Gemini, and tests/gemini-web-search.test.ts still passes unchanged. Closes #2457 --- .../000_baseline_and_scope.md | 72 +++++++++ .../001_current_state_inventory.md | 130 ++++++++++++++++ .../002_plan_audit.md | 90 +++++++++++ ...p2_issue2457_sidecar_backend_resolution.md | 129 ++++++++++++++++ ...wp3_issue2411_status_routing_visibility.md | 145 ++++++++++++++++++ .../030_wp4_issue2412_version_manager_shim.md | 142 +++++++++++++++++ .../management/agent-settings-routes.ts | 20 ++- src/server/management/config-routes.ts | 15 +- .../sidecar-settings-web-search-gate.test.ts | 122 +++++++++++++++ 9 files changed, 855 insertions(+), 10 deletions(-) create mode 100644 devlog/_plan/260825_operator_visibility_train/000_baseline_and_scope.md create mode 100644 devlog/_plan/260825_operator_visibility_train/001_current_state_inventory.md create mode 100644 devlog/_plan/260825_operator_visibility_train/002_plan_audit.md create mode 100644 devlog/_plan/260825_operator_visibility_train/010_wp2_issue2457_sidecar_backend_resolution.md create mode 100644 devlog/_plan/260825_operator_visibility_train/020_wp3_issue2411_status_routing_visibility.md create mode 100644 devlog/_plan/260825_operator_visibility_train/030_wp4_issue2412_version_manager_shim.md diff --git a/devlog/_plan/260825_operator_visibility_train/000_baseline_and_scope.md b/devlog/_plan/260825_operator_visibility_train/000_baseline_and_scope.md new file mode 100644 index 00000000000..542f15a1943 --- /dev/null +++ b/devlog/_plan/260825_operator_visibility_train/000_baseline_and_scope.md @@ -0,0 +1,72 @@ +# 000 — Operator visibility train: baseline, scope, and work-phase map + +Unit opened 2026-08-25. Session `01a03688-c5ee-76c2-bb0f-a7a9213345d5`. +Goalplan slug `fix-three-opencodex-operator-visibility-defects`. + +## Baseline + +Verified live at unit open, immediately after the v2.32.1 publish: + +| Ref | SHA | Meaning | +|-----|-----|---------| +| `origin/dev` | `bb89eafbe` | devlog: pin the report to the code SHA its gates describe (#2506) | +| `origin/main` | `71c57ea64` | `release: v2.32.1` | +| `origin/preview` | `f4cb9f800` | `release: v2.32.1-preview.20260825` | + +`git merge-base --is-ancestor origin/dev origin/main` exits 0, so `dev` is an +ancestor of the shipped release and this unit starts from published code. +npm `latest` is `2.32.1`, `preview` is `2.32.1-preview.20260825`. + +## What this unit is + +Three defects that share one shape: **OpenCodex knows the truth and does not +tell the operator.** None of them is a routing or execution bug. In all three +the runtime is already correct and the surface that reports to a human is +wrong, stale, or silent. + +| # | Surface | The lie | +|---|---------|---------| +| #2457 | Management write | The picker offers Gemini, then the save rejects it as an OpenAI model | +| #2411 | `ocx status` | Green proxy while nothing routes through it | +| #2412 | Shim auto-restore | A destroyed shim returns an ineligible verdict with no message | + +That shared shape is why they travel together and why none of them may be +"fixed" by changing behavior. Every fix in this unit is a reporting fix. + +## Work-phase map + +| Phase | Doc | Issue | Deliverable | +|-------|-----|-------|-------------| +| WP1 | this unit | — | Docs-only roadmap at diff-level precision | +| WP2 | `010` | #2457 | Submitted backend is what the pair check validates | +| WP3 | `020` | #2411 | `ocx status` prints routing and warns on unused proxy | +| WP4 | `030` | #2412 | Version-manager shim destruction is detected and reported | + +One work-phase is one full PABCD cycle. WP2, WP3, and WP4 each produce one PR +against `dev`. + +## Scope boundary + +Out of scope, stated once so no later phase reopens it: + +- Merging other contributors' PRs, or another npm release. +- `src/lab/` — the core-lab boundary test exists for a reason. +- The undeclared-tool guard, and any auth, OAuth, credential, workflow, or + release-automation surface. +- Auto-wrapping a version-manager-owned `codex` binary as a new original. + This is the one that is tempting and wrong; see `030`. +- The Codex-side namespaced-model error message in #2411's reproduction. That + is upstream copy, not ours. + +## Evidence rule + +A remembered pass is not evidence. Every completion claim in this unit carries +exact command output, the PR number and head SHA, and the CI run id and +conclusion on that SHA. + +## Prior art consulted + +- `260824_v2_32_1_hotfix_train/` — the freeze/GO discipline this unit inherits. +- `tests/repo-hygiene.test.ts` — no gitlinks, no vendored clones. +- `AGENTS.md` — focused checks during implementation, full suite before a + non-trivial PR goes review-ready. diff --git a/devlog/_plan/260825_operator_visibility_train/001_current_state_inventory.md b/devlog/_plan/260825_operator_visibility_train/001_current_state_inventory.md new file mode 100644 index 00000000000..611edb5af94 --- /dev/null +++ b/devlog/_plan/260825_operator_visibility_train/001_current_state_inventory.md @@ -0,0 +1,130 @@ +# 001 — Current-state inventory + +Read at `bb89eafbe`. Every line anchor below was opened and read, not inferred. + +## #2457 — the pair check discards the union + +The accepted union is complete. `src/server/management/config-routes.ts:591`: + +```ts +const WEB_SEARCH_BACKENDS_UNION = ["openai", "anthropic", "xai", "gemini", "exa"] as const; +``` + +The pair check nineteen lines later throws it away. `config-routes.ts:668`: + +```ts +const effectiveBackend = body.webSearch.backend === "anthropic" + ? "anthropic" + : body.webSearch.backend === "openai" || body.webSearch.backend === null + ? "openai" + : config.webSearchSidecar?.backend ?? "openai"; +``` + +A submitted `"gemini"` is not `"anthropic"`, not `"openai"`, not `null`. +It falls to the final arm and the request is validated against the **stored** +backend. With stored `openai` (or unset), `webSearchModelIsRejected("openai", +"gemini-3.7-flash", candidates)` is true, and the route returns 400 before the +persistence block at `:687` — which does honor the full union — ever runs. + +`src/server/management/agent-settings-routes.ts:1121` carries the same stale +ternary with a different null policy: + +```ts +const effectiveBackend = section.backend === "anthropic" + ? "anthropic" + : section.backend === "openai" + ? "openai" + : section.backend === null + ? config.webSearchSidecar?.backend ?? "openai" + : stored?.backend ?? config.webSearchSidecar?.backend ?? "openai"; +``` + +The comment directly above that block reads: *"Same module as +/api/sidecar-settings — a gate on one route and a stale copy on the other is no +gate at all."* The gate is shared; the backend resolution is not, and it drifted +exactly as the comment feared. + +`xai` and `exa` have the identical hole. They escape notice because +backend-only submissions short-circuit on `effectiveModel` being empty. + +The executor is already correct and must not be touched: +`resolveSidecarBackend("gemini")` returns `"gemini"` +(`src/web-search/index.ts:162`), and `planWebSearch` already defaults Gemini to +`gemini-3.7-flash` (`:285`). Writing the pair directly into `config.json` +works today, which is the reporter's own proof that only the write gate is wrong. + +## #2411 — status has the routing kind and never prints it + +`collectStatus()` already computes it. `src/cli/status.ts:188`: + +```ts +const startup = collectStartupHealth(config, { + service, + shim: codexShim, + routingKind: getCodexRoutingKind(), +}); +``` + +`startup` lands on `json.startup` at `src/cli/status.ts:316`, so +`ocx status --json` **already exposes** `startup.routingKind`. The human +renderer is what drops it. `src/cli/index.ts:845`: + +```ts +if (status.json.proxy.pid || status.json.proxy.health.ok) { + console.log(`✅ Proxy: ${status.proxyLabel}`); +} +``` + +That boolean never consults `startup.routingKind`. A live PID or a good +`/healthz` is sufficient for the green check. + +Worse, the next line reinforces it. `startupHealthSummary` +(`src/codex/autostart-health.ts:143`) renders native routing as *"native Codex +routing (no opencodex restart dependency)"*, and `deriveStartupHealth` marks it +`rebootSafe: true`. That is correct on its own terms — there is genuinely no +restart dependency when nothing routes — but printed under a green proxy it +reads as a second all-clear. + +`ocx doctor` already prints the missing token. `src/cli/doctor.ts:986`: + +```ts +console.log(` routing=${startup.routingKind}, service=${...}, shim=${...}`); +``` + +So the fix is not new computation. It is routing the value that already exists +to the surface people actually run. + +## #2412 — the ineligible verdict carries no message + +`src/codex/shim.ts:2043`: + +```ts +if (!existsSync(file.wrapperPath) || !hasUsableBackingPath(file)) return { status: "ineligible" }; +``` + +No `message` field. That is why the condition is invisible: the CLI warns only +when one exists. `src/cli/codex-shim-autorestore.ts:35`: + +```ts +} else if ((result.status === "deferred" || result.status === "ineligible") && result.message) { + deps.warn(`⚠️ ${result.message}`); +} +``` + +A mise/asdf/volta upgrade rewrites the install tree in place, destroying both +`codex` and its sibling `codex.opencodex-real` (`backupPathFor`, +`src/codex/shim.ts:601`). `hasUsableBackingPath` (`:481`) then returns false, +the silent ineligible fires, and `ocx start` / `ocx ensure` / +`ocx service repair` all proceed to report success. + +`diagnoseCodexShim` (`src/codex/shim.ts:2156`) already produces the exact +diagnostic string the reporter pasted. The information exists; nothing routes it +to the commands that matter. + +## The common root + +In all three, the correct value is computed and then discarded on the way to the +human: a validated union collapsed into a two-arm ternary, a routing kind +carried in JSON but not printed, a diagnosis produced by one command and absent +from three others. None of the three fixes changes what OpenCodex does. They +change what it admits. diff --git a/devlog/_plan/260825_operator_visibility_train/002_plan_audit.md b/devlog/_plan/260825_operator_visibility_train/002_plan_audit.md new file mode 100644 index 00000000000..2e3882fa591 --- /dev/null +++ b/devlog/_plan/260825_operator_visibility_train/002_plan_audit.md @@ -0,0 +1,90 @@ +# 002 — Plan audit (A phase, WP1) + +The dispatched read-only auditor produced nothing across four wait cycles and +was retired under the loop's failed-dispatch rule. The audit below was performed +directly by the main agent against source at `bb89eafbe`. Every anchor cited in +`001`, `010`, `020`, and `030` was re-opened and confirmed. + +## Anchor verification + +| Doc claim | Verified | +|-----------|----------| +| `config-routes.ts:591` union of five backends | yes, exact | +| `config-routes.ts:668` two-arm ternary falling back to stored | yes, exact | +| `config-routes.ts:688` persistence honors the full union | yes | +| `agent-settings-routes.ts:1121` five-arm ternary | yes | +| `cli/status.ts:188` computes `routingKind` | yes | +| `cli/status.ts:316` `startup` lands in JSON | yes | +| `cli/index.ts:845` green check ignores routing | yes | +| `cli/doctor.ts:986` prints `routing=` | yes | +| `shim.ts:481` `hasUsableBackingPath` | yes | +| `shim.ts:1887` `allowFreshInstall` guard | yes | +| `cli/codex-shim-autorestore.ts:35` warns only with a message | yes | +| `autostart-health.ts:143` `startupHealthSummary` | yes | + +One correction: `030` cites the destroyed-shim bail as `shim.ts:2043`. The +actual line is **`2045`**; `2043` is inside the `preserveOnly` branch. The +quoted code is right, the number is off by two. + +## Blocking findings + +**A1 — `030` targets only one of six `ineligible` returns.** +`rg 'status: "ineligible"' src/codex/shim.ts` finds returns at `2028`, `2031`, +`2039`, `2042`, `2045`, `2049`, and `2085`. Only `2028` and `2085` carry a +message today. The plan attaches one to `2045`, but `2042` is the +`preserveOnly` sibling case and `2049` is `isHealthyShimProbe` — both are +reachable in a version-manager overwrite and both would stay silent. + +Correction: WP4 must attach messages to the reachable silent returns, not just +the one the reporter happened to hit. The `preserveOnly` branch at `2042` +deserves its own wording — its condition is a missing backup **or** a resurrected +original, which is a different story from a destroyed wrapper. + +**A2 — `020`'s truth table omits `custom-local` and `unknown`.** +`CodexRoutingKind` (`inject.ts:314`) has five members. The table covers +`opencodex-local`, `native`, and `custom-remote`. The predicate as written +returns `[]` for `custom-local` and `unknown`, which is the correct behavior — +`startupHealthSummary` already renders both as `AT RISK after restart` with a +remedy command (`autostart-health.ts:149-150`), so a second warning would be +noise. But the plan does not say so, and a later reader could "fix" the omission. + +Correction: state the five-member coverage explicitly and record that +`custom-local`/`unknown` are intentionally silent **because** they are already +loud elsewhere. Add both to the helper's test cases so the intent is pinned. + +## Non-blocking findings + +**B1 — `010`'s cast.** `WEB_SEARCH_BACKENDS_UNION.includes(x as ...)` does not +narrow `x` in TypeScript; `includes` returns `boolean`, not a type predicate. +The proposed `submittedBackend as typeof WEB_SEARCH_BACKENDS_UNION[number]` +cast in the true arm is therefore load-bearing, not decorative. It is sound +because `:591` already rejected non-members, but the doc should say that the +cast is doing real work rather than reading as noise. + +**B2 — `webSearchModelIsRejected`'s `backend` parameter type.** If it is typed +as the narrow union, passing the widened value type-checks only because both +resolve to the same union. Confirm at implementation time; if it is narrower, +the signature is the thing to widen, not the call site to cast. + +**B3 — line-number drift.** `030` says `2043`, actual `2045`. Corrected in this +document rather than by rewriting `030`, so the drift stays visible. + +## Verified correct + +- The #2457 mechanism, end to end: union at `:591`, ternary at `:668`, + persistence at `:688`. A submitted `gemini` provably reaches the stored-backend + arm. +- Both null policies genuinely differ between the two routes. `010`'s refusal to + unify them is right. +- `startup.routingKind` is already in `status --json`. `020`'s claim that no + schema change is needed holds. +- `allowFreshInstall: false` at `1887` is the invariant that blocks adoption. + `030`'s refusal to relax it is correct, and it is what makes A1 a + message-plumbing fix rather than a behavior change. + +## Verdict + +**PASS with two required amendments.** A1 and A2 are corrections to WP4 and WP3 +scope respectively; neither invalidates the plan's shape, and both are folded +into this document rather than silently patched into the originals. B1–B3 are +notes for the implementer. diff --git a/devlog/_plan/260825_operator_visibility_train/010_wp2_issue2457_sidecar_backend_resolution.md b/devlog/_plan/260825_operator_visibility_train/010_wp2_issue2457_sidecar_backend_resolution.md new file mode 100644 index 00000000000..dd91d9f3037 --- /dev/null +++ b/devlog/_plan/260825_operator_visibility_train/010_wp2_issue2457_sidecar_backend_resolution.md @@ -0,0 +1,129 @@ +# 010 — WP2: the submitted sidecar backend is what the pair check validates (#2457) + +## The change in one sentence + +Both management write paths must validate the requested model against the +**backend the caller submitted**, not against a two-member subset with the +stored backend as fallback. + +## Hunk 1 — `src/server/management/config-routes.ts` (`PUT /api/sidecar-settings`) + +Before, at `:668`: + +```ts +const effectiveBackend = body.webSearch.backend === "anthropic" + ? "anthropic" + : body.webSearch.backend === "openai" || body.webSearch.backend === null + ? "openai" + : config.webSearchSidecar?.backend ?? "openai"; +``` + +After: + +```ts +const submittedBackend = body.webSearch.backend; +const effectiveBackend = + typeof submittedBackend === "string" + && WEB_SEARCH_BACKENDS_UNION.includes(submittedBackend as typeof WEB_SEARCH_BACKENDS_UNION[number]) + ? submittedBackend as typeof WEB_SEARCH_BACKENDS_UNION[number] + : submittedBackend === null + ? "openai" + : config.webSearchSidecar?.backend ?? "openai"; +``` + +`WEB_SEARCH_BACKENDS_UNION` is already in scope at `:591`; an unknown literal +was already rejected there, so by this point a string is either a union member +or the request is dead. + +## Hunk 2 — `src/server/management/agent-settings-routes.ts` (`PUT /api/claude-code`) + +Before, at `:1121`: the five-arm ternary quoted in `001`. + +After, reusing the local `allowedBackends` built at `:1081`: + +```ts +const submittedBackend = section.backend; +const effectiveBackend = + typeof submittedBackend === "string" && allowedBackends.includes(submittedBackend) + ? submittedBackend as WebSearchBackend + : submittedBackend === null + ? config.webSearchSidecar?.backend ?? "openai" + : stored?.backend ?? config.webSearchSidecar?.backend ?? "openai"; +``` + +## The two null policies are different and both stay + +This is the part a careless fix breaks. They are not the same rule: + +| Route | `backend: null` means | Resolves to | +|-------|------------------------|-------------| +| `/api/sidecar-settings` | unset the global backend | `"openai"` (the resolver's own default for unset) | +| `/api/claude-code` | drop the Claude override | inherit `config.webSearchSidecar?.backend ?? "openai"` | + +Do not unify them. A shared helper that collapses both to one fallback would +silently change what clearing the Claude override means. + +## Shape decision + +Two shapes were considered: + +- **A (chosen):** inline the union membership check in both writers. +- **B:** extract `submittedWebSearchBackend()` into + `web-search-sidecar-options.ts`. + +B reads better as drift protection, which is exactly what failed here. But the +two null policies above cannot live in one helper, so B would extract only the +string arm and leave the divergent part behind — the appearance of unification +without the substance. A is five lines per route with the union named locally. +If a reviewer prefers B, the helper must take the null fallback as a parameter. + +## What must NOT change + +- `webSearchModelIsRejected` / `webSearchModelRejection` + (`src/server/management/web-search-sidecar-options.ts:91`). The helper is + correct; only its `backend` argument was wrong. +- The runtime executor: `src/web-search/index.ts`, `src/web-search/backends.ts`. +- The raw `config.json` escape hatch, which deliberately skips this gate. +- Vision sidecar validation, which has a different three-member union ending in + `"routed"`, not `"exa"`. +- `GET /api/sidecar-settings` and its `webSearchModels` rows. + +## Must still return 400 after the fix + +These are the assertions that prove the gate was not merely widened: + +1. `{ backend: "openai", model: "claude-haiku-4-5" }` — real mismatch. +2. `{ model: "gemini-3.7-flash" }` with backend omitted and stored `openai` — + preserved-backend semantics survive. +3. `{ backend: "gemini", model: "gpt-5.6-luna" }` — inverse mismatch. +4. `{ backend: "zen" }` — still fails the union gate at `:591`. + +## Regression tests + +All in `tests/sidecar-settings-web-search-gate.test.ts`, which already mocks +`getAccountSet` and `listManagementModelRows`. A Gemini pair placed in +`tests/web-search-backend-union.test.ts` would still be rejected after the fix +because that file has no candidate rows — the pair check would correctly find no +matching row. Wrong file, false failure. + +| Test | Setup | Assertion | Fails before? | +|------|-------|-----------|---------------| +| `PUT persists openai/luna -> gemini/gemini-3.7-flash` | stored `{openai, gpt-5.6-luna}`, `google-antigravity` oauth + healthy account set with `projectId`, management row `gemini-3.7-flash` | 200, config holds the Gemini pair | **Yes** — 400 today | +| `each leftover union member persists its own pair` (`test.each(["xai","gemini"])`) | matching candidate per backend | 200 each | **Yes** | +| `omitted backend still validates against the stored backend` | Gemini row live, PUT model only | 400, stored pair unchanged | No — guards the fix | +| `PUT /api/claude-code persists a gemini override` | stored override `{openai, gpt-5.6-luna}` | 200, `claudeCode.webSearchSidecar` is the Gemini pair | **Yes** — 400 today | + +## Existing tests that must stay green + +- `PUT rejects a backend/model mismatch and does not persist it` (`:139`) +- `PUT validates a backend-only update against the preserved effective model` (`:150`) +- `PUT persists the Anthropic auth-slot pair exactly as offered` (`:160`) +- `tests/claude-management-api.test.ts` sidecar round-trip (`:370`) +- `tests/gemini-web-search.test.ts` executor plan test (`:75`) — untouched, and + its continued passing is the proof the executor needed no change. + +## Acceptance + +`bun test tests/sidecar-settings-web-search-gate.test.ts tests/web-search-backend-union.test.ts tests/claude-management-api.test.ts tests/gemini-web-search.test.ts` +green; `bun x tsc --noEmit` exit 0; `bun run privacy:scan` pass; new tests +demonstrated red before the patch. diff --git a/devlog/_plan/260825_operator_visibility_train/020_wp3_issue2411_status_routing_visibility.md b/devlog/_plan/260825_operator_visibility_train/020_wp3_issue2411_status_routing_visibility.md new file mode 100644 index 00000000000..fbf78ea2756 --- /dev/null +++ b/devlog/_plan/260825_operator_visibility_train/020_wp3_issue2411_status_routing_visibility.md @@ -0,0 +1,145 @@ +# 020 — WP3: `ocx status` reports routing and warns on an unused proxy (#2411) + +## The change in one sentence + +`ocx status` prints the routing kind it already computes, and says so plainly +when a healthy proxy is paired with native routing. + +## The design question, settled + +Two shapes: + +- **A (chosen):** keep `✅` on the proxy line, always print `routing=`, and add + a warning only for the healthy-proxy + native-routing combination. +- **B:** flip the first line to `⚠️` for that combination. + +B is tempting because the reporter's complaint is literally "the check is +green." But the proxy line makes a narrow claim — the process is up and +`/healthz` answered — and that claim is **true** in this state. The reporter +proved it himself by curling the proxy directly and getting `ok`. Turning that +line yellow would make the one honest signal lie in order to compensate for a +missing one. It also collides with the existing `❌` path, whose remedy text +("Restart with 'ocx start'") is wrong for this failure: the proxy does not need +restarting, Codex needs re-pointing. + +So: add the missing signal, do not corrupt the present one. + +## Hunk 1 — extract the routing detail so status and doctor cannot drift + +`src/codex/autostart-health.ts`, next to `startupHealthSummary` at `:143`: + +```ts +export function formatStartupRoutingDetail(health: StartupHealth): string { + const service = health.serviceViable + ? "viable" + : health.serviceInstalled ? "installed-but-unhealthy" : "absent"; + const shim = health.shimHealthy + ? "healthy" + : health.shimInstalled ? "stale" : "absent"; + return `routing=${health.routingKind}, service=${service}, shim=${shim}`; +} +``` + +`src/cli/doctor.ts:986` then becomes a call to it, emitting byte-identical +output. This matters: #2457 exists because two routes computed the same thing +separately and drifted. Do not introduce a second copy of doctor's line. + +## Hunk 2 — the warning predicate + +`src/cli/status.ts`, pure and exported for direct testing, in the manner of +`src/cli/status-oauth.ts:55`: + +```ts +export function unusedProxyWarningLines(input: { + proxyUp: boolean; + routingKind: StartupHealth["routingKind"]; +}): string[] { + if (!input.proxyUp || input.routingKind !== "native") return []; + return [ + "⚠️ Codex routing is native — the running proxy is unused.", + " Codex requests go to OpenAI, not this proxy. Re-point with: ocx restore back", + ]; +} +``` + +A pure function is the point: the interesting behavior is a two-input truth +table, and it should be testable without spawning a CLI. + +## Hunk 3 — render + +`src/cli/index.ts`, after the Health line at `:850`: + +```ts +const proxyUp = Boolean(status.json.proxy.pid || status.json.proxy.health.ok); +for (const line of unusedProxyWarningLines({ + proxyUp, + routingKind: status.json.startup.routingKind, +})) { + console.log(` ${line}`); +} +``` + +and after `Restart safety` at `:869`: + +```ts +console.log(` ${formatStartupRoutingDetail(status.json.startup)}`); +``` + +Placing the routing detail directly under restart safety is deliberate. That +summary line is the one that reads as a second all-clear ("no opencodex restart +dependency"); the routing token immediately below it supplies the missing +context for why there is no dependency. + +## Truth table + +| Proxy | Routing | First line | Warning | `routing=` | +|-------|---------|-----------|---------|------------| +| up | `opencodex-local` | ✅ | no | yes | +| up | `native` | ✅ | **yes** | yes | +| up | `custom-remote` | ✅ | no | yes | +| down | `native` | ❌ | no | yes | + +`custom-local` / `custom-remote` are also "this proxy is unused," but they are +a deliberate operator choice and `startupHealthSummary` already names them as a +remote gateway. Warning there would train people to ignore the warning. Native +is the accidental state, and the only one #2411 reports. + +Proxy down plus native routing must not warn: the operator has two problems and +the `❌` line with its restart remedy is the correct lead. + +## JSON + +No schema change, no `schemaVersion` bump. `startup.routingKind` is already in +the payload — the gap was never the data. Adding a derived +`proxyUnusedByCodex` boolean was considered and rejected: consumers can +combine two fields they already have, and `tests/cli-status-json.test.ts:21` +pins `schemaVersion === 1`. + +## What must NOT change + +- `classifyCodexRouting`, `getCodexRoutingKind`, `deriveStartupHealth`, + `startupHealthSummary`. This phase reads them; it does not touch them. +- `rebootSafe: true` for native routing. `tests/autostart-health.test.ts:108` + pins it, and it is correct: there really is no restart dependency. +- The `❌` branch and its `ocx start` / `ocx service repair` guidance. +- Redaction behavior of `status --json`. +- Anything in #2412's shim territory. The two issues are related as cause and + symptom but ship as separate PRs, per the maintainer's own split. + +## Regression tests + +| Test | File | Assertion | Fails before? | +|------|------|-----------|---------------| +| `unusedProxyWarningLines covers the four routing states` | `tests/cli-status-json.test.ts` | the truth table above | **Yes** — helper absent | +| `status prints routing=native without starting the proxy` | `tests/cli-help.test.ts` (extend `:139`) | stdout has `routing=native`, and does **not** have the unused-proxy warning while the proxy is down | **Yes** | +| `status --json exposes startup.routingKind` | `tests/cli-status-json.test.ts` | `parsed.startup.routingKind === "native"` | No — pins existing data against future removal | +| `formatStartupRoutingDetail matches doctor's line` | `tests/autostart-health.test.ts` | `routing=native, service=absent, shim=absent` | **Yes** | + +The CLI tests need a temp `CODEX_HOME` holding a `config.toml` without +`openai_base_url`; `tests/codex-plugins-doctor.test.ts:356` is the pattern. + +## Acceptance + +`bun test tests/cli-status-json.test.ts tests/cli-help.test.ts tests/autostart-health.test.ts tests/codex-plugins-doctor.test.ts` +green; `bun x tsc --noEmit` exit 0; `bun run privacy:scan` pass; doctor's +output byte-identical before and after the extraction. diff --git a/devlog/_plan/260825_operator_visibility_train/030_wp4_issue2412_version_manager_shim.md b/devlog/_plan/260825_operator_visibility_train/030_wp4_issue2412_version_manager_shim.md new file mode 100644 index 00000000000..561197decf8 --- /dev/null +++ b/devlog/_plan/260825_operator_visibility_train/030_wp4_issue2412_version_manager_shim.md @@ -0,0 +1,142 @@ +# 030 — WP4: detect and report version-manager shim destruction (#2412) + +## The change in one sentence + +When a version manager has overwritten the shim and its backup, say so with an +actionable message — and refuse to adopt the new binary as a replacement +original. + +## The temptation, and why it is wrong + +The obvious fix is to make auto-restore work: a backup is missing, so take the +current `codex` binary, rename it to `codex.opencodex-real`, and write a fresh +shim over it. It would make the symptom disappear immediately. + +It is wrong twice over. + +First, it is a lie about provenance. The binary now sitting at that path is the +version manager's newly installed `codex`, not the original OpenCodex wrapped. +Recording it as `.opencodex-real` asserts a history that did not happen. + +Second, it does not survive. The next `mise upgrade codex` rewrites the same +install tree and destroys shim and backup again. The fix would re-arm itself +every upgrade, so the operator gets a repair that silently un-repairs on a +schedule — the worst possible failure shape, because it looks solved. + +The install tree belongs to the version manager. OpenCodex should not be +installing files into it, and the supported route for these users is +`openai_base_url` injection plus `ocx service install`, which is what +`ocx start` already configures. + +So: detect, report, document. Never adopt. + +## Hunk 1 — the ownership heuristic + +`src/codex/shim.ts`, exported for direct unit tests: + +```ts +export function isVersionManagerOwnedCodexPath(path: string): boolean { + const n = path.replace(/\\/g, "/").toLowerCase(); + return n.includes("/mise/installs/") || n.includes("/mise/shims/") + || n.includes("/.asdf/installs/") || n.includes("/.asdf/shims/") + || n.includes("/.volta/"); +} +``` + +Backslash normalization is for Windows, where volta is common. Scope is the +three managers named in #2412; nvm/fnm/npm-prefix are deliberately excluded +until someone reports them, because a false positive here refuses a repair that +would otherwise be correct. + +## Hunk 2 — carry a message, and refuse VM-owned adoption + +`src/codex/shim.ts:2043`, before: + +```ts +if (!existsSync(file.wrapperPath) || !hasUsableBackingPath(file)) return { status: "ineligible" }; +``` + +After: compute `vmOwned` across wrapper/original/backup paths, include it in the +bail condition, and attach a message built from +`diagnoseCodexShim().summary` — the string `ocx codex-shim status` already +prints — plus, when `vmOwned`, this guidance: + +> This Codex binary is owned by a version manager (mise/asdf/volta). OpenCodex +> will not wrap it as a new original, because the next upgrade would overwrite +> the shim and its backup again. Keep routing through Codex `openai_base_url` +> (`ocx start`) and use `ocx service install` for autostart. + +The replacement path at `:2076` needs the same guard. If a stale +`.opencodex-real` happens to survive an upgrade, the existing code would +cheerfully re-wrap the new version-manager binary — the adoption this phase +forbids, arriving through the back door. + +## Hunk 3 — no CLI changes needed for start/ensure/repair + +This is the satisfying part. `src/cli/codex-shim-autorestore.ts:35` already +warns on an ineligible result **if it carries a message**: + +```ts +} else if ((result.status === "deferred" || result.status === "ineligible") && result.message) { + deps.warn(`⚠️ ${result.message}`); +} +``` + +and `src/cli/root.ts:83` runs that preflight before every command except +uninstall and `codex-shim install`. So attaching the message lights up +`ocx start`, `ocx ensure`, `ocx service repair`, and `ocx status` at once. +The mechanism was built correctly; one field was missing. + +## Hunk 4 — docs + +`docs-site/src/content/docs/reference/cli/lifecycle.md`, after the paragraph at +~`:357` promising that a completed Codex update restores the shim. That promise +is false for version-manager installs, and leaving it unqualified is how someone +concludes OpenCodex is broken rather than unsupported here. State plainly: the +install tree is not a supported shim target, upgrades destroy shim and backup, +and the supported configuration is service + `openai_base_url`. + +English is authoritative; translated locales must not keep promising restore for +this case. + +## What must NOT change + +- Healthy shims stay `{ status: "healthy" }` on the zero-overhead path + (`:2058`), including version-manager-owned ones that are currently intact. + Detection gates repair, not operation. +- Non-VM overwrite with a surviving backup still auto-restores and still warns + "automatic repair after Codex update". +- `allowFreshInstall: false`. The never-fresh-install rule at `:1887` is the + invariant this phase reinforces, not one it relaxes. +- `repairService()` semantics. It reports on the background service, and that + report is accurate; the shim warning arrives from the preflight instead. +- The first-line proxy badge. That is #2411's territory. + +## Regression tests + +| Test | File | Assertion | Fails before? | +|------|------|-----------|---------------| +| `version-manager overwrite with missing backup is ineligible and names the paths` | `tests/codex-shim.test.ts` | `ineligible` **with** a message naming wrapper state, missing backup, and the version manager; wrapper bytes unchanged | **Yes** — message is undefined | +| `version-manager-owned replacement is not adopted as a new original` | `tests/codex-shim.test.ts` | backup present but VM-owned path → ineligible; wrapper, backup, and state bytes all unchanged | **Yes** — today this restores | +| `ineligible destroyed shim warns on ordinary commands` | `tests/codex-shim-autorestore.test.ts` | one `⚠️` containing the diagnostic | **Yes** | +| `isVersionManagerOwnedCodexPath classifies known trees` | `tests/codex-shim.test.ts` | mise/asdf/volta true; `/usr/local/bin/codex`, `~/.npm-global/bin/codex` false | **Yes** — helper absent | + +`tests/codex-shim.test.ts:1921` (`missing backup, missing wrapper, corrupt +state, and platform mismatch never fresh-install`) asserts only on `status`, so +adding a message does not break it — and it is the test that would catch an +adoption regression. + +## Acceptance + +`bun test tests/codex-shim.test.ts tests/codex-shim-autorestore.test.ts tests/codex-shim-readiness.test.ts` +green; `bun x tsc --noEmit` exit 0; `bun run privacy:scan` pass; docs build not +required for a Markdown-only change but the page must render in review. + +## Open question for review + +Explicit `ocx codex-shim install` against a version-manager-owned PATH: +warn-and-allow, or refuse outright? Auto-restore must refuse — that is settled +above and is what this issue asks for. An explicit operator command is a +different act. Recommendation: warn, allow, and let the operator own it; a hard +refusal removes a workaround someone may be relying on. This does not block the +phase either way. diff --git a/src/server/management/agent-settings-routes.ts b/src/server/management/agent-settings-routes.ts index b1cbb8cd44b..bddbd37fe14 100644 --- a/src/server/management/agent-settings-routes.ts +++ b/src/server/management/agent-settings-routes.ts @@ -60,6 +60,7 @@ import { webSearchCandidateRows, webSearchModelIsRejected, webSearchModelRejection, + type WebSearchBackend, } from "./web-search-sidecar-options"; import { drainAndShutdown } from "../lifecycle"; import { filterRequestLogs, getRequestLogEntries, type RequestLogEntry } from "../request-log"; @@ -1118,13 +1119,18 @@ export async function handleAgentSettingsRoutes(ctx: ManagementContext): Promise if (field === "webSearchSidecar" && (section.model !== undefined || section.backend !== undefined)) { const stored = config.claudeCode?.webSearchSidecar; - const effectiveBackend = section.backend === "anthropic" - ? "anthropic" - : section.backend === "openai" - ? "openai" - : section.backend === null - ? config.webSearchSidecar?.backend ?? "openai" - : stored?.backend ?? config.webSearchSidecar?.backend ?? "openai"; + // Validate against the SUBMITTED backend across the whole union, not + // just openai/anthropic (#2457). allowedBackends above already refused + // unknown literals; Array.includes does not narrow, hence the cast. + // null keeps its own meaning here — drop the override and inherit the + // global backend — which is deliberately NOT the sidecar-settings rule. + const submittedBackend = section.backend; + const effectiveBackend = typeof submittedBackend === "string" + && allowedBackends.includes(submittedBackend) + ? submittedBackend as WebSearchBackend + : submittedBackend === null + ? config.webSearchSidecar?.backend ?? "openai" + : stored?.backend ?? config.webSearchSidecar?.backend ?? "openai"; const effectiveModel = section.model === "" ? config.webSearchSidecar?.model : typeof section.model === "string" diff --git a/src/server/management/config-routes.ts b/src/server/management/config-routes.ts index e54cb5a3c04..d4d1cb72c6a 100644 --- a/src/server/management/config-routes.ts +++ b/src/server/management/config-routes.ts @@ -665,9 +665,18 @@ export async function handleConfigRoutes(ctx: ManagementContext): Promise { }); }); +// #2457: the picker offers every union backend, but the pair check used to +// collapse the union into openai/anthropic and fall back to the STORED backend +// for anything else. A submitted `gemini` was therefore validated as an OpenAI +// model and 400'd, while writing the identical pair straight into config.json +// worked. The executor was never the problem; only the write gate was. +const antigravityOAuth: OcxProviderConfig = { + adapter: "google", + baseUrl: "https://cloudcode-pa.googleapis.com", + authMode: "oauth", +}; + +function armAntigravity(): void { + accountSets["google-antigravity"] = { + accounts: [{ id: "acct-antigravity" }], + activeAccountId: "acct-antigravity", + }; + const set = accountSets["google-antigravity"] as unknown as { + accounts: Array<{ id: string; credential?: { projectId: string } }>; + }; + set.accounts[0].credential = { projectId: "proj-1" }; +} + +describe("submitted backend drives the pair check (#2457)", () => { + test("PUT switches openai/luna to a gemini pair the picker offers", async () => { + armAntigravity(); + managementRows = [{ provider: "google-antigravity", id: "gemini-3.7-flash", disabled: false }]; + const cfg = config({ + providers: { openai: forward, claude: anthropicOAuth, "google-antigravity": antigravityOAuth }, + webSearchSidecar: { backend: "openai", model: "gpt-5.6-luna" }, + }); + const response = await sidecarSettings(cfg, { + method: "PUT", + body: { webSearch: { backend: "gemini", model: "gemini-3.7-flash" } }, + }); + expect(response.status).toBe(200); + expect(cfg.webSearchSidecar).toMatchObject({ backend: "gemini", model: "gemini-3.7-flash" }); + }); + + test("PUT accepts a gemini pair when no sidecar is configured at all", async () => { + armAntigravity(); + managementRows = [{ provider: "google-antigravity", id: "gemini-3.7-flash", disabled: false }]; + const cfg = config({ + providers: { openai: forward, claude: anthropicOAuth, "google-antigravity": antigravityOAuth }, + }); + const response = await sidecarSettings(cfg, { + method: "PUT", + body: { webSearch: { backend: "gemini", model: "gemini-3.7-flash" } }, + }); + expect(response.status).toBe(200); + expect(cfg.webSearchSidecar).toMatchObject({ backend: "gemini", model: "gemini-3.7-flash" }); + }); + + test("PUT still rejects a gemini model submitted without its backend", async () => { + armAntigravity(); + managementRows = [{ provider: "google-antigravity", id: "gemini-3.7-flash", disabled: false }]; + const cfg = config({ + providers: { openai: forward, claude: anthropicOAuth, "google-antigravity": antigravityOAuth }, + webSearchSidecar: { backend: "openai", model: "gpt-5.6-luna" }, + }); + const response = await sidecarSettings(cfg, { + method: "PUT", + body: { webSearch: { model: "gemini-3.7-flash" } }, + }); + expect(response.status).toBe(400); + expect(cfg.webSearchSidecar).toEqual({ backend: "openai", model: "gpt-5.6-luna" }); + }); + + test("PUT still rejects a real mismatch inside the widened union", async () => { + armAntigravity(); + managementRows = [{ provider: "google-antigravity", id: "gemini-3.7-flash", disabled: false }]; + const cfg = config({ + providers: { openai: forward, claude: anthropicOAuth, "google-antigravity": antigravityOAuth }, + }); + const response = await sidecarSettings(cfg, { + method: "PUT", + body: { webSearch: { backend: "gemini", model: "gpt-5.6-luna" } }, + }); + expect(response.status).toBe(400); + expect(cfg.webSearchSidecar).toBeUndefined(); + }); +}); + describe("xSearch config round-trip (review High)", () => { test("PUT validates doc limits (400) and persists+echoes a valid block; GET carries it; null clears", async () => { usableCodexAccounts.add(MAIN_CODEX_ACCOUNT_ID); @@ -285,3 +367,43 @@ describe("xSearch config round-trip (review High)", () => { expect(cfg.webSearchSidecar).toEqual({ backend: "openai" }); }); }); + +// The Claude override is the second writer named in web-search-sidecar-options.ts: +// "a gate on one route and a stale copy on the other is the same as no gate at +// all." It carried the same collapsed ternary, so it needs the same proof (#2457). +describe("claude-code webSearchSidecar override honors the submitted backend (#2457)", () => { + async function claudeCode(cfg: OcxConfig, body: unknown): Promise { + const url = new URL("http://localhost/api/claude-code"); + const request = new Request(url, { + method: "PUT", + headers: { "content-type": "application/json" }, + body: JSON.stringify(body), + }); + const response = await handleManagementAPI(request, url, cfg); + if (!response) throw new Error("claude-code route did not handle the request"); + return response; + } + + function geminiConfig(): OcxConfig { + armAntigravity(); + managementRows = [{ provider: "google-antigravity", id: "gemini-3.7-flash", disabled: false }]; + return config({ + providers: { openai: forward, claude: anthropicOAuth, "google-antigravity": antigravityOAuth }, + claudeCode: { webSearchSidecar: { backend: "openai", model: "gpt-5.6-luna" } }, + }); + } + + test("PUT persists a gemini override over a stored openai pair", async () => { + const cfg = geminiConfig(); + const response = await claudeCode(cfg, { webSearchSidecar: { backend: "gemini", model: "gemini-3.7-flash" } }); + expect(response.status).toBe(200); + expect(cfg.claudeCode?.webSearchSidecar).toMatchObject({ backend: "gemini", model: "gemini-3.7-flash" }); + }); + + test("PUT still rejects a mismatched override pair", async () => { + const cfg = geminiConfig(); + const response = await claudeCode(cfg, { webSearchSidecar: { backend: "gemini", model: "gpt-5.6-luna" } }); + expect(response.status).toBe(400); + expect(cfg.claudeCode?.webSearchSidecar).toEqual({ backend: "openai", model: "gpt-5.6-luna" }); + }); +}); From b06cb1b4dc2b7b9bfff3b2b2ce063d49d64a06fc Mon Sep 17 00:00:00 2001 From: JUN Date: Tue, 25 Aug 2026 12:17:15 +0900 Subject: [PATCH 02/29] fix(cli): report Codex routing in ocx status and name an unused proxy (#2518) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * devlog: operator visibility train roadmap unit (260825) Docs-only roadmap for three operator-visibility defects that share one shape: OpenCodex computes the truth and does not report it. - 010 (#2457): the sidecar pair check collapses a five-member union into a two-arm ternary, so a submitted gemini backend is validated against the stored openai backend and the dashboard save 400s. - 020 (#2411): collectStatus already computes routingKind and ships it in status --json, but the human renderer never prints it, so a healthy proxy reads green while nothing routes through it. - 030 (#2412): a version-manager overwrite leaves a message-less ineligible verdict, and the CLI warns only when a message exists. 002 records the plan audit, including two amendments: WP4 must plumb messages to every reachable silent ineligible return rather than only the one the reporter hit, and WP3 must pin custom-local/unknown as intentionally silent. * fix(cli): report Codex routing in ocx status and name an unused proxy ocx status greened on process liveness alone, so a proxy answering /healthz read healthy even when ~/.codex/config.toml had no openai_base_url and every routed request went to OpenAI instead. The next line made it worse: native routing is genuinely rebootSafe, so restart safety printed 'no opencodex restart dependency' — a second all-clear for a broken setup. ocx doctor already knew, and printed routing=native, but status is the command people run first. status now prints the routing token doctor prints, and warns when a live proxy is paired with native routing. The proxy line keeps its check. That claim is narrow and true: the listener is up, which the reporter proved by curling it successfully. Turning it yellow would make the one honest signal lie to compensate for a missing one, and it would collide with the not-running branch whose remedy ('ocx start') is wrong here — the proxy does not need restarting, Codex needs re-pointing. Only native warns. custom-local and unknown are equally unused, but startupHealthSummary already renders both as AT RISK with a remedy, and custom-remote is a deliberate choice; warning on all four would train operators to ignore the line that matters. All five routing kinds are pinned in the test. The routing detail is extracted into formatStartupRoutingDetail so status and doctor cannot drift. Two callers computing the same string separately is exactly how #2457 happened. No routing decision, classifier, or restart-safety verdict changes. Closes #2411 --- src/cli/doctor.ts | 4 ++-- src/cli/index.ts | 11 +++++++++-- src/cli/status.ts | 23 ++++++++++++++++++++++ src/codex/autostart-health.ts | 16 +++++++++++++++ tests/autostart-health.test.ts | 36 +++++++++++++++++++++++++++++++++- tests/cli-help.test.ts | 5 +++++ 6 files changed, 90 insertions(+), 5 deletions(-) diff --git a/src/cli/doctor.ts b/src/cli/doctor.ts index 72a9e6202f9..aaadb343559 100644 --- a/src/cli/doctor.ts +++ b/src/cli/doctor.ts @@ -42,7 +42,7 @@ import { resolveEffectiveUserIdentity, } from "../codex/user-identity"; import { collectProjectCodexConfigWarnings, formatProjectCodexConfigWarningsForDoctor } from "../codex/project-config-warnings"; -import { collectStartupHealth, startupHealthSummary } from "../codex/autostart-health"; +import { collectStartupHealth, formatStartupRoutingDetail, startupHealthSummary } from "../codex/autostart-health"; import { displayCodexRuntimePath, loadLastEffortClamp, @@ -983,7 +983,7 @@ export async function runDoctor(args: string[] = []): Promise { const startup = collectStartupHealth(doctorConfig); console.log("\nCodex restart safety"); console.log(` ${startup.rebootSafe ? "ok " : "!! "} ${startupHealthSummary(startup)}`); - console.log(` routing=${startup.routingKind}, service=${startup.serviceViable ? "viable" : startup.serviceInstalled ? "installed-but-unhealthy" : "absent"}, shim=${startup.shimHealthy ? "healthy" : startup.shimInstalled ? "stale" : "absent"}`); + console.log(` ${formatStartupRoutingDetail(startup)}`); console.log("\nCodex runtime selection"); { diff --git a/src/cli/index.ts b/src/cli/index.ts index 57d5c85c532..85d4b61a94e 100755 --- a/src/cli/index.ts +++ b/src/cli/index.ts @@ -25,7 +25,7 @@ import { writePid, writeRuntimePort, } from "../config/process-state"; -import { collectStatus } from "./status"; +import { collectStatus, unusedProxyWarningLines } from "./status"; import { discoverStableProxyForRestart, @@ -46,7 +46,7 @@ import { runCli } from "./root"; import { ProxyOwnershipRefusedError, stopProxy } from "../lib/process-control"; import { loadServiceTokenFromFile } from "../lib/service-secrets"; import { diagnoseService, isServiceOwnershipError, serviceCommand, serviceEnvironmentOwnedHere, serviceStartableFromTray, serviceStatusSummary, stopServiceIfInstalled, uninstallServiceIfInstalled } from "../service"; -import { startupHealthSummary } from "../codex/autostart-health"; +import { formatStartupRoutingDetail, startupHealthSummary } from "../codex/autostart-health"; import { drainAndShutdown, isRecyclingForExit, startServer } from "../server"; import { injectSystemEnv, reconcileShellHook, revertSystemEnv, uninstallShellHook } from "../server/system-env"; import { buildDesktop3pRegistry } from "../claude/desktop-3p"; @@ -848,6 +848,12 @@ async function handleStatus() { console.log(`❌ Proxy: ${status.proxyLabel}`); } console.log(` Health: ${status.healthLabel}`); + for (const line of unusedProxyWarningLines({ + proxyUp: Boolean(status.json.proxy.pid || status.json.proxy.health.ok), + routingKind: status.json.startup.routingKind, + })) { + console.log(` ${line}`); + } if (!(status.json.proxy.pid || status.json.proxy.health.ok)) { console.log(" ↳ Not running — Codex/Claude requests will fail with connection errors."); // The service summary a few lines below already tells a registered-but-not-serving @@ -867,6 +873,7 @@ async function handleStatus() { console.log(` Default provider: ${status.json.defaultProvider}`); console.log(` Codex autostart: ${status.json.codexAutostart ? "enabled" : "disabled"}`); console.log(` Restart safety: ${startupHealthSummary(status.json.startup)}`); + console.log(` ${formatStartupRoutingDetail(status.json.startup)}`); console.log(` Service: ${status.json.service.summary}`); console.log(` ${status.json.codexShim.summary}`); console.log(` Codex runtime: ${status.json.codexRuntime.path}`); diff --git a/src/cli/status.ts b/src/cli/status.ts index 9c9f577e0e9..1ba11f96a33 100644 --- a/src/cli/status.ts +++ b/src/cli/status.ts @@ -118,6 +118,29 @@ export function proxyHealthFailureReason(error: unknown, signal: AbortSignal): " : "unreachable"; } +/** + * `ocx status` greens on process liveness alone, so a proxy that answers + * /healthz reads healthy even when Codex is not pointed at it and every routed + * request goes to OpenAI instead (#2411). The proxy line is not wrong — the + * listener really is up — so it keeps its check, and this supplies the signal + * that was missing rather than corrupting the one that was already honest. + * + * Only `native` warns. `custom-local` and `unknown` are also "this proxy is + * unused", but startupHealthSummary already renders both as AT RISK with a + * remedy command, and `custom-remote` is a deliberate operator choice. Warning + * on all four would teach operators to skip the line that matters. + */ +export function unusedProxyWarningLines(input: { + proxyUp: boolean; + routingKind: StartupHealth["routingKind"]; +}): string[] { + if (!input.proxyUp || input.routingKind !== "native") return []; + return [ + "⚠️ Codex routing is native — the running proxy is unused.", + " Codex requests go to OpenAI, not this proxy. Re-point with: ocx start", + ]; +} + async function checkProxyHealth(target: ListenTarget): Promise { const url = target.healthUrl; const controller = new AbortController(); diff --git a/src/codex/autostart-health.ts b/src/codex/autostart-health.ts index b3496696e17..95df491ac3f 100644 --- a/src/codex/autostart-health.ts +++ b/src/codex/autostart-health.ts @@ -154,3 +154,19 @@ export function startupHealthSummary(health: StartupHealth): string { if (health.serviceInstalled && !health.serviceViable) return `AT RISK after restart (installed service is disabled, stopped, or unhealthy; run '${command}')`; return `AT RISK after restart (no viable background service; run '${command}')`; } + +/** + * The routing/service/shim token `ocx doctor` prints under restart safety. + * Extracted so `ocx status` can show the same string rather than growing a + * second copy that drifts (#2411). Two management routes computing the same + * thing separately is exactly how #2457 happened. + */ +export function formatStartupRoutingDetail(health: StartupHealth): string { + const service = health.serviceViable + ? "viable" + : health.serviceInstalled ? "installed-but-unhealthy" : "absent"; + const shim = health.shimHealthy + ? "healthy" + : health.shimInstalled ? "stale" : "absent"; + return `routing=${health.routingKind}, service=${service}, shim=${shim}`; +} diff --git a/tests/autostart-health.test.ts b/tests/autostart-health.test.ts index 157f1e95abb..c3fe6aa9911 100644 --- a/tests/autostart-health.test.ts +++ b/tests/autostart-health.test.ts @@ -1,5 +1,6 @@ import { describe, expect, test } from "bun:test"; -import { deriveStartupHealth, startupHealthSummary } from "../src/codex/autostart-health"; +import { deriveStartupHealth, formatStartupRoutingDetail, startupHealthSummary } from "../src/codex/autostart-health"; +import { unusedProxyWarningLines } from "../src/cli/status"; import { classifyCodexRouting, hasInjectedCodexRouting } from "../src/codex/inject"; import { handleManagementAPI } from "../src/server/management-api"; import { invalidateStartupHealthCache, markStartupHealthDiagnosticStale } from "../src/server/startup-health-cache"; @@ -231,3 +232,36 @@ describe("Codex startup health", () => { }, 40_000); }); import { ManagementRequest as Request } from "./helpers/management-auth"; + +describe("routing visibility (#2411)", () => { + test("formatStartupRoutingDetail renders the token doctor already prints", () => { + expect(formatStartupRoutingDetail(deriveStartupHealth({ ...base, routingKind: "native" }))) + .toBe("routing=native, service=absent, shim=absent"); + expect(formatStartupRoutingDetail(deriveStartupHealth({ + ...base, + serviceInstalled: true, + serviceViable: true, + shimInstalled: true, + shimHealthy: true, + }))).toBe("routing=opencodex-local, service=viable, shim=healthy"); + expect(formatStartupRoutingDetail(deriveStartupHealth({ ...base, serviceInstalled: true }))) + .toBe("routing=opencodex-local, service=installed-but-unhealthy, shim=absent"); + expect(formatStartupRoutingDetail(deriveStartupHealth({ ...base, shimInstalled: true }))) + .toBe("routing=opencodex-local, service=absent, shim=stale"); + }); + + // A healthy proxy paired with native routing is the state #2411 reports: the + // process answers /healthz truthfully while no Codex request reaches it. + // custom-local and unknown stay silent on purpose — startupHealthSummary + // already renders both as AT RISK with a remedy, so a second warning would + // train operators to ignore this one. + test("unusedProxyWarningLines fires only for a live proxy on native routing", () => { + expect(unusedProxyWarningLines({ proxyUp: true, routingKind: "native" }).length).toBeGreaterThan(0); + expect(unusedProxyWarningLines({ proxyUp: true, routingKind: "native" }).join(" ")).toContain("unused"); + expect(unusedProxyWarningLines({ proxyUp: false, routingKind: "native" })).toEqual([]); + expect(unusedProxyWarningLines({ proxyUp: true, routingKind: "opencodex-local" })).toEqual([]); + expect(unusedProxyWarningLines({ proxyUp: true, routingKind: "custom-remote" })).toEqual([]); + expect(unusedProxyWarningLines({ proxyUp: true, routingKind: "custom-local" })).toEqual([]); + expect(unusedProxyWarningLines({ proxyUp: true, routingKind: "unknown" })).toEqual([]); + }); +}); diff --git a/tests/cli-help.test.ts b/tests/cli-help.test.ts index 4fa55e2da1c..1edfb67a992 100644 --- a/tests/cli-help.test.ts +++ b/tests/cli-help.test.ts @@ -174,6 +174,11 @@ describe("CLI subcommand help", () => { expect(result.stdout).toContain("Service:"); expect(result.stdout).toContain(join(opencodexHome, "service.log")); expect(result.stdout).toContain("Codex autostart shim"); + // #2411: status must name the routing kind it already computes. The + // proxy is down in this fixture, so the unused-proxy warning must stay + // quiet — that warning is for a LIVE proxy nothing routes through. + expect(result.stdout).toContain("routing="); + expect(result.stdout).not.toContain("the running proxy is unused"); } finally { rmSync(opencodexHome, { recursive: true, force: true }); } From 47a31d76ed2f96496650cb41a85a5377fe1fb94f Mon Sep 17 00:00:00 2001 From: JUN Date: Tue, 25 Aug 2026 12:17:23 +0900 Subject: [PATCH 03/29] fix(codex): report a destroyed shim instead of bailing silently (#2519) * devlog: operator visibility train roadmap unit (260825) Docs-only roadmap for three operator-visibility defects that share one shape: OpenCodex computes the truth and does not report it. - 010 (#2457): the sidecar pair check collapses a five-member union into a two-arm ternary, so a submitted gemini backend is validated against the stored openai backend and the dashboard save 400s. - 020 (#2411): collectStatus already computes routingKind and ships it in status --json, but the human renderer never prints it, so a healthy proxy reads green while nothing routes through it. - 030 (#2412): a version-manager overwrite leaves a message-less ineligible verdict, and the CLI warns only when a message exists. 002 records the plan audit, including two amendments: WP4 must plumb messages to every reachable silent ineligible return rather than only the one the reporter hit, and WP3 must pin custom-local/unknown as intentionally silent. * fix(codex): report a destroyed shim instead of bailing silently A version manager (mise/asdf/volta) rewrites its install tree in place on upgrade, destroying both the opencodex shim and the sibling .opencodex-real backup that auto-restore needs. Auto-restore returned a bare ineligible with no message, and the CLI warns only when a message exists, so ocx start, ocx ensure, and ocx service repair all reported success while routing quietly stayed native. That is the cause behind the misleading green status in #2411. Auto-restore now explains itself. The message names the wrapper state, the missing backup, and both paths, and for a version-manager tree it says why no repair is coming and what to do instead. Restoring by adopting the newly installed binary as a replacement original would be wrong twice: it records a provenance that never happened, and the next upgrade overwrites it again, so the repair would silently un-repair on the version manager's schedule. allowFreshInstall stays false. The replacement path is also refused on a version-manager tree even when a stale backup survives, since that is the same adoption arriving through the back door. Per the plan audit, every reachable silent bail now carries a message, not only the one in the report: the preserveOnly sibling case and the unhealthy-probe case were equally invisible. Healthy shims are untouched, including version-manager-owned ones that are currently intact. Detection gates repair, not operation. Docs record that a version-manager install tree is not a supported shim target and point those users at openai_base_url routing plus the service. Closes #2412 --- .../content/docs/reference/cli/lifecycle.md | 13 ++++ src/codex/shim.ts | 59 ++++++++++++++++++- tests/codex-shim.test.ts | 34 ++++++++++- 3 files changed, 102 insertions(+), 4 deletions(-) diff --git a/docs-site/src/content/docs/reference/cli/lifecycle.md b/docs-site/src/content/docs/reference/cli/lifecycle.md index f406c3cea4f..930368c91c4 100644 --- a/docs-site/src/content/docs/reference/cli/lifecycle.md +++ b/docs-site/src/content/docs/reference/cli/lifecycle.md @@ -360,6 +360,19 @@ changing is left untouched and retried later. Repair failures warn without faili command; manual fallback: `ocx codex-shim install`. Set `codexShimAutoRestore` to `false`, or set `OPENCODEX_CODEX_SHIM_AUTO_RESTORE=0` for a process-level opt-out. +That restore needs the original launcher OpenCodex saved next to the shim. A version manager — +mise, asdf, volta — rewrites its whole install tree on upgrade, which destroys the shim *and* that +backup, so there is nothing left to restore from. **A version-manager install tree is not a +supported shim target.** OpenCodex reports the condition and stops rather than wrapping the newly +installed binary as a replacement original: doing so would record a history that never happened, and +the next upgrade would overwrite it again, so the repair would silently undo itself on the version +manager's schedule. + +If your `codex` is owned by a version manager, route through Codex configuration instead of the +launcher: `ocx start` writes `openai_base_url`, and `ocx service install` provides autostart. Run +`ocx status` to confirm — it reports the active routing, and warns when a running proxy is not the +one Codex is pointed at. + | Subcommand | Action | | --- | --- | | `install` | Install the shim (or repair if stale). | diff --git a/src/codex/shim.ts b/src/codex/shim.ts index c4c1b834000..4f2ba3a6dbd 100644 --- a/src/codex/shim.ts +++ b/src/codex/shim.ts @@ -603,6 +603,47 @@ function backupPathFor(path: string): string { return ext ? `${path.slice(0, -ext.length)}.opencodex-real${ext}` : `${path}.opencodex-real`; } +/** + * True when a Codex binary lives inside a version manager's install tree. + * + * These trees are rewritten in place on upgrade, which destroys both the shim + * and the sibling .opencodex-real backup it restores from (#2412). The tempting + * repair — adopt the newly installed binary as a fresh original — is wrong + * twice: it records a provenance that never happened, and the next upgrade wipes + * it again, so the repair silently un-repairs on the version manager's schedule. + * + * Scope is the three managers named in the report. nvm/fnm/npm-prefix are + * deliberately excluded: a false positive here refuses a restore that would + * otherwise be correct. + */ +export function isVersionManagerOwnedCodexPath(path: string): boolean { + const normalized = path.replace(/\\/g, "/").toLowerCase(); + return normalized.includes("/mise/installs/") + || normalized.includes("/mise/shims/") + || normalized.includes("/.asdf/installs/") + || normalized.includes("/.asdf/shims/") + || normalized.includes("/.volta/"); +} + +/** + * Why auto-restore refused, in the operator's own terms. Auto-restore used to + * return a bare `{ status: "ineligible" }`, and the CLI warns only when a + * message is present, so `ocx start`, `ocx ensure`, and `ocx service repair` all + * reported success while routing quietly stayed native (#2412, the cause behind + * the misleading green status in #2411). + */ +function destroyedShimMessage(file: ShimFileState): string { + const wrapper = existsSync(file.wrapperPath) + ? isShim(file.wrapperPath) ? "present but unusable" : "present but not an opencodex shim" + : "missing"; + const backup = existsSync(file.backupPath) ? "present" : "missing"; + const base = `Codex autostart shim not restored: wrapper ${wrapper} at ${file.wrapperPath}; original backup ${backup} at ${file.backupPath}.`; + if (!isVersionManagerOwnedCodexPath(file.wrapperPath)) { + return `${base} Re-run 'ocx codex-shim install' once the Codex binary is stable.`; + } + return `${base} This Codex binary is owned by a version manager (mise/asdf/volta), so opencodex will not wrap it as a new original — the next upgrade would overwrite the shim and its backup again. Route through Codex instead with 'ocx start', and use 'ocx service install' for autostart.`; +} + function shQuote(value: string): string { return `'${value.replace(/'/g, "'\\''")}'`; } @@ -2039,14 +2080,20 @@ export function autoRestoreCodexShim(options: { if (seen.has(file.wrapperPath)) return { status: "ineligible" }; seen.add(file.wrapperPath); if (file.preserveOnly) { - if (!existsSync(file.backupPath) || existsSync(file.originalPath)) return { status: "ineligible" }; + if (!existsSync(file.backupPath) || existsSync(file.originalPath)) { + return { status: "ineligible", message: destroyedShimMessage(file) }; + } continue; } - if (!existsSync(file.wrapperPath) || !hasUsableBackingPath(file)) return { status: "ineligible" }; + if (!existsSync(file.wrapperPath) || !hasUsableBackingPath(file)) { + return { status: "ineligible", message: destroyedShimMessage(file) }; + } const probe = stableShimPathProbe(file.wrapperPath); if (!probe) return { status: "deferred" }; if (probe.prefix.includes(SHIM_MARKER)) { - if (!isHealthyShimProbe(probe, state.platform)) return { status: "ineligible" }; + if (!isHealthyShimProbe(probe, state.platform)) { + return { status: "ineligible", message: destroyedShimMessage(file) }; + } if (state.platform !== "win32" && !isCurrentUnixShimProbe(probe)) { obsoleteShimProbes.set(file.wrapperPath, probe); continue; @@ -2054,6 +2101,12 @@ export function autoRestoreCodexShim(options: { healthyCount += 1; continue; } + // A surviving backup would otherwise let the replacement path below wrap the + // version manager's NEW binary as a fresh original — the same adoption the + // missing-backup case refuses, arriving through the back door. + if (isVersionManagerOwnedCodexPath(file.wrapperPath)) { + return { status: "ineligible", message: destroyedShimMessage(file) }; + } replacementProbes.set(file.wrapperPath, probe); } diff --git a/tests/codex-shim.test.ts b/tests/codex-shim.test.ts index b9ed2a5663e..7ab67f595dc 100644 --- a/tests/codex-shim.test.ts +++ b/tests/codex-shim.test.ts @@ -3,7 +3,7 @@ import { spawnSync } from "node:child_process"; import { chmodSync, copyFileSync, existsSync, lstatSync, mkdirSync, mkdtempSync, readFileSync, readdirSync, renameSync, rmSync, statSync, symlinkSync, utimesSync, writeFileSync } from "node:fs"; import { delimiter, dirname, join } from "node:path"; import { tmpdir } from "node:os"; -import { autoRestoreCodexShim, buildUnixCodexShim, buildWindowsCodexShim, buildWindowsPowerShellCodexShim, diagnoseCodexShim, findCodexOnPath, installCodexShim, isWindowsInteropDir, lastCodexDiscoveryError, setCodexShimFreshWriteHookForTests, setCodexShimGuardedWriteHookForTests, setCodexShimProbeHookForTests, setCodexShimProbeObservationMsForTests, setCodexShimProbeShellForTests, setCodexShimRollbackRestoreHookForTests, uninstallCodexShim } from "../src/codex/shim"; +import { autoRestoreCodexShim, buildUnixCodexShim, buildWindowsCodexShim, buildWindowsPowerShellCodexShim, diagnoseCodexShim, findCodexOnPath, installCodexShim, isVersionManagerOwnedCodexPath, isWindowsInteropDir, lastCodexDiscoveryError, setCodexShimFreshWriteHookForTests, setCodexShimGuardedWriteHookForTests, setCodexShimProbeHookForTests, setCodexShimProbeObservationMsForTests, setCodexShimProbeShellForTests, setCodexShimRollbackRestoreHookForTests, uninstallCodexShim } from "../src/codex/shim"; const SHIM_MARKER = "opencodex codex autostart shim"; const UNIX_SHIM_REVISION_MARKER = "opencodex unix codex shim revision 2"; @@ -1999,3 +1999,35 @@ describe("WSL PATH interop guard", () => { expect(found).toBe(`${dir}/codex`); }); }); + +// #2412: a version-manager upgrade (mise/asdf/volta) rewrites its install tree +// in place, destroying both the shim and its sibling .opencodex-real backup. +// The bail was silent, and the CLI warns only when a message exists, so start / +// ensure / service repair all reported success while routing stayed native. +describe("version-manager shim destruction (#2412)", () => { + test("classifies version-manager install trees without catching ordinary paths", () => { + expect(isVersionManagerOwnedCodexPath("/home/u/.local/share/mise/installs/codex/latest/bin/codex")).toBe(true); + expect(isVersionManagerOwnedCodexPath("/home/u/.local/share/mise/shims/codex")).toBe(true); + expect(isVersionManagerOwnedCodexPath("/home/u/.asdf/installs/codex/1.0/bin/codex")).toBe(true); + expect(isVersionManagerOwnedCodexPath("/home/u/.asdf/shims/codex")).toBe(true); + expect(isVersionManagerOwnedCodexPath("/home/u/.volta/bin/codex")).toBe(true); + expect(isVersionManagerOwnedCodexPath("C:\\Users\\u\\.volta\\bin\\codex.cmd")).toBe(true); + expect(isVersionManagerOwnedCodexPath("/usr/local/bin/codex")).toBe(false); + expect(isVersionManagerOwnedCodexPath("/home/u/.npm-global/bin/codex")).toBe(false); + expect(isVersionManagerOwnedCodexPath("/opt/homebrew/bin/codex")).toBe(false); + }); + + test("a destroyed shim reports the paths instead of bailing silently", () => { + withInstalledShim(({ wrappers, backups }) => { + writeFileSync(wrappers[0], "#!/bin/sh\necho version-manager codex\n", "utf8"); + if (process.platform !== "win32") chmodSync(wrappers[0], 0o755); + rmSync(backups[0]); + const result = autoRestoreCodexShim({ enabled: () => true, stabilitySleep: skipStabilityWait }); + expect(result.status).toBe("ineligible"); + // The silent bail is the whole defect: cli/codex-shim-autorestore.ts warns + // only when a message exists. + expect(result.message).toBeTruthy(); + expect(result.message).toContain("backup"); + }); + }); +}); From fff86110a790b4658d785958b7130f901baa2160 Mon Sep 17 00:00:00 2001 From: Olddonkey Date: Mon, 24 Aug 2026 20:21:23 -0700 Subject: [PATCH 04/29] fix(xai): stop the undeclared-tool guard from killing hosted x_search turns (#2425) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * fix(xai): stop the undeclared-tool guard from killing hosted x_search turns xAI executes hosted `x_search` itself and reports the activity as a `custom_tool_call` whose name is deliberately absent from the request catalog. The guard treats every `custom_tool_call` as client-executed, so once a request also declares any NAMED client tool the guard trips on xAI's own hosted call and fails the whole turn. Reproduced 2026-08-22 with the identical request ({type:"function",name:"shell"} + {type:"x_search"}): direct to xAI 200, custom_tool_call + message, 10 annotations through opencodex response.failed, no response.completed, reasoning item only, 0 output chars A request declaring ONLY x_search passes, because the guard activates only once a named client tool exists — which is why this is easy to miss with a minimal repro and why every realistic Codex request would hit it. The fix mirrors the existing NAMELESS_CLIENT_DECLARATION_CALL_TYPES in the other direction: PROVIDER_EXECUTED_DECLARATION_CALL_TYPES maps a hosted declaration to the item type the provider emits for it, and those items need no client name to be authorized. Two gates, both required, so this cannot widen into a blanket exemption: destination core.ts passes an empty set unless the route actually terminates at xAI (isXaiResponsesDestination: exact host, https, standard port — lookalikes and odd ports excluded) declaration the turn must actually declare x_search Names are never matched. One turn emitted `x_keyword_search` and `x_semantic_search`, and the other xAI host emits `x_user_search` — three literals for one tool, so the name channel carries no signal. RESIDUAL RISK, accepted deliberately and documented at the branch: inside a turn that declared x_search on xAI, a hallucinated client custom tool is exempted too, precisely because names cannot be trusted. The alternative is failing every hosted-search turn. #1700's protection is untouched for every other turn, provider and item type. Tests pin the two gates in the negative direction as well as the positive: no declaration still refuses, empty authorization still refuses, and apply_patch is still refused INSIDE an authorized turn. Gate: 14438 pass / 1 fail; that failure also fails on untouched upstream/dev at the same commit (baseline: 4 fail, a superset). Zero regressions. * fix(xai): require the hosted call-id prefix, not just the item type Review caught that the previous shape exempted EVERY custom_tool_call inside an authorized turn — and apply_patch, the tool #1700 exists to protect, arrives as a custom_tool_call (see the repo's own 'never blocks apply_patch' test). The regression test written to prove #1700 survived used a function_call, which the exemption never touched, so it asserted something true but irrelevant. Measured 2026-08-23 against cli-chat-proxy.grok.com, a hosted x_search item is {type:custom_tool_call, name:x_keyword_search, call_id:xs_call-428a4403-...}. Authorization now requires the item type AND that call-id prefix, on top of the existing destination and declaration gates. Names are still never matched — three literals have been observed for this one tool. The #1700 test now uses the real shape: an undeclared custom_tool_call named apply_patch with an ordinary call_id, inside an authorized turn, still refused. core.ts passes a correctly-typed empty set on the non-xAI branch, so a Set is now a compile error rather than a silent no-op. * test(xai): pin hosted-call destination boundary --- src/providers/xai-transport.ts | 21 +++ src/server/responses-undeclared-tool-guard.ts | 86 +++++++++++- src/server/responses/core.ts | 13 +- tests/responses-undeclared-tool-guard.test.ts | 128 ++++++++++++++++++ tests/xai-transport.test.ts | 22 +++ 5 files changed, 265 insertions(+), 5 deletions(-) diff --git a/src/providers/xai-transport.ts b/src/providers/xai-transport.ts index 1e56506ecca..e005abdf4e6 100644 --- a/src/providers/xai-transport.ts +++ b/src/providers/xai-transport.ts @@ -4,6 +4,27 @@ import { resolveGithubCopilotTransport } from "./github-copilot-transport"; export const XAI_GROK_CLI_BASE_URL = "https://cli-chat-proxy.grok.com/v1"; +/** The two hosts that serve xAI's Responses API: the public API and the Grok CLI proxy. */ +const XAI_RESPONSES_HOSTS = new Set(["api.x.ai", "cli-chat-proxy.grok.com"]); + +/** + * True when this provider's Responses traffic terminates at xAI itself. + * + * Probed 2026-08-22 one field per request: the two hosts accept and refuse exactly the same + * web_search fields, so they are one dialect rather than two. Matching is exact-host over + * https, which keeps lookalikes (`api.x.ai.evil.test`) and nonstandard ports out. + */ +export function isXaiResponsesDestination(provider: Pick): boolean { + try { + const url = new URL(provider.baseUrl); + return url.protocol === "https:" + && XAI_RESPONSES_HOSTS.has(url.hostname.toLowerCase()) + && (url.port === "" || url.port === "443"); + } catch { + return false; + } +} + export const XAI_GROK_COMPATIBILITY = { version: "0.2.93", userAgent: "opencodex-grok/0.2.93", diff --git a/src/server/responses-undeclared-tool-guard.ts b/src/server/responses-undeclared-tool-guard.ts index 658a1c6bab5..158e6585b89 100644 --- a/src/server/responses-undeclared-tool-guard.ts +++ b/src/server/responses-undeclared-tool-guard.ts @@ -4,6 +4,28 @@ import { sseDataPayload, type SseBlockRewrite } from "./sse-payload-rewrite"; /** Item types the client executes through a request-declared wire name. */ const CLIENT_EXECUTED_CALL_TYPES = new Set(["function_call", "custom_tool_call"]); +/** + * Hosted declarations whose response items the PROVIDER executes, keyed by the request + * declaration type. These need no client answer, so their names are deliberately absent from + * the request catalog and must not be read as an undeclared client tool. + * + * xAI surfaces hosted `x_search` as `custom_tool_call`. Probed 2026-08-23 against the OAuth CLI + * destination: its hosted calls use an `xs_call-` call-id prefix. Observed call names were + * `x_keyword_search`, `x_semantic_search`, and `x_user_search` — three literals for one tool, + * which is why authorization keys on the declaration, item type, and call-id prefix, never on + * the name. + */ +export type ProviderExecutedCallType = Readonly<{ + itemType: string; + callIdPrefix: string; +}>; + +type ProviderExecutedCallTypes = ReadonlySet; + +export const PROVIDER_EXECUTED_DECLARATION_CALL_TYPES = new Map([ + ["x_search", { itemType: "custom_tool_call", callIdPrefix: "xs_call-" }], +]); + /** Nameless declaration kinds whose response items still require client execution. */ const NAMELESS_CLIENT_DECLARATION_CALL_TYPES = new Map([ ["local_shell", "local_shell_call"], @@ -19,6 +41,7 @@ const NAMELESS_CLIENT_CALL_DISPLAY_NAMES = new Map([ ]); const EMPTY_DECLARED_NAMELESS_CLIENT_CALL_TYPES: ReadonlySet = new Set(); +const EMPTY_PROVIDER_EXECUTED_CALL_TYPES: ReadonlySet = new Set(); /** Supported hosted/private declarations that carry no client-executable wire name. */ const NAMELESS_TOOL_SPEC_TYPES = new Set([ @@ -127,6 +150,53 @@ function addNamelessClientCallTypes(callTypes: Set, specs: unknown): voi } } +function addProviderExecutedCallTypes( + callTypes: Set, + specs: unknown, +): void { + if (!Array.isArray(specs)) return; + for (const spec of specs) { + if (!isPlainObject(spec) || typeof spec.type !== "string") continue; + const callType = PROVIDER_EXECUTED_DECLARATION_CALL_TYPES.get(spec.type); + if (callType) callTypes.add(callType); + } +} + +/** + * Item types this turn's hosted declarations authorize the PROVIDER to emit unnamed. + * + * Caller must gate this on the destination actually being that provider; a declaration alone + * is not authority, or any upstream could claim a hosted shape it never serves. + */ +export function collectProviderExecutedCallTypes(body: unknown): Set { + const callTypes = new Set(); + if (!isPlainObject(body)) return callTypes; + addProviderExecutedCallTypes(callTypes, body.tools); + if (Array.isArray(body.input)) { + for (const item of body.input) { + if ( + isPlainObject(item) + && (item.type === "additional_tools" || item.type === "tool_search_output") + ) addProviderExecutedCallTypes(callTypes, item.tools); + } + } + return callTypes; +} + +function isAuthorizedProviderExecutedCall( + item: Record, + callTypes: ProviderExecutedCallTypes, +): boolean { + if (typeof item.call_id !== "string") return false; + for (const callType of callTypes) { + if ( + item.type === callType.itemType + && item.call_id.startsWith(callType.callIdPrefix) + ) return true; + } + return false; +} + /** Nameless client-call item types authorized by supported request tool declarations. */ export function collectDeclaredNamelessClientCallTypes(body: unknown): Set { const callTypes = new Set(); @@ -190,9 +260,14 @@ function undeclaredNameInItem( item: unknown, declared: ReadonlySet, declaredNamelessClientCallTypes: ReadonlySet, + providerExecutedCallTypes: ProviderExecutedCallTypes = EMPTY_PROVIDER_EXECUTED_CALL_TYPES, ): string | undefined { if (!isPlainObject(item)) return undefined; if (typeof item.type !== "string") return undefined; + // The provider executes this exact measured shape itself, so there is no client name to + // authorize. The caller supplies these signatures only for the matching destination and + // declarations; the item must additionally carry the hosted call-id prefix. + if (isAuthorizedProviderExecutedCall(item, providerExecutedCallTypes)) return undefined; const namelessDisplayName = NAMELESS_CLIENT_CALL_DISPLAY_NAMES.get(item.type); if (namelessDisplayName !== undefined) { // Only Codex's explicit `execution: "client"` form delegates tool search to the client. @@ -214,14 +289,15 @@ export function undeclaredToolCallName( payload: unknown, declared: ReadonlySet, declaredNamelessClientCallTypes: ReadonlySet = EMPTY_DECLARED_NAMELESS_CLIENT_CALL_TYPES, + providerExecutedCallTypes: ProviderExecutedCallTypes = EMPTY_PROVIDER_EXECUTED_CALL_TYPES, ): string | undefined { if (!isPlainObject(payload)) return undefined; if (payload.type === "response.output_item.added" || payload.type === "response.output_item.done") { - return undeclaredNameInItem(payload.item, declared, declaredNamelessClientCallTypes); + return undeclaredNameInItem(payload.item, declared, declaredNamelessClientCallTypes, providerExecutedCallTypes); } // Sparse gateways skip incremental items and only ever ship the terminal snapshot. if (payload.type === "response.completed" || payload.type === "response.incomplete") { - return undeclaredToolCallNameInResponse(payload.response, declared, declaredNamelessClientCallTypes); + return undeclaredToolCallNameInResponse(payload.response, declared, declaredNamelessClientCallTypes, providerExecutedCallTypes); } return undefined; } @@ -231,10 +307,11 @@ export function undeclaredToolCallNameInResponse( response: unknown, declared: ReadonlySet, declaredNamelessClientCallTypes: ReadonlySet = EMPTY_DECLARED_NAMELESS_CLIENT_CALL_TYPES, + providerExecutedCallTypes: ProviderExecutedCallTypes = EMPTY_PROVIDER_EXECUTED_CALL_TYPES, ): string | undefined { if (!isPlainObject(response) || !Array.isArray(response.output)) return undefined; for (const item of response.output) { - const name = undeclaredNameInItem(item, declared, declaredNamelessClientCallTypes); + const name = undeclaredNameInItem(item, declared, declaredNamelessClientCallTypes, providerExecutedCallTypes); if (name !== undefined) return name; } return undefined; @@ -274,6 +351,7 @@ function failedBlocks(name: string, newline: string): readonly string[] { export function createUndeclaredToolCallGuardBlockRewrite( declared: ReadonlySet, declaredNamelessClientCallTypes: ReadonlySet = EMPTY_DECLARED_NAMELESS_CLIENT_CALL_TYPES, + providerExecutedCallTypes: ProviderExecutedCallTypes = EMPTY_PROVIDER_EXECUTED_CALL_TYPES, ): SseBlockRewrite { let tripped = false; return (block: string) => { @@ -286,7 +364,7 @@ export function createUndeclaredToolCallGuardBlockRewrite( } catch { return [block]; } - const name = undeclaredToolCallName(parsed, declared, declaredNamelessClientCallTypes); + const name = undeclaredToolCallName(parsed, declared, declaredNamelessClientCallTypes, providerExecutedCallTypes); if (name === undefined) return [block]; tripped = true; return failedBlocks(name, block.includes("\r\n") ? "\r\n" : "\n"); diff --git a/src/server/responses/core.ts b/src/server/responses/core.ts index 67005be579d..3aed07ef0f1 100644 --- a/src/server/responses/core.ts +++ b/src/server/responses/core.ts @@ -191,7 +191,7 @@ import { rotateProviderTransportOn429, } from "../../providers/key-failover"; import { shouldAttemptImageTierRetry } from "../image-retry"; -import { resolveProviderTransport } from "../../providers/xai-transport"; +import { isXaiResponsesDestination, resolveProviderTransport } from "../../providers/xai-transport"; import type { WsData } from "../ws-bridge"; import { codexAccountSelectionForTurn, registerTurn, trackStreamLifetime, unregisterTurn } from "../lifecycle"; import { redactSecretString, sanitizeLogMetadataString } from "../../lib/redact"; @@ -308,12 +308,14 @@ import { import { collectDeclaredNamelessClientCallTypes, collectDeclaredWireToolNames, + collectProviderExecutedCallTypes, createUndeclaredToolCallGuardBlockRewrite, currentTurnWireToolCatalogBody, hasExplicitWireToolCatalog, undeclaredToolCallMessage, undeclaredToolCallName, undeclaredToolCallNameInResponse, + type ProviderExecutedCallType, } from "../responses-undeclared-tool-guard"; import { createGithubCopilotResponsesBlockRewrite } from "../github-copilot-responses-repair"; import { responsesJsonToSseStream } from "../responses-json-events"; @@ -2942,6 +2944,11 @@ async function handleResponsesInner( const clientDeclaredNamelessCallTypes = collectDeclaredNamelessClientCallTypes( clientToolAuthorizationBody, ); + // Hosted calls the PROVIDER runs itself. Gated on the destination actually being xAI, so a + // declaration alone cannot buy the exemption on some other upstream that never serves it. + const providerExecutedCallTypes = isXaiResponsesDestination(route.provider) + ? collectProviderExecutedCallTypes(clientToolAuthorizationBody) + : new Set(); let request: Awaited>; try { request = await adapter.buildRequest(parsed, { headers: selectedForwardHeaders, translatorBudget }); @@ -3061,6 +3068,7 @@ async function handleResponsesInner( payload, declaredWireToolNames, declaredNamelessClientCallTypes, + providerExecutedCallTypes, ) !== undefined) { inspectionSawUndeclaredTool = true; } @@ -3074,6 +3082,7 @@ async function handleResponsesInner( response, declaredWireToolNames, declaredNamelessClientCallTypes, + providerExecutedCallTypes, ) !== undefined ) { return; @@ -3725,6 +3734,7 @@ async function handleResponsesInner( ? createUndeclaredToolCallGuardBlockRewrite( declaredWireToolNames, declaredNamelessClientCallTypes, + providerExecutedCallTypes, ) : undefined, ].filter((rewrite): rewrite is NonNullable => rewrite !== undefined); @@ -3946,6 +3956,7 @@ async function handleResponsesInner( JSON.parse(clientJson), declaredWireToolNames, declaredNamelessClientCallTypes, + providerExecutedCallTypes, ); } catch { return undefined; diff --git a/tests/responses-undeclared-tool-guard.test.ts b/tests/responses-undeclared-tool-guard.test.ts index f92b9fd0838..1e93fb714dd 100644 --- a/tests/responses-undeclared-tool-guard.test.ts +++ b/tests/responses-undeclared-tool-guard.test.ts @@ -8,11 +8,13 @@ import { describe, expect, test } from "bun:test"; import { collectDeclaredNamelessClientCallTypes, collectDeclaredWireToolNames, + collectProviderExecutedCallTypes, createUndeclaredToolCallGuardBlockRewrite, currentTurnWireToolCatalogBody, hasExplicitWireToolCatalog, undeclaredToolCallNameInResponse, UNDECLARED_TOOL_CALL_ERROR_CODE, + type ProviderExecutedCallType, } from "../src/server/responses-undeclared-tool-guard"; import { relaySseWithBlockRewrite } from "../src/server/sse-payload-rewrite"; import { handleResponses } from "../src/server/responses"; @@ -1248,3 +1250,129 @@ describe("undeclaredToolCallNameInResponse", () => { )).toBeUndefined(); }); }); + +/** + * xAI runs hosted `x_search` itself and reports the activity as a `custom_tool_call` whose name + * is absent from the request catalog. Probed 2026-08-23 against the OAuth CLI destination: the + * provider's hosted calls carry an `xs_call-` call-id prefix. Observed names were + * `x_keyword_search`, `x_semantic_search`, and `x_user_search`, so authorization keys on the + * declaration, item type, and call-id prefix, never on the name. + */ +describe("provider-executed hosted calls", () => { + const declared = new Set(["shell"]); + const nameless = new Set(); + const xSearchAuthorized = collectProviderExecutedCallTypes({ + tools: [{ type: "function", name: "shell" }, { type: "x_search" }], + }); + + function hostedCall(name: string) { + return { output: [{ type: "custom_tool_call", name, call_id: "xs_call-1" }] }; + } + + test("authorizes the provider's hosted call under any of its observed names", () => { + expect(collectProviderExecutedCallTypes({ tools: [{ type: "x_search" }] })) + .toEqual(new Set([{ itemType: "custom_tool_call", callIdPrefix: "xs_call-" }])); + for (const name of ["x_keyword_search", "x_semantic_search", "x_user_search"]) { + expect(undeclaredToolCallNameInResponse( + hostedCall(name), declared, nameless, xSearchAuthorized, + )).toBeUndefined(); + } + }); + + test("without the x_search declaration the same item is still refused", () => { + const noHostedDeclaration = collectProviderExecutedCallTypes({ + tools: [{ type: "function", name: "shell" }], + }); + expect(noHostedDeclaration.size).toBe(0); + expect(undeclaredToolCallNameInResponse( + hostedCall("x_keyword_search"), declared, nameless, noHostedDeclaration, + )).toBe("x_keyword_search"); + }); + + test("the caller gates on destination: an empty authorization set refuses the same item", () => { + // core.ts passes an empty set unless the route actually terminates at xAI, so a declaration + // alone cannot buy the exemption on an upstream that never serves the hosted tool. + expect(undeclaredToolCallNameInResponse( + hostedCall("x_keyword_search"), declared, nameless, new Set(), + )).toBe("x_keyword_search"); + }); + + test("#1700 still holds: an undeclared client tool is refused inside an authorized turn", () => { + expect(undeclaredToolCallNameInResponse( + { output: [{ type: "custom_tool_call", name: "apply_patch", call_id: "call_patch" }] }, + declared, nameless, xSearchAuthorized, + )).toBe("apply_patch"); + }); + + test("a declared client tool is unaffected", () => { + expect(undeclaredToolCallNameInResponse( + { output: [{ type: "function_call", name: "shell", call_id: "c1" }] }, + declared, nameless, xSearchAuthorized, + )).toBeUndefined(); + }); +}); + +describe("xAI hosted-call authorization through handleResponses", () => { + const hostedCall = { + type: "custom_tool_call", + id: "ctc_search", + call_id: "xs_call-1", + name: "x_keyword_search", + input: "{}", + status: "completed", + }; + + async function post(baseUrl: string): Promise { + const config = { + port: 0, + defaultProvider: "fixture", + providers: { + fixture: { + adapter: "openai-responses", + baseUrl, + authMode: "key", + apiKey: "fixture-key", + }, + }, + } as OcxConfig; + const savedFetch = globalThis.fetch; + globalThis.fetch = (async () => Response.json({ + id: "resp_search", + status: "completed", + output: [hostedCall], + })) as typeof fetch; + try { + return await handleResponses(new Request("http://localhost/v1/responses", { + method: "POST", + headers: { "content-type": "application/json" }, + body: JSON.stringify({ + model: "fixture/grok-4.6", + stream: false, + input: [{ role: "user", content: [{ type: "input_text", text: "search" }] }], + tools: [ + { type: "function", name: "shell", parameters: { type: "object" } }, + { type: "x_search" }, + ], + }), + }), config, { model: "", provider: "" }); + } finally { + globalThis.fetch = savedFetch; + } + } + + test("accepts the measured xs_call shape for an exact xAI destination", async () => { + const response = await post("https://api.x.ai/v1"); + + expect(response.status).toBe(200); + const body = await response.json() as { output: Array> }; + expect(body.output[0]).toMatchObject(hostedCall); + }); + + test("rejects the identical item for a lookalike destination", async () => { + const response = await post("https://api.x.ai.evil.test/v1"); + + expect(response.status).toBe(502); + const body = await response.json() as { error: { message: string } }; + expect(body.error.message).toContain('undeclared client tool "x_keyword_search"'); + }); +}); diff --git a/tests/xai-transport.test.ts b/tests/xai-transport.test.ts index 79e6cbf5e1c..38cb29142f3 100644 --- a/tests/xai-transport.test.ts +++ b/tests/xai-transport.test.ts @@ -3,6 +3,7 @@ import { createOpenAIChatAdapter } from "../src/adapters/openai-chat"; import { parseRequest } from "../src/responses/parser"; import { buildModelsRequest } from "../src/oauth"; import { + isXaiResponsesDestination, resolveProviderTransport, deriveXaiConvId, XAI_CONV_ID_HEADER, @@ -51,6 +52,27 @@ function parsed(): OcxParsedRequest { }; } +describe("xAI Responses destination detection", () => { + test.each([ + "https://api.x.ai/v1", + "https://api.x.ai:443/v1", + XAI_GROK_CLI_BASE_URL, + "https://CLI-CHAT-PROXY.GROK.COM:443/v1", + ])("accepts the exact xAI HTTPS destination %s", baseUrl => { + expect(isXaiResponsesDestination({ baseUrl })).toBe(true); + }); + + test.each([ + "http://api.x.ai/v1", + "https://api.x.ai:444/v1", + "https://api.x.ai.evil.test/v1", + "https://cli-chat-proxy.grok.com.evil.test/v1", + "not a URL", + ])("rejects a non-xAI or malformed destination %s", baseUrl => { + expect(isXaiResponsesDestination({ baseUrl })).toBe(false); + }); +}); + describe("xAI auth-mode transport selection", () => { test("OAuth selects the Grok CLI subscription transport and required headers", () => { const effective = resolveProviderTransport("xai", provider("oauth")); From c09f040db348377cbc81e916dc5fee1a01f2b300 Mon Sep 17 00:00:00 2001 From: luvs01 Date: Tue, 25 Aug 2026 12:23:27 +0900 Subject: [PATCH 05/29] fix(gui): guard quota reset date formatting (#2405) Co-authored-by: luvs01 <27862058+luvs01@users.noreply.github.com> --- gui/src/components/QuotaBars.tsx | 19 ++++++++++++++----- tests/quota-bars-rows.test.ts | 1 + 2 files changed, 15 insertions(+), 5 deletions(-) diff --git a/gui/src/components/QuotaBars.tsx b/gui/src/components/QuotaBars.tsx index 0402936791f..3c80e0c5200 100644 --- a/gui/src/components/QuotaBars.tsx +++ b/gui/src/components/QuotaBars.tsx @@ -349,10 +349,19 @@ function StackedQuotaRow({ row, threshold, t, locale, incomplete }: { ); } -function formatResetAt(resetAt: number | undefined, t: TFn, locale: Locale): { day: string; time: string } { - if (typeof resetAt !== "number" || !Number.isFinite(resetAt)) return { day: "", time: "" }; +/** Normalize seconds-or-milliseconds epochs and reject values outside JavaScript Date's range. */ +function resetDate(resetAt: number | undefined): { date: Date; ms: number } | null { + if (typeof resetAt !== "number" || !Number.isFinite(resetAt)) return null; const ms = resetAt < 10_000_000_000 ? resetAt * 1000 : resetAt; const date = new Date(ms); + if (!Number.isFinite(date.getTime())) return null; + return { date, ms }; +} + +function formatResetAt(resetAt: number | undefined, t: TFn, locale: Locale): { day: string; time: string } { + const normalized = resetDate(resetAt); + if (!normalized) return { day: "", time: "" }; + const { date } = normalized; const now = new Date(); const tag = bcp47(locale); const time = new Intl.DateTimeFormat(tag, { hour: "2-digit", minute: "2-digit", hour12: false }).format(date); @@ -371,9 +380,9 @@ export function formatResetFuture( locale: Locale = "en", now = Date.now(), ): string { - if (typeof resetAt !== "number" || !Number.isFinite(resetAt)) return ""; - const ms = resetAt < 10_000_000_000 ? resetAt * 1000 : resetAt; - const date = new Date(ms); + const normalized = resetDate(resetAt); + if (!normalized) return ""; + const { date, ms } = normalized; const tag = bcp47(locale); const time = new Intl.DateTimeFormat(tag, { hour: "2-digit", minute: "2-digit", hour12: false }).format(date); const nowDate = new Date(now); diff --git a/tests/quota-bars-rows.test.ts b/tests/quota-bars-rows.test.ts index b7ec117d121..2f8b8de5bff 100644 --- a/tests/quota-bars-rows.test.ts +++ b/tests/quota-bars-rows.test.ts @@ -143,6 +143,7 @@ describe("formatResetFuture", () => { expect(formatResetFuture(NOW - 86_400_000, t, "en", NOW)).toContain("quota.resetsAt"); expect(formatResetFuture(undefined, t, "en", NOW)).toBe(""); expect(formatResetFuture(Number.NaN, t, "en", NOW)).toBe(""); + expect(formatResetFuture(Number.MAX_VALUE, t, "en", NOW)).toBe(""); }); test("seconds-epoch inputs are normalized to milliseconds", () => { From 312b3e7b66d9af9b86cd607c4162dbc7db3bad98 Mon Sep 17 00:00:00 2001 From: luvs01 Date: Tue, 25 Aug 2026 12:23:34 +0900 Subject: [PATCH 06/29] fix(gui): ignore stale Startup secondary responses (#2416) Co-authored-by: luvs01 <27862058+luvs01@users.noreply.github.com> --- gui/src/pages/Startup.tsx | 8 ++++- gui/tests/startup-revisit-cache.test.tsx | 45 ++++++++++++++++++++++++ 2 files changed, 52 insertions(+), 1 deletion(-) diff --git a/gui/src/pages/Startup.tsx b/gui/src/pages/Startup.tsx index 8bb03166c01..53b997fac3c 100644 --- a/gui/src/pages/Startup.tsx +++ b/gui/src/pages/Startup.tsx @@ -88,8 +88,14 @@ export default function Startup({ apiBase }: { apiBase: string }) { /** True while settings (runtime notice) are still in flight — reserves notice slot height. */ const [runtimeNoticePending, setRuntimeNoticePending] = useState(() => !cached?.data); const paintedRef = useRef(Boolean(cached?.data)); + const secondaryGenerationRef = useRef(0); + + useEffect(() => () => { + secondaryGenerationRef.current += 1; + }, [apiBase]); const fetchStartup = useCallback(async (signal: AbortSignal): Promise => { + const secondaryGeneration = ++secondaryGenerationRef.current; const keepSecondary = paintedRef.current; // Keep prior notice/tray visible on revalidation; only reserve empty slots on first paint. if (!keepSecondary) { @@ -154,7 +160,7 @@ export default function Startup({ apiBase }: { apiBase: string }) { // Health drives the main page, so publish it before the lower-priority settings/tray // requests finish. Their result updates the existing reserved slots independently. void Promise.all([settingsPromise, trayPromise]).then(([settings, trayResult]) => { - if (signal.aborted) return; + if (signal.aborted || secondaryGeneration !== secondaryGenerationRef.current) return; const nextTray = next.platform === "win32" ? trayResult.tray : null; if (next.platform === "win32") { setTray(nextTray); diff --git a/gui/tests/startup-revisit-cache.test.tsx b/gui/tests/startup-revisit-cache.test.tsx index 11a0ef7f6c3..44674f170ec 100644 --- a/gui/tests/startup-revisit-cache.test.tsx +++ b/gui/tests/startup-revisit-cache.test.tsx @@ -108,3 +108,48 @@ test("a revisit with session cache keeps Action required visible without a loadi await act(async () => { root.unmount(); }); container.remove(); }); + +test("a superseded settings response cannot overwrite newer Startup cache", async () => { + const { createRoot } = await import("react-dom/client"); + const container = document.createElement("div"); + document.body.append(container); + + let settingsCalls = 0; + let resolveStaleSettings!: (response: Response) => void; + globalThis.fetch = (async (input: RequestInfo | URL) => { + const url = String(input); + if (url.includes("/api/startup-health")) return Response.json(atRiskHealth()); + if (!url.includes("/api/settings")) return new Response(null, { status: 404 }); + settingsCalls += 1; + if (settingsCalls === 1) { + return await new Promise(resolve => { resolveStaleSettings = resolve; }); + } + return Response.json({ codexRuntime: { version: "fresh", newerAvailable: { version: "new" } } }); + }) as typeof fetch; + + let root!: Root; + await act(async () => { + root = createRoot(container); + root.render(); + }); + await act(async () => { await new Promise(r => testWindow.setTimeout(r, 20)); }); + + const refresh = Array.from(container.querySelectorAll("button")) + .find(button => button.textContent?.includes("Refresh")); + expect(refresh).toBeDefined(); + await act(async () => { refresh?.click(); }); + await act(async () => { await new Promise(r => testWindow.setTimeout(r, 20)); }); + expect(settingsCalls).toBe(2); + expect(testWindow.sessionStorage.getItem(CACHE_KEY)).toContain("fresh"); + + await act(async () => { + resolveStaleSettings(Response.json({ codexRuntime: { version: "stale", newerAvailable: { version: "new" } } })); + await new Promise(r => testWindow.setTimeout(r, 20)); + }); + + expect(testWindow.sessionStorage.getItem(CACHE_KEY)).toContain("fresh"); + expect(testWindow.sessionStorage.getItem(CACHE_KEY)).not.toContain("stale"); + + await act(async () => { root.unmount(); }); + container.remove(); +}); From d402272189264806486d53fb572f73fefaa6badb Mon Sep 17 00:00:00 2001 From: Bohdan Date: Tue, 25 Aug 2026 05:24:18 +0200 Subject: [PATCH 07/29] fix(responses): retry pre-output EOFs affecting Ox Alpha (#2486) * fix(responses): retry pre-output EOFs * docs: sync empty completion retry translations --- .../docs/ja/reference/configuration/server.md | 2 +- .../docs/ko/reference/configuration/server.md | 2 +- .../docs/reference/configuration/server.md | 2 +- .../docs/ru/reference/configuration/server.md | 2 +- .../zh-cn/reference/configuration/server.md | 2 +- .../responses/empty-completion-guard.ts | 34 +++++++++++--- structure/04_transports-and-sidecars.md | 18 ++++---- tests/empty-completion-guard.test.ts | 46 +++++++++++++++++-- 8 files changed, 86 insertions(+), 22 deletions(-) diff --git a/docs-site/src/content/docs/ja/reference/configuration/server.md b/docs-site/src/content/docs/ja/reference/configuration/server.md index 72e372e8f2c..d31c8db6fb0 100644 --- a/docs-site/src/content/docs/ja/reference/configuration/server.md +++ b/docs-site/src/content/docs/ja/reference/configuration/server.md @@ -12,7 +12,7 @@ description: リスナー、リモート アクセス、アドミッション | `port` | `number` | `10100` |プロキシリッスンポート。 | | `hostname?` | `string` | `"127.0.0.1"` |バインドアドレス。非ループバック バインドには `OPENCODEX_API_AUTH_TOKEN` が必要です。 | | `proxy?` | `string` | — |送信 HTTP(S) プロキシ URL または `${ENV_VAR}`。これらの変数が設定されていない場合にのみ、`HTTP_PROXY` / `HTTPS_PROXY` に適用されます。ループバックは `NO_PROXY` に残ります。 | -| `emptyCompletionRetry?` | `boolean` | `false` | テキストもツール呼び出しもない Responses 完了を、同一リクエストで 1 回再試行するよう明示的に有効化します。再試行は課金対象になる場合があります。`OCX_EMPTY_COMPLETION_RETRY=0` で設定を変更せず無効化できます。combo と routed-compaction turn は対象外です。 | +| `emptyCompletionRetry?` | `boolean` | `false` | テキストもツール呼び出しもない Responses ターンを、ターミナルイベント前にストリームが終了した場合も含め、同一リクエストで 1 回再試行するよう明示的に有効化します。再試行は課金対象になる場合があります。`OCX_EMPTY_COMPLETION_RETRY=0` で設定を変更せず無効化できます。combo と routed-compaction turn は対象外です。 | | `stallTimeoutSec?` | `number` | `300` | `response.incomplete` より前にアップストリーム データがない秒数。最小 1。 | `connectTimeoutMs?` | `number` | `200000` |試行ごとの DNS/TCP/TLS/最終ヘッダーの期限。本体が生成される前に終了します。 | | `shutdownTimeoutMs?` | `number` | `5000` |アクティブなターンが中止される前の正常な排出期限。 | diff --git a/docs-site/src/content/docs/ko/reference/configuration/server.md b/docs-site/src/content/docs/ko/reference/configuration/server.md index 28a358862a1..79caa1fe872 100644 --- a/docs-site/src/content/docs/ko/reference/configuration/server.md +++ b/docs-site/src/content/docs/ko/reference/configuration/server.md @@ -12,7 +12,7 @@ description: 리스너, 원격 접근, admission 키, 타임아웃, 저장소, | `port` | `number` | `10100` | 프록시 수신 포트입니다. | | `hostname?` | `string` | `"127.0.0.1"` | 바인드 주소입니다. 루프백이 아닌 바인드에는 `OPENCODEX_API_AUTH_TOKEN`이 필요합니다. | | `proxy?` | `string` | — | 송신용 HTTP(S) 프록시 URL 또는 `${ENV_VAR}`입니다. 해당 변수가 비어 있을 때만 `HTTP_PROXY` / `HTTPS_PROXY`에 적용되며, 루프백은 `NO_PROXY`에 그대로 남습니다. | -| `emptyCompletionRetry?` | `boolean` | `false` | 텍스트나 도구 호출 없이 완료된 Responses 요청을 한 번 동일하게 재시도하도록 선택합니다. 재시도에는 비용이 발생할 수 있습니다. `OCX_EMPTY_COMPLETION_RETRY=0`은 설정을 바꾸지 않고 비활성화하며, combo 및 routed-compaction turn은 제외됩니다. | +| `emptyCompletionRetry?` | `boolean` | `false` | 텍스트나 도구 호출이 없는 Responses 턴을, 터미널 이벤트 전에 스트림이 종료된 경우를 포함해 동일한 요청으로 한 번 재시도하도록 선택합니다. 재시도에는 비용이 발생할 수 있습니다. `OCX_EMPTY_COMPLETION_RETRY=0`은 설정을 바꾸지 않고 비활성화하며, combo 및 routed-compaction turn은 제외됩니다. | | `stallTimeoutSec?` | `number` | `300` | 업스트림 데이터가 없을 때 `response.incomplete`가 되기까지의 초 수입니다. 최소 1입니다. | | `connectTimeoutMs?` | `number` | `200000` | 시도별 DNS/TCP/TLS/최종 헤더 기한입니다. 본문 생성 전에 끝납니다. | | `shutdownTimeoutMs?` | `number` | `5000` | 진행 중인 turn을 중단하기 전에 허용하는 정상 종료 드레인 기한입니다. | diff --git a/docs-site/src/content/docs/reference/configuration/server.md b/docs-site/src/content/docs/reference/configuration/server.md index fe12dc58a37..a5e2a3af614 100644 --- a/docs-site/src/content/docs/reference/configuration/server.md +++ b/docs-site/src/content/docs/reference/configuration/server.md @@ -13,7 +13,7 @@ runs helper features around provider requests. | `port` | `number` | `10100` | Proxy listen port. | | `hostname?` | `string` | `"127.0.0.1"` | Bind address. Non-loopback binds require `OPENCODEX_API_AUTH_TOKEN`. | | `proxy?` | `string` | — | Outbound HTTP(S) proxy URL or `${ENV_VAR}`. Applied to `HTTP_PROXY` / `HTTPS_PROXY` only when those variables are unset; loopback remains in `NO_PROXY`. | -| `emptyCompletionRetry?` | `boolean` | `false` | Opt in to one identical Responses retry when a completion has no text or tool call. The retry may be billable. `OCX_EMPTY_COMPLETION_RETRY=0` disables it without changing config; combo and routed-compaction turns remain excluded. | +| `emptyCompletionRetry?` | `boolean` | `false` | Opt in to one identical Responses retry when a turn has no text or tool call, including a stream that ends before a terminal event. The retry may be billable. `OCX_EMPTY_COMPLETION_RETRY=0` disables it without changing config; combo and routed-compaction turns remain excluded. | | `stallTimeoutSec?` | `number` | `300` | Seconds without upstream data before `response.incomplete`. Minimum 1. | | `connectTimeoutMs?` | `number` | `200000` | Per-attempt DNS/TCP/TLS/final-header deadline; it ends before body generation. | | `shutdownTimeoutMs?` | `number` | `5000` | Graceful drain deadline before active turns are aborted. | diff --git a/docs-site/src/content/docs/ru/reference/configuration/server.md b/docs-site/src/content/docs/ru/reference/configuration/server.md index d8d7b5d6f95..3e650d0b7f2 100644 --- a/docs-site/src/content/docs/ru/reference/configuration/server.md +++ b/docs-site/src/content/docs/ru/reference/configuration/server.md @@ -13,7 +13,7 @@ description: Listener, удалённый доступ, admission key, тайм | `port` | `number` | `10100` | Порт, который слушает прокси. | | `hostname?` | `string` | `"127.0.0.1"` | Адрес bind'а. Не-loopback bind требует `OPENCODEX_API_AUTH_TOKEN`. | | `proxy?` | `string` | — | URL исходящего HTTP(S)-прокси или `${ENV_VAR}`. Применяется к `HTTP_PROXY` / `HTTPS_PROXY` только когда эти переменные не заданы; loopback всегда остаётся в `NO_PROXY`. | -| `emptyCompletionRetry?` | `boolean` | `false` | Явно включает один идентичный повтор Responses, если completion не содержит ни текста, ни tool call. Повтор может тарифицироваться. `OCX_EMPTY_COMPLETION_RETRY=0` отключает его без изменения config; combo и routed-compaction turn исключены. | +| `emptyCompletionRetry?` | `boolean` | `false` | Явно включает один идентичный повтор Responses, если в turn нет ни текста, ни tool call, включая случай, когда stream завершается до terminal event. Повтор может тарифицироваться. `OCX_EMPTY_COMPLETION_RETRY=0` отключает его без изменения config; combo и routed-compaction turn исключены. | | `stallTimeoutSec?` | `number` | `300` | Секунды без upstream-данных до `response.incomplete`. Минимум 1. | | `connectTimeoutMs?` | `number` | `200000` | Дедлайн одной попытки DNS/TCP/TLS/final-header; он завершается до генерации тела ответа. | | `shutdownTimeoutMs?` | `number` | `5000` | Дедлайн graceful-drain до принудительного прерывания активных turn'ов. | diff --git a/docs-site/src/content/docs/zh-cn/reference/configuration/server.md b/docs-site/src/content/docs/zh-cn/reference/configuration/server.md index d4a6a3fc643..c9753f58bb6 100644 --- a/docs-site/src/content/docs/zh-cn/reference/configuration/server.md +++ b/docs-site/src/content/docs/zh-cn/reference/configuration/server.md @@ -13,7 +13,7 @@ description: 监听、远程访问、准入密钥、超时、存储、侧车、 | `port` | `number` | `10100` | 代理监听端口。 | | `hostname?` | `string` | `"127.0.0.1"` | 绑定地址。非回环绑定需要 `OPENCODEX_API_AUTH_TOKEN`。 | | `proxy?` | `string` | — | 出站 HTTP(S) 代理 URL,或 `${ENV_VAR}`。仅当 `HTTP_PROXY` / `HTTPS_PROXY` 未设置时才会应用;回环地址始终保留在 `NO_PROXY` 中。 | -| `emptyCompletionRetry?` | `boolean` | `false` | 显式启用:当 Responses 完成时既无文本也无工具调用,使用相同请求重试一次。重试可能产生费用。`OCX_EMPTY_COMPLETION_RETRY=0` 可在不修改配置的情况下禁用;combo 与 routed-compaction turn 不参与。 | +| `emptyCompletionRetry?` | `boolean` | `false` | 显式启用:当 Responses turn 既无文本也无工具调用时,使用相同请求重试一次,包括流在终止事件之前结束的情况。重试可能产生费用。`OCX_EMPTY_COMPLETION_RETRY=0` 可在不修改配置的情况下禁用;combo 与 routed-compaction turn 不参与。 | | `stallTimeoutSec?` | `number` | `300` | 在上游没有数据之前可等待的秒数,超过后返回 `response.incomplete`。最小值为 1。 | | `connectTimeoutMs?` | `number` | `200000` | 每次尝试的 DNS/TCP/TLS/最终响应头截止时间;它在正文生成之前结束。 | | `shutdownTimeoutMs?` | `number` | `5000` | 优雅停机截止时间,超过后会中止仍在进行中的请求。 | diff --git a/src/server/responses/empty-completion-guard.ts b/src/server/responses/empty-completion-guard.ts index 84bbb5632e4..6195ad0b8b6 100644 --- a/src/server/responses/empty-completion-guard.ts +++ b/src/server/responses/empty-completion-guard.ts @@ -153,9 +153,10 @@ export interface EmptyCompletionGuardOptions { * Watch an adapter event stream for the empty-completion failure mode. Events * are held until the turn produces content or ends: reasoning and other * pre-content events stay buffered (released in order on first content), the - * terminal is withheld, and an empty terminal triggers one identical-turn - * retry through `continuation`. Usage is merged across attempts so the bridge - * and request log meter the whole turn, not just the attempt that succeeded. + * terminal is withheld, and an empty terminal or pre-output EOF triggers one + * identical-turn retry through `continuation`. Usage is merged across attempts + * so the bridge and request log meter the whole turn, not just the attempt that + * succeeded. * * Heartbeats always pass through untouched: they feed the bridge's stall * watchdog, so holding them behind the content gate would trip false @@ -194,7 +195,11 @@ export async function* guardEmptyCompletionEventStream( if (sawContent || passthrough) { // Buffered content is already flowing; everything downstream passes // through. Every terminal carries usage merged across every attempt. - yield isTerminalEvent(event) ? withUsage(event) : event; + if (isTerminalEvent(event)) { + yield withUsage(event); + return; + } + yield event; continue; } if (isContentEvent(event)) { @@ -267,8 +272,25 @@ export async function* guardEmptyCompletionEventStream( if (isReasoningEvent(event)) yield { type: "heartbeat" }; } if (!terminalSeen) { - // The source ended without a terminal event (truncated stream). Release - // what was held so the bridge can mark the stream incomplete. + // A terminal-less EOF before text or a tool call is replay-safe: nothing + // actionable reached the client. Retry once, then surface a stated error + // instead of letting the bridge reduce the turn to adapter_eof. + if (!sawContent && !passthrough && retries < maxRetries) { + retries += 1; + try { + source = await options.continuation(); + } catch { + yield emptyCompletionRetryFailedEvent(usage, true); + return; + } + continue; + } + if (!sawContent && retries > 0) { + yield emptyCompletionRetryFailedEvent(usage, true); + return; + } + // Post-output EOF remains incomplete; replaying could duplicate text or + // executable tool calls. yield* releaseHeld(); return; } diff --git a/structure/04_transports-and-sidecars.md b/structure/04_transports-and-sidecars.md index 51d06726213..f7b9b1859bc 100644 --- a/structure/04_transports-and-sidecars.md +++ b/structure/04_transports-and-sidecars.md @@ -491,14 +491,16 @@ configurable via `stallTimeoutSec`, checked on the 2 s heartbeat tick) closes th `response.incomplete` / `upstream_stall_timeout` and cancels the upstream request if no real adapter events arrive. Adapter-yielded `{ type: "heartbeat" }` events DO reset the watchdog. -Top-level `emptyCompletionRetry: true` opts Responses turns into one identical replay when a -successful upstream completion contains neither output text nor a tool call. The default is off -because the replay may be billable; `OCX_EMPTY_COMPLETION_RETRY=0` is a disable-only emergency -override. Streaming and buffered HTTP adapters plus `runTurn` transports share the same guard, -while combo attempts and routed compaction stay excluded. Pre-content reasoning is retained under -named event-count and byte caps and emits liveness heartbeats while held. A second empty result or -retry failure becomes typed 502 `empty_completion_retry_failed`; usage is merged across sends, and -the Logs attempt records recovery kind `empty-completion`. +Top-level `emptyCompletionRetry: true` opts Responses turns into one identical replay when an +upstream turn produces neither output text nor a tool call, including a stream that ends before a +terminal event. A terminal-less stream is replayed only before actionable output; post-output EOF +remains incomplete so text or tool calls cannot be duplicated. The default is off because the replay +may be billable; `OCX_EMPTY_COMPLETION_RETRY=0` is a disable-only emergency override. Streaming and +buffered HTTP adapters plus `runTurn` transports share the same guard, while combo attempts and +routed compaction stay excluded. Pre-content reasoning is retained under named event-count and byte +caps and emits liveness heartbeats while held. A second empty result or retry failure becomes typed +502 `empty_completion_retry_failed`; usage is merged across sends, and the Logs attempt records +recovery kind `empty-completion`. The web-search loop requests `stream: true` for every routed-model iteration, but buffers the events needed to decide whether to intercept a synthetic search call. Text explicitly phased as diff --git a/tests/empty-completion-guard.test.ts b/tests/empty-completion-guard.test.ts index 3265521b3a1..2b786722bf4 100644 --- a/tests/empty-completion-guard.test.ts +++ b/tests/empty-completion-guard.test.ts @@ -322,17 +322,57 @@ describe("empty-completion guard retry", () => { ]); }); - test("a truncated first source (no terminal) releases held events and ends", async () => { + test("a pre-output EOF retries once and succeeds", async () => { let continuations = 0; const events = await collect(guardEmptyCompletionEventStream({ firstEvents: eventsOf({ type: "thinking_delta", thinking: "..." }), continuation: () => { continuations += 1; - return eventsOf(); + return eventsOf( + { type: "text_delta", text: "recovered" }, + { type: "done" }, + ); + }, + })); + + expect(continuations).toBe(1); + expect(withoutHeartbeats(events)).toEqual([ + { type: "thinking_delta", thinking: "..." }, + { type: "text_delta", text: "recovered" }, + { type: "done" }, + ]); + }); + + test("a second pre-output EOF surfaces empty_completion_retry_failed", async () => { + let continuations = 0; + const events = await collect(guardEmptyCompletionEventStream({ + firstEvents: eventsOf({ type: "thinking_delta", thinking: "first" }), + continuation: () => { + continuations += 1; + return eventsOf({ type: "thinking_delta", thinking: "second" }); + }, + })); + + expect(continuations).toBe(1); + expect(withoutHeartbeats(events)).toEqual([ + expect.objectContaining({ + type: "error", + code: EMPTY_COMPLETION_RETRY_FAILED_CODE, + }), + ]); + }); + + test("a post-output EOF is not retried", async () => { + let continuations = 0; + const events = await collect(guardEmptyCompletionEventStream({ + firstEvents: eventsOf({ type: "text_delta", text: "partial" }), + continuation: () => { + continuations += 1; + return eventsOf({ type: "text_delta", text: "duplicate" }, { type: "done" }); }, })); expect(continuations).toBe(0); - expect(withoutHeartbeats(events)).toEqual([{ type: "thinking_delta", thinking: "..." }]); + expect(events).toEqual([{ type: "text_delta", text: "partial" }]); }); }); From 0f30b3959f5c29315bf98af01ee197be606fd88a Mon Sep 17 00:00:00 2001 From: Nguyen Thanh Dat Date: Tue, 25 Aug 2026 10:24:24 +0700 Subject: [PATCH 08/29] fix(claude): do not let a non-registering row veto a bare context key (#2485) buildClaudeContextWindows registers a bare routed id only when it is unambiguous across providers. The count that decides that is taken over every routed model, including the ones the loop right below then skips: for (const m of routedModels) bareCounts.set(m.id, ...) for (const m of routedModels) { if (typeof window !== "number" || window <= 0) continue; if (m.provider === "anthropic" && window < ONE_MILLION) continue; ... if (bareCounts.get(m.id) === 1) put(m.id, window); } A skipped row contributes no window, so it cannot disagree with anything - but it still pushes the count to 2 and withholds the key. Measured: [{a/m: 1_000_000}, {b/m: no contextWindow}] -> bare "m" absent [{openrouter/claude-x: 1M}, {anthropic/claude-x: 200k}] -> bare absent The map then has exactly one authoritative answer and refuses to give it. A Claude Code slot set to the bare id resolves to no window, so shouldMarkOneMillion returns false and a genuine 1M model loses its [1m] marker and its 1M accounting, because some unrelated provider happens to list the same id. Count over the rows that can actually claim the key. Two providers that both register still withhold the bare id - that ambiguity is real, and the audit #5 case is unchanged. --- src/claude/context-windows.ts | 25 ++++++++++++------- tests/claude-context-windows.test.ts | 37 ++++++++++++++++++++++++++++ 2 files changed, 53 insertions(+), 9 deletions(-) diff --git a/src/claude/context-windows.ts b/src/claude/context-windows.ts index da2023f1365..dd22fc93a9a 100644 --- a/src/claude/context-windows.ts +++ b/src/claude/context-windows.ts @@ -118,17 +118,24 @@ export function buildClaudeContextWindows( put(desktop3pAlias("native", slug), window); put(aliasForNative(slug), window); } + // Anthropic passthrough guard (audit 021 #3): canonical claude ids ride the + // subscription passthrough — marking a sub-1M one would strap [1m]/1M-beta onto + // a model that cannot host it. Register anthropic rows only at >=1M. + const registrable = routedModels.filter( + m => + typeof m.contextWindow === "number" && + m.contextWindow > 0 && + !(m.provider === "anthropic" && m.contextWindow < ONE_MILLION), + ); // Bare routed ids are registered only when unambiguous across providers (audit - // 021 #5) — natives are registered first, so a native slug always wins the bare key. + // 021 #5) — natives are registered first, so a native slug always wins the bare + // key. Counted over the rows that can actually claim the key: a row this loop + // skips contributes no window, so letting it veto the bare key withholds an + // answer that was never in doubt. const bareCounts = new Map(); - for (const m of routedModels) bareCounts.set(m.id, (bareCounts.get(m.id) ?? 0) + 1); - for (const m of routedModels) { - const window = m.contextWindow; - if (typeof window !== "number" || window <= 0) continue; - // Anthropic passthrough guard (audit 021 #3): canonical claude ids ride the - // subscription passthrough — marking a sub-1M one would strap [1m]/1M-beta onto - // a model that cannot host it. Register anthropic rows only at >=1M. - if (m.provider === "anthropic" && window < ONE_MILLION) continue; + for (const m of registrable) bareCounts.set(m.id, (bareCounts.get(m.id) ?? 0) + 1); + for (const m of registrable) { + const window = m.contextWindow as number; put(`${m.provider}/${m.id}`, window); put(desktop3pAlias(m.provider, m.id), window); put(aliasForRoute(m.provider, m.id), window); diff --git a/tests/claude-context-windows.test.ts b/tests/claude-context-windows.test.ts index 8f364b7636d..85c70d321e8 100644 --- a/tests/claude-context-windows.test.ts +++ b/tests/claude-context-windows.test.ts @@ -1,6 +1,7 @@ import { describe, expect, test } from "bun:test"; import { AUTO_COMPACT_WINDOW_DEFAULT, boundedContextWindows, buildClaudeContextWindows, effectiveModelEnv, resolveAutoContext, shouldMarkOneMillion, withOneMillionMarker } from "../src/claude/context-windows"; import { desktop3pAlias } from "../src/claude/desktop-3p"; +import type { CatalogModel } from "../src/codex/catalog"; describe("claude context-window map (devlog 260712 B2)", () => { const routed = [ @@ -132,6 +133,42 @@ describe("auto-context (devlog 260712 020 + audit 021)", () => { expect(map["gpt-5.6-sol"]).toBe(272_000); // native default, not 999k }); + test("a row that registers nothing does not make a bare id ambiguous", () => { + // Only one of these two rows can claim the bare key, so there is nothing to + // be ambiguous about — withholding it left a 1M model with no window, and a + // slot set to the bare id lost its [1m] marker. + const noWindow = buildClaudeContextWindows([], [ + { provider: "a", id: "shared-model", contextWindow: 1_000_000 }, + { provider: "b", id: "shared-model" } as CatalogModel, + ]); + expect(noWindow["shared-model"]).toBe(1_000_000); + + // Same for a row the anthropic sub-1M guard skips. + const anthropicSkipped = buildClaudeContextWindows([], [ + { provider: "openrouter", id: "claude-x", contextWindow: 1_000_000 }, + { provider: "anthropic", id: "claude-x", contextWindow: 200_000 }, + ]); + expect(anthropicSkipped["claude-x"]).toBe(1_000_000); + expect(anthropicSkipped["anthropic/claude-x"]).toBeUndefined(); + + // A zero or negative window is not a claim either. + const zeroWindow = buildClaudeContextWindows([], [ + { provider: "a", id: "shared-model", contextWindow: 400_000 }, + { provider: "b", id: "shared-model", contextWindow: 0 }, + ]); + expect(zeroWindow["shared-model"]).toBe(400_000); + }); + + test("two providers that both register keep the bare id withheld", () => { + const map = buildClaudeContextWindows([], [ + { provider: "a", id: "shared-model", contextWindow: 300_000 }, + { provider: "b", id: "shared-model", contextWindow: 900_000 }, + ]); + expect(map["shared-model"]).toBeUndefined(); + expect(map["a/shared-model"]).toBe(300_000); + expect(map["b/shared-model"]).toBe(900_000); + }); + test("auto-context marks a wide native slot, and turning it off unmarks anything under 1M", () => { const windows = buildClaudeContextWindows(["gpt-5.6-sol"], []); const env = effectiveModelEnv({ model: "gpt-5.6-sol" }, windows); From 09062014ed4ff9ff2e200b8ce8970a4e22a4f4a1 Mon Sep 17 00:00:00 2001 From: Michael KIM Date: Tue, 25 Aug 2026 12:25:00 +0900 Subject: [PATCH 09/29] fix(kiro): prioritize tool search results within catalog budget (#2475) --- src/adapters/kiro-tools.ts | 29 ++++++++++++++++++++--------- tests/kiro-adapter.test.ts | 35 +++++++++++++++++++++++++++++++++++ 2 files changed, 55 insertions(+), 9 deletions(-) diff --git a/src/adapters/kiro-tools.ts b/src/adapters/kiro-tools.ts index 6aa8bb2a426..960a534387d 100644 --- a/src/adapters/kiro-tools.ts +++ b/src/adapters/kiro-tools.ts @@ -172,6 +172,12 @@ function omittedToolCatalogNotice(kept: number, omitted: readonly OcxTool[], reg return `[opencodex] Kiro's outbound catalog budget allows ${kept} of ${kept + omitted.length} client tools this turn. Omitted and unavailable this turn: ${summary}.`; } +function boundedCatalogPriority(tool: OcxTool): number { + if (tool.loadedFromToolSearch) return 0; + if (tool.toolSearch) return 1; + return 2; +} + export function convertKiroToolContext( parsed: OcxParsedRequest, registry: KiroToolNameRegistry = createKiroToolNameRegistry(), @@ -181,9 +187,7 @@ export function convertKiroToolContext( // Validate every listed name even when tool_choice:none emulates a tool-free turn. for (const tool of tools) registry.alias(namespacedToolName(tool.namespace, tool.name)); const effectiveTools = parsed.options.toolChoice === "none" ? [] : tools; - const convertedTools: unknown[] = []; - let omittedAt = effectiveTools.length; - for (const [index, tool] of effectiveTools.entries()) { + const convertedEntries = effectiveTools.map((tool, index) => { const description = tool.description || `Tool: ${tool.name}`; // Send the full namespaced wire name (e.g. mcp__chrome-devtools__navigate_page) so Kiro echoes // it back; the bridge's toolNsMap is keyed by this name and restores the MCP namespace Codex @@ -198,19 +202,26 @@ export function convertKiroToolContext( inputSchema: { json: ensureRootObjectType(sanitizeKiroSchema(tool.parameters ?? {})) }, }, }; - // Preserve declaration order and only omit a suffix. Ranking tools would make a catalog change - // silently alter which capability disappears; this deterministic policy is paired with a - // model-visible omission notice so unavailable tools are explicit rather than assumed absent. + return { tool, index, converted }; + }); + const exceedsBudget = convertedEntries.length > MAX_KIRO_TOOL_COUNT + || serializedToolCatalogBytes(convertedEntries.map(entry => entry.converted)) > MAX_KIRO_TOOL_CATALOG_BYTES; + const candidates = exceedsBudget + ? convertedEntries.toSorted((a, b) => boundedCatalogPriority(a.tool) - boundedCatalogPriority(b.tool) || a.index - b.index) + : convertedEntries; + const convertedTools: unknown[] = []; + let omittedAt = candidates.length; + for (const [index, entry] of candidates.entries()) { if ( convertedTools.length >= MAX_KIRO_TOOL_COUNT - || serializedToolCatalogBytes([...convertedTools, converted]) > MAX_KIRO_TOOL_CATALOG_BYTES + || serializedToolCatalogBytes([...convertedTools, entry.converted]) > MAX_KIRO_TOOL_CATALOG_BYTES ) { omittedAt = index; break; } - convertedTools.push(converted); + convertedTools.push(entry.converted); } - const omittedTools = effectiveTools.slice(omittedAt); + const omittedTools = candidates.slice(omittedAt).map(entry => entry.tool); return { tools: convertedTools, systemAdditions: omittedTools.length > 0 ? [omittedToolCatalogNotice(convertedTools.length, omittedTools, registry)] : [], diff --git a/tests/kiro-adapter.test.ts b/tests/kiro-adapter.test.ts index 770de3d8ded..9483ec9b3e1 100644 --- a/tests/kiro-adapter.test.ts +++ b/tests/kiro-adapter.test.ts @@ -724,6 +724,41 @@ describe("kiro adapter — buildRequest", () => { expect(current.content).toContain("Omitted and unavailable this turn"); }); + test("large catalogs prioritize tool-search discoveries and the search gateway", async () => { + const ordinaryTools = Array.from({ length: MAX_KIRO_TOOL_COUNT + 20 }, (_, index) => ({ + name: `ordinary_tool_${String(index).padStart(3, "0")}`, + description: `Ordinary tool ${index}`, + parameters: { type: "object" }, + })); + const searchGateway = { + name: "tool_search", + description: "Search deferred tools", + parameters: { type: "object" }, + toolSearch: true, + }; + const loadedTool = { + name: "codex_app__send_message_to_thread", + description: "Send a message to a task", + parameters: { type: "object" }, + loadedFromToolSearch: true, + }; + const tools = [...ordinaryTools, searchGateway, loadedTool]; + + const current = JSON.parse((await createKiroAdapter(provider).buildRequest( + parsedWith([{ role: "user", content: "hi" }], tools), + )).body).conversationState.currentMessage.userInputMessage; + const ordinary = current.userInputMessageContext.tools.slice(0, -1); + const names = ordinary.map((tool: { toolSpecification: { name: string } }) => tool.toolSpecification.name); + const omissionNotice = current.content.split("\n\n", 1)[0]; + + expect(ordinary).toHaveLength(MAX_KIRO_TOOL_COUNT); + expect(names.slice(0, 2)).toEqual([loadedTool.name, searchGateway.name]); + expect(names.slice(2)).toEqual(ordinaryTools.slice(0, MAX_KIRO_TOOL_COUNT - 2).map(tool => tool.name)); + expect(omissionNotice).toContain("ordinary_tool_046"); + expect(omissionNotice).not.toContain(loadedTool.name); + expect(omissionNotice).not.toContain(searchGateway.name); + }); + test("large catalogs retain the declared prefix within Kiro's serialized byte budget", async () => { // Top-level descriptions stay small, so existing description truncation cannot make this pass. // The repeated schema descriptions instead make the aggregate converted catalog exceed 96 KiB. From d659c542f7e33523747ec30a6b609329cf839fcf Mon Sep 17 00:00:00 2001 From: luvs01 Date: Tue, 25 Aug 2026 12:26:05 +0900 Subject: [PATCH 10/29] feat(codex): add per-model ChatGPT compaction budgets (#1905) * feat(codex): add per-model compaction budgets * docs(config): document per-model compact budgets * fix(codex): redact compaction budget validation names * fix(codex): require context window for combo compact budget * fix(codex): preserve lower native compact limits --------- Co-authored-by: luvs01 <27862058+luvs01@users.noreply.github.com> --- .../fr/reference/configuration/providers.md | 1 + .../ja/reference/configuration/providers.md | 1 + .../ko/reference/configuration/providers.md | 1 + .../docs/reference/configuration/providers.md | 1 + .../ru/reference/configuration/providers.md | 1 + .../tr/reference/configuration/providers.md | 1 + .../reference/configuration/providers.md | 1 + .../reference/configuration/providers.md | 1 + src/codex/catalog/aggregation.ts | 12 ++ src/codex/catalog/effort.ts | 21 ++- src/codex/catalog/metadata.ts | 28 ++- src/codex/catalog/parsing.ts | 55 +++--- src/codex/catalog/provider-fetch.ts | 168 ++++++++++++++--- src/codex/catalog/sync.ts | 2 +- src/codex/convergence.ts | 5 + src/config.ts | 12 ++ src/providers/auto-compact-budget.ts | 65 +++++++ src/server/auth-cors.ts | 11 ++ src/server/management/model-rows.ts | 4 + src/server/management/provider-routes.ts | 86 +++++++-- src/types/provider.ts | 5 + tests/auto-compact-budget.test.ts | 57 ++++++ tests/codex-catalog.test.ts | 173 +++++++++++++++++- ...odex-convergence-account-selectors.test.ts | 15 ++ tests/config.test.ts | 34 ++++ tests/management-provider-validation.test.ts | 89 ++++++++- tests/native-model-toggle.test.ts | 43 ++++- 27 files changed, 813 insertions(+), 80 deletions(-) create mode 100644 src/providers/auto-compact-budget.ts create mode 100644 tests/auto-compact-budget.test.ts diff --git a/docs-site/src/content/docs/fr/reference/configuration/providers.md b/docs-site/src/content/docs/fr/reference/configuration/providers.md index feedf5ad0ca..822c7ba47ce 100644 --- a/docs-site/src/content/docs/fr/reference/configuration/providers.md +++ b/docs-site/src/content/docs/fr/reference/configuration/providers.md @@ -84,6 +84,7 @@ sauvegarde dont le contenu diffère, puis réécrit en identifiants sans préfix | `modelContextWindows?` | `Record` | Valeurs de repli ou plafonds de contexte par modèle. Ils remplacent `contextWindow` : une fenêtre inconnue utilise la valeur configurée, tandis que des métadonnées actives plus faibles restent déterminantes. | | `modelInputModalities?` | `Record` | Conseils de saisie par modèle tels que `["text"]` ou `["text", "image"]`. | | `modelMaxInputTokens?` | `Record` | Limites d'entrée maximales positives par modèle utilisées pour les conseils de compactage automatique du catalogue. | +| `modelAutoCompactTokenLimits?` | `Record` | Budgets souples de compactage automatique par modèle, sous forme d'entiers sûrs positifs. Ils peuvent uniquement abaisser l'enveloppe effective de 90 % du contexte ou de l'entrée maximale et sont omis lorsqu'aucune fenêtre de contexte faisant autorité n'est connue. Pour le fournisseur canonique `openai`, les clés doivent être les identifiants exacts de modèles natifs pris en charge, sans préfixe de fournisseur ni de sélecteur de compte. PATCH fusionne les entrées ; `null` supprime une clé, tandis que `null` pour le champ entier efface la table. Ces marqueurs `null` sont réservés à PATCH. | | `defaultMaxOutputTokens?` | `number` | Solution de secours `openai-chat` à l’échelle du fournisseur lorsque le client omet `max_output_tokens`. | | `modelMaxOutputTokens?` | `Record` | Budgets de repli `openai-chat` positifs par modèle ; les correspondances exactes ou par motif priment sur la valeur par défaut du fournisseur. | | `modelCosts?` | `Record` | Prix affichés par modèle (USD par 1M de jetons), indexés par l'identifiant exact du modèle en amont de ce fournisseur — et non par un identifiant de fournisseur ni par une étiquette routée `provider/model`, par exemple `{ "deepseek-v4-flash": { "input": 0.14, "output": 0.28, "cacheRead": 0.0028, "cacheWrite": 0 } }`. Tout identifiant de modèle constitue une clé valide : les fournisseurs personnalisés peuvent cibler n'importe quel point de terminaison compatible avec OpenAI au moyen de l'adaptateur `openai-chat`, et les identifiants de fournisseur locaux ou internes fonctionnent même s'ils sont absents des catalogues intégrés. Les prix configurés par l'utilisateur priment sur les catalogues intégrés dans les estimations des pages Journaux (`~$`) et Utilisation. Les entrées historiques sont recalculées à partir de la surcharge actuelle ; modifier un prix peut donc changer les totaux antérieurs. L'ordre de repli est le suivant : `modelCosts` défini par l'utilisateur → catalogue jawcode → surcharge des prix attendus → repli propre au fournisseur au niveau du modèle. Une entrée entièrement nulle passe à la source suivante. Chaque tarif doit être un nombre fini positif ou nul, inférieur ou égal à 1 000 000 (USD par 1M de jetons) ; les lignes hors plage sont rejetées par l'interface de gestion et ignorées au chargement. Ces valeurs servent uniquement à l'estimation lors de l'affichage : les surcharges n'affectent jamais le routage, la sélection des comptes, les quotas ni la facturation. | diff --git a/docs-site/src/content/docs/ja/reference/configuration/providers.md b/docs-site/src/content/docs/ja/reference/configuration/providers.md index 27cf1406ecc..f1210892a57 100644 --- a/docs-site/src/content/docs/ja/reference/configuration/providers.md +++ b/docs-site/src/content/docs/ja/reference/configuration/providers.md @@ -72,6 +72,7 @@ account を削除しても mapping は保持され、同じ id を再追加す | `modelContextWindows?` | `Record` | モデルごとのコンテキスト値および上限。`contextWindow` より優先され、ウィンドウが不明なら設定値を使い、より小さいライブメタデータがあればそちらが優先されます。 | | `modelInputModalities?` | `Record` | `["text"]` や `["text", "image"]` などのモデルごとの入力ヒント。 | | `modelMaxInputTokens?` | `Record` |カタログの自動圧縮ヒントに使用されるモデルごとの正の最大入力制限。 | +| `modelAutoCompactTokenLimits?` | `Record` | モデルごとの正の安全な整数によるソフト自動圧縮予算。実効値であるコンテキストまたは最大入力の 90% の上限を下げることだけができ、信頼できるコンテキストウィンドウが不明な場合は出力されません。canonical `openai` では、キーは provider や account-selector の接頭辞を含まない、サポート対象の正確なネイティブモデル ID でなければなりません。provider PATCH はエントリをマージし、キーを `null` にするとそのキーを削除し、フィールド全体を `null` にするとマップを消去します。これらの `null` tombstone は PATCH 専用です。 | | `defaultMaxOutputTokens?` | `number` |クライアントが `max_output_tokens` を省略した場合の、プロバイダー全体の `openai-chat` フォールバック。 | | `modelMaxOutputTokens?` | `Record` |モデルごとの `openai-chat` フォールバック バジェットがプラスになります。正確な/パターン一致はプロバイダーのデフォルトを上回ります。 | | `modelCosts?` | `Record` | モデルごとの表示価格(100万トークンあたりの米ドル)。そのプロバイダーの正確なアップストリーム モデル ID をキーにします(プロバイダー識別子やルーティングされた `provider/model` ラベルではありません)。値は `input`, `output`, `cacheRead`, `cacheWrite` の 4 フィールドです(例: `{ "deepseek-v4-flash": { "input": 0.14, "output": 0.28, "cacheRead": 0.0028, "cacheWrite": 0 } }`)。組み込みカタログにないモデル ID も、任意の OpenAI 互換エンドポイントを対象とするカスタムプロバイダーや、ローカル・内部プロバイダーで有効です。ユーザー設定の価格は Logs の `~$` と Usage の見積もりで組み込みカタログより優先されます。過去のエントリも現在のオーバーレイで再計算されるため、価格を編集すると過去の合計が変わることがあります(フォールバック順: ユーザー設定 → jawcode カタログ → expected-price オーバーレイ → モデル別ベンダー価格)。全ゼロのエントリは次のソースにフォールバックします。各レートは 0 以上の有限数で、最大 1,000,000(100万トークンあたりの米ドル)です。範囲外の行は管理境界で拒否され、読み込み時に破棄されます。表示専用の見積もりであり、ルーティング・アカウント選択・クォータ・請求には影響しません。 | diff --git a/docs-site/src/content/docs/ko/reference/configuration/providers.md b/docs-site/src/content/docs/ko/reference/configuration/providers.md index 707129b2ed1..ccacb0a94f8 100644 --- a/docs-site/src/content/docs/ko/reference/configuration/providers.md +++ b/docs-site/src/content/docs/ko/reference/configuration/providers.md @@ -72,6 +72,7 @@ managed map을 활성화하면 privacy-safe selector를 만들고, 이후 계정 | `modelContextWindows?` | `Record` | 모델별 컨텍스트 값이자 상한입니다. `contextWindow`보다 우선하며, 창 크기를 알 수 없으면 설정값을 쓰고 더 작은 라이브 메타데이터가 있으면 그쪽을 따릅니다. | | `modelInputModalities?` | `Record` | `["text"]` 또는 `["text", "image"]` 같은 모델별 입력 힌트입니다. | | `modelMaxInputTokens?` | `Record` | 카탈로그 자동 압축 힌트에 쓰는 양수 모델별 최대 입력 한도입니다. | +| `modelAutoCompactTokenLimits?` | `Record` | 모델별 양의 안전 정수형 소프트 자동 압축 예산입니다. 유효한 컨텍스트 또는 최대 입력의 90% 한도를 낮출 수만 있으며, 신뢰할 수 있는 컨텍스트 창을 알 수 없으면 내보내지 않습니다. canonical `openai`에서는 키가 공급자나 계정 선택자 접두사가 없는 정확한 지원 네이티브 모델 ID여야 합니다. 공급자 PATCH는 항목을 병합하며, 키를 `null`로 지정하면 해당 키를 삭제하고 필드 전체를 `null`로 지정하면 맵을 지웁니다. 이 `null` tombstone은 PATCH에서만 사용할 수 있습니다. | | `defaultMaxOutputTokens?` | `number` | 클라이언트가 `max_output_tokens`를 생략했을 때 쓰는 공급자 전반의 `openai-chat` 폴백입니다. | | `modelMaxOutputTokens?` | `Record` | 양수 모델별 `openai-chat` 폴백 예산입니다. 정확한 일치와 패턴 일치가 공급자 기본값보다 우선합니다. | | `modelCosts?` | `Record` | 모델별 표시 가격(100만 토큰당 USD). 해당 공급자의 정확한 업스트림 모델 ID를 키로 사용하며(공급자 식별자나 라우팅된 `provider/model` 레이블이 아님) 값은 `input`, `output`, `cacheRead`, `cacheWrite` 네 필드입니다(예: `{ "deepseek-v4-flash": { "input": 0.14, "output": 0.28, "cacheRead": 0.0028, "cacheWrite": 0 } }`). 커스텀 공급자는 `openai-chat` 어댑터로 임의의 OpenAI 호환 엔드포인트를 대상으로 할 수 있으며, 내장 카탈로그에 없는 로컬·내부 공급자 ID도 유효합니다. 사용자 구성 가격은 Logs `~$` 및 Usage 추정에서 내장 카탈로그보다 우선합니다. 기존 항목도 현재 오버레이로 다시 계산되므로 가격을 편집하면 과거 합계가 바뀔 수 있습니다(폴백 순서: 사용자 설정 → jawcode 카탈로그 → expected-price 오버레이 → 모델별 벤더 가격). 전부 0인 항목은 다음 소스로 폴백합니다. 각 요율은 0 이상의 유한한 숫자이며 최대 1,000,000(100만 토큰당 USD)입니다. 범위를 벗어난 행은 관리 경계에서 거부되고 로드 시 삭제됩니다. 표시 전용 추정이며 라우팅·계정 선택·할당량·청구에는 영향을 주지 않습니다. | diff --git a/docs-site/src/content/docs/reference/configuration/providers.md b/docs-site/src/content/docs/reference/configuration/providers.md index c44b628714d..85cef14e1e4 100644 --- a/docs-site/src/content/docs/reference/configuration/providers.md +++ b/docs-site/src/content/docs/reference/configuration/providers.md @@ -85,6 +85,7 @@ differing backup and rewrites known legacy namespaced selected ids to bare ids. | `modelContextWindows?` | `Record` | Per-model context fallbacks/caps. These override `contextWindow`: an unknown window uses the configured value, while smaller live metadata remains authoritative. | | `modelInputModalities?` | `Record` | Per-model input hints such as `["text"]` or `["text", "image"]`. | | `modelMaxInputTokens?` | `Record` | Positive per-model max input limits used for catalog auto-compaction hints. | +| `modelAutoCompactTokenLimits?` | `Record` | Positive safe-integer per-model soft auto-compaction budgets. Values can only lower the effective 90%-of-context/max-input envelope and are omitted when no authoritative context window is known. For canonical `openai`, keys must be exact supported native model IDs without provider or account-selector prefixes. Provider PATCH merges entries; set a key to `null` to delete it or the whole field to `null` to clear the map. These `null` tombstones are PATCH-only. | | `defaultMaxOutputTokens?` | `number` | Provider-wide `openai-chat` fallback when the client omits `max_output_tokens`. | | `modelMaxOutputTokens?` | `Record` | Positive per-model `openai-chat` fallback budgets; exact/pattern matches beat the provider default. | | `modelCosts?` | `Record` | Per-model display prices (USD per 1M tokens), keyed by that provider's exact upstream model id — not a provider identifier or a routed `provider/model` label, e.g. `{ "deepseek-v4-flash": { "input": 0.14, "output": 0.28, "cacheRead": 0.0028, "cacheWrite": 0 } }`. Any model id is a valid key — custom providers may target any OpenAI-compatible endpoint through the `openai-chat` adapter, and local or internal provider ids work even when they are absent from the built-in catalogs. User-configured prices win over the built-in catalogs in the Logs `~$` and Usage estimates; historical entries are repriced from the current overlay, so editing a price can move past totals. The fallback order is user `modelCosts` → exact official correction → jawcode catalog → expected-price overlay → model-level vendor fallback, and an all-zero entry falls through to the next source in that sequence. Each rate must be a non-negative finite number at most 1,000,000 (USD per 1M tokens); out-of-range rows are rejected by the management boundary and dropped on load. Display-time estimation only: overlays never affect routing, account selection, quotas, or billing. | diff --git a/docs-site/src/content/docs/ru/reference/configuration/providers.md b/docs-site/src/content/docs/ru/reference/configuration/providers.md index 21b70bebb5a..c4155170741 100644 --- a/docs-site/src/content/docs/ru/reference/configuration/providers.md +++ b/docs-site/src/content/docs/ru/reference/configuration/providers.md @@ -85,6 +85,7 @@ cross-route credential fallback не существует. Строки API GPT- | `modelContextWindows?` | `Record` | Значения и cap'ы контекста по отдельным моделям. Перекрывают `contextWindow`: если окно неизвестно, берётся заданное значение, а более маленькая live-metadata остаётся авторитетной. | | `modelInputModalities?` | `Record` | Подсказки modality по модели, например `["text"]` или `["text", "image"]`. | | `modelMaxInputTokens?` | `Record` | Положительные лимиты max input по моделям, используемые для подсказок auto-compaction в каталоге. | +| `modelAutoCompactTokenLimits?` | `Record` | Мягкие бюджеты автосжатия по моделям в виде положительных безопасных целых чисел. Они могут только уменьшать эффективную границу в 90 % контекста или максимального ввода и не выдаются, если авторитетное окно контекста неизвестно. Для канонического `openai` ключами могут быть только точные поддерживаемые ID нативных моделей без префиксов провайдера или селектора аккаунта. PATCH провайдера объединяет записи: `null` для ключа удаляет его, а `null` для всего поля очищает карту. Такие маркеры `null` допустимы только в PATCH. | | `defaultMaxOutputTokens?` | `number` | Provider-wide fallback для `openai-chat`, когда клиент не передал `max_output_tokens`. | | `modelMaxOutputTokens?` | `Record` | Положительные fallback-budget'ы `openai-chat` по моделям; exact/pattern-match имеет приоритет над provider-default. | | `modelCosts?` | `Record` | Отображаемые цены по моделям (USD за 1M токенов), ключ — точный upstream id модели этого провайдера (не идентификатор провайдера и не маршрутизируемая метка `provider/model`), значение — четыре поля: `input`, `output`, `cacheRead`, `cacheWrite` (пример: `{ "deepseek-v4-flash": { "input": 0.14, "output": 0.28, "cacheRead": 0.0028, "cacheWrite": 0 } }`). Любой id допустим — кастомный провайдер может указывать на любой OpenAI-совместимый endpoint через адаптер `openai-chat`, а локальные и внутренние провайдеры работают даже без строки во встроенных каталогах. Пользовательские цены имеют приоритет над встроенными каталогами в оценках `~$` в Logs и Usage; исторические записи пересчитываются по текущему оверлею, поэтому изменение цены может сдвинуть прошлые суммы (порядок: пользователь → каталог jawcode → expected-price overlay → вендорская цена модели); полностью нулевая запись переходит к следующему источнику. Каждая ставка должна быть неотрицательным конечным числом не более 1 000 000 (USD за 1M токенов); строки вне диапазона отклоняются на управляющей границе и отбрасываются при загрузке. Только оценка для отображения: оверлеи не влияют на маршрутизацию, выбор аккаунта, квоты или биллинг. | diff --git a/docs-site/src/content/docs/tr/reference/configuration/providers.md b/docs-site/src/content/docs/tr/reference/configuration/providers.md index d4e414700a7..4db3211bc51 100644 --- a/docs-site/src/content/docs/tr/reference/configuration/providers.md +++ b/docs-site/src/content/docs/tr/reference/configuration/providers.md @@ -91,6 +91,7 @@ alanlı seçilmiş kimlikleri yalın kimliklere yeniden yazar. | `modelContextWindows?` | `Record` | Model başına bağlam geri dönüşleri/sınırları. Bunlar `contextWindow`'u geçersiz kılar: bilinmeyen bir pencere yapılandırılmış değeri kullanırken, daha küçük canlı meta veriler yetkili kalır. | | `modelInputModalities?` | `Record` | Model başına girdi ipuçları, örn. `["text"]` veya `["text", "image"]`. | | `modelMaxInputTokens?` | `Record` | Katalog otomatik sıkıştırma ipuçları için kullanılan pozitif model başına maksimum girdi sınırları. | +| `modelAutoCompactTokenLimits?` | `Record` | Model başına pozitif güvenli tamsayı biçiminde yumuşak otomatik sıkıştırma bütçeleri. Değerler yalnızca bağlamın veya maksimum girdinin etkin %90 zarfını düşürebilir ve yetkili bir bağlam penceresi bilinmiyorsa yayımlanmaz. Canonical `openai` için anahtarlar, sağlayıcı veya hesap seçici öneki olmadan desteklenen tam yerel model kimlikleri olmalıdır. Sağlayıcı PATCH girdileri birleştirir; bir anahtarı `null` yapmak o anahtarı siler, alanın tamamını `null` yapmak haritayı temizler. Bu `null` silme işaretleri yalnızca PATCH içindir. | | `defaultMaxOutputTokens?` | `number` | İstemci `max_output_tokens` değerini atladığında sağlayıcı genelinde `openai-chat` geri dönüşü. | | `modelMaxOutputTokens?` | `Record` | Pozitif model başına `openai-chat` geri dönüş bütçeleri; tam/kalıp eşleşmeleri sağlayıcı varsayılanını yener. | | `modelCosts?` | `Record` | Sağlayıcının tam yukarı akış model kimliğine göre anahtarlanan model başına görüntüleme fiyatları (1M token başına USD) — bir sağlayıcı tanımlayıcısı veya yönlendirilen `provider/model` etiketi değil, örn. `{ "deepseek-v4-flash": { "input": 0.14, "output": 0.28, "cacheRead": 0.0028, "cacheWrite": 0 } }`. Herhangi bir model kimliği geçerli bir anahtardır — özel sağlayıcılar `openai-chat` adaptörü aracılığıyla herhangi bir OpenAI uyumlu uç noktayı hedefleyebilir ve yerel veya dahili sağlayıcı kimlikleri yerleşik kataloglarda bulunmasalar bile çalışır. Kullanıcı tarafından yapılandırılan fiyatlar Günlükler `~$` ve Kullanım tahminlerinde yerleşik katalogları yener; geçmiş girdiler geçerli katmandan yeniden fiyatlandırılır, bu nedenle bir fiyatı düzenlemek geçmiş toplamları değiştirebilir. Geri dönüş sırası: kullanıcı `modelCosts` → jawcode kataloğu → beklenen fiyat katmanı → model düzeyinde satıcı geri dönüşü ve tamamen sıfır bir girdi bu dizideki bir sonraki kaynağa düşer. Her oran en fazla 1.000.000 (1M token başına USD) olan negatif olmayan sonlu bir sayı olmalıdır; aralık dışı satırlar yönetim sınırı tarafından reddedilir ve yükleme sırasında bırakılır. Yalnızca görüntüleme zamanı tahmini: katmanlar yönlendirmeyi, hesap seçimini, kotaları veya faturalandırmayı asla etkilemez. | diff --git a/docs-site/src/content/docs/zh-cn/reference/configuration/providers.md b/docs-site/src/content/docs/zh-cn/reference/configuration/providers.md index 1564842cbb7..3630a9ba6ce 100644 --- a/docs-site/src/content/docs/zh-cn/reference/configuration/providers.md +++ b/docs-site/src/content/docs/zh-cn/reference/configuration/providers.md @@ -72,6 +72,7 @@ selector,而不是分配一个新名称。 | `modelContextWindows?` | `Record` | 按模型设置的上下文数值与上限。优先于 `contextWindow`:窗口未知时采用所配置的数值,而更小的实时元数据仍然优先。 | | `modelInputModalities?` | `Record` | 按模型设置的输入提示,例如 `["text"]` 或 `["text", "image"]`。 | | `modelMaxInputTokens?` | `Record` | 正数型、按模型设置的最大输入限制,用于目录自动压缩提示。 | +| `modelAutoCompactTokenLimits?` | `Record` | 按模型设置的正安全整数软自动压缩预算。该值只能降低“上下文或最大输入的 90%”这一有效上限;没有已知的权威上下文窗口时不会输出。对于规范 `openai`,键必须是受支持的精确原生模型 ID,且不得包含提供者或账户选择器前缀。提供者 PATCH 会合并条目;将某个键设为 `null` 会删除该键,将整个字段设为 `null` 会清空映射。这些 `null` 删除标记仅适用于 PATCH。 | | `defaultMaxOutputTokens?` | `number` | 当客户端省略 `max_output_tokens` 时,`openai-chat` 的提供者级回退值。 | | `modelMaxOutputTokens?` | `Record` | 正数型、按模型设置的 `openai-chat` 回退预算;精确/模式匹配优先于提供者默认值。 | | `modelCosts?` | `Record` | 按模型设置的显示价格(每 100 万 token 的美元数),以该提供者的精确上游模型 ID 为键(不是提供者标识符或路由后的 `provider/model` 标签),值为四个字段:`input`、`output`、`cacheRead`、`cacheWrite`(示例:`{ "deepseek-v4-flash": { "input": 0.14, "output": 0.28, "cacheRead": 0.0028, "cacheWrite": 0 } }`)。任何模型 ID 都是有效键——自定义提供者可以通过 `openai-chat` 适配器指向任意 OpenAI 兼容端点,即使不存在于内置目录中,本地 OpenAI 兼容和内部提供者的 ID 同样有效。用户配置的价格在 Logs 的 `~$` 和 Usage 估算中优先于内置目录;历史条目也会按当前覆盖项重新计价,因此修改价格可能改变过去的总额(回退顺序:用户配置 → jawcode 目录 → expected-price 覆盖 → 模型级厂商价格);全零条目会回退到该顺序中的下一个来源。每个费率必须是大于等于 0 的有限数字,且不超过 1,000,000(每 100 万 token 的美元数);超出范围的条目会在管理边界被拒绝,并在加载时被丢弃。仅用于显示的估算:覆盖项不影响路由、账户选择、配额或计费。 | diff --git a/docs-site/src/content/docs/zh-tw/reference/configuration/providers.md b/docs-site/src/content/docs/zh-tw/reference/configuration/providers.md index f47b5bef059..b0a46f49ecc 100644 --- a/docs-site/src/content/docs/zh-tw/reference/configuration/providers.md +++ b/docs-site/src/content/docs/zh-tw/reference/configuration/providers.md @@ -54,6 +54,7 @@ description: 供應商項目、認證、端點、模型目錄、配額、context | `modelContextWindows?` | `Record` | Per-model context 上限。這些覆寫 `contextWindow` 且永不提高較小的即時中繼資料。 | | `modelInputModalities?` | `Record` | Per-model 輸入提示,如 `["text"]` 或 `["text", "image"]`。 | | `modelMaxInputTokens?` | `Record` | 用於目錄自動壓縮提示的正數 per-model max input 限制。 | +| `modelAutoCompactTokenLimits?` | `Record` | Per-model 正安全整數型 soft 自動壓縮預算。此值只能降低「context 或 max input 的 90%」這個有效上限;沒有已知的權威 context window 時不會輸出。對 canonical `openai` 而言,key 必須是受支援的精確 native model ID,且不得含 provider 或 account-selector 前綴。Provider PATCH 會合併項目;將單一 key 設為 `null` 會刪除該 key,將整個欄位設為 `null` 會清空 map。這些 `null` tombstone 僅供 PATCH 使用。 | | `defaultMaxOutputTokens?` | `number` | 當客戶端省略 `max_output_tokens` 時的供應商範圍 `openai-chat` 後備。 | | `modelMaxOutputTokens?` | `Record` | 正數 per-model `openai-chat` 後援預算;精確/模式比對勝過供應商預設。 | | `headers?` | `Record` | 額外上游標頭。Authorization、cookie、API-key 標頭、內嵌換行與無效名稱被拒絕。 | diff --git a/src/codex/catalog/aggregation.ts b/src/codex/catalog/aggregation.ts index a6055342278..73c17459544 100644 --- a/src/codex/catalog/aggregation.ts +++ b/src/codex/catalog/aggregation.ts @@ -13,6 +13,7 @@ import { getModelMetadata, getModelMetadataCaseInsensitive, listModelMetadata, r import { enrichProviderFromRegistry, shouldCaseFoldMetadataModelId } from "../../providers/derive"; import { getProviderRegistryEntry } from "../../providers/registry"; import { applyProviderContextCap, providerContextCap } from "../../providers/context-cap"; +import { clampAutoCompactTokenLimit } from "../../providers/auto-compact-budget"; import { routedSlug, slugEquals, slugEquivalenceKey, slugsEquivalent } from "../../providers/slug-codec"; import { CODEX_GPT5_IDENTITY_LINE } from "../../adapters/identity"; import { filterCursorConfiguredModelsByLiveDiscovery } from "../../adapters/cursor/discovery"; @@ -154,8 +155,16 @@ export function deriveComboCatalogModel( // combo would have the same window even without the cap. const contextCapped = limitingMembers.every(member => member.contextCapped === true); const maxInputTokens = Math.min( + contextWindow, ...members.map(member => member.maxInputTokens ?? member.contextWindow!), ); + const autoCompactTokenLimit = Math.min( + ...members.map(member => clampAutoCompactTokenLimit( + member.contextWindow!, + member.maxInputTokens, + member.autoCompactTokenLimit, + )), + ); const defaultReasoningEffort = effectiveComboDefault( combo.defaultEffort, reasoningEfforts, @@ -167,6 +176,7 @@ export function deriveComboCatalogModel( owned_by: COMBO_NAMESPACE, contextWindow, maxInputTokens, + autoCompactTokenLimit, ...(hasLimitingContextCapMetadata ? { contextCapped } : {}), inputModalities, reasoningEfforts, @@ -210,6 +220,7 @@ export function comboCatalogWarningSignature( key, contextWindow: member?.contextWindow ?? null, maxInputTokens: member?.maxInputTokens ?? null, + autoCompactTokenLimit: member?.autoCompactTokenLimit ?? null, inputModalities: [...new Set(member?.inputModalities ?? [])].sort(), reasoningEfforts: [...new Set(member?.reasoningEfforts ?? [])].sort(), parallelToolCalls: member?.parallelToolCalls === true, @@ -299,6 +310,7 @@ export function normalizedOpenAiApiSignature(model: CatalogModel): string { id: model.id, contextWindow: model.contextWindow ?? null, maxInputTokens: model.maxInputTokens ?? null, + autoCompactTokenLimit: model.autoCompactTokenLimit ?? null, inputModalities: [...new Set(model.inputModalities ?? [])].sort(), reasoningEfforts: [...new Set(model.reasoningEfforts ?? [])].sort(), ownedBy: model.owned_by ?? null, diff --git a/src/codex/catalog/effort.ts b/src/codex/catalog/effort.ts index 0648b64d17e..2d6494e2fd3 100644 --- a/src/codex/catalog/effort.ts +++ b/src/codex/catalog/effort.ts @@ -13,6 +13,7 @@ import { getModelMetadata, getModelMetadataCaseInsensitive, listModelMetadata, r import { enrichProviderFromRegistry, shouldCaseFoldMetadataModelId } from "../../providers/derive"; import { getProviderRegistryEntry } from "../../providers/registry"; import { applyProviderContextCap, providerContextCap, resolveUnknownRoutedContextWindow } from "../../providers/context-cap"; +import { clampAutoCompactTokenLimit } from "../../providers/auto-compact-budget"; import { routedSlug, slugEquals, slugsEquivalent } from "../../providers/slug-codec"; import { CODEX_GPT5_IDENTITY_LINE } from "../../adapters/identity"; import { filterCursorConfiguredModelsByLiveDiscovery } from "../../adapters/cursor/discovery"; @@ -128,9 +129,23 @@ export function applyCatalogModelMetadata(entry: RawEntry, model?: CatalogModel) if (typeof resolvedContext === "number" && resolvedContext > 0) { entry.context_window = resolvedContext; entry.max_context_window = resolvedContext; - entry.auto_compact_token_limit = Math.min( - Math.floor(resolvedContext * 0.9), - model.maxInputTokens ?? Number.POSITIVE_INFINITY, + entry.auto_compact_token_limit = clampAutoCompactTokenLimit( + resolvedContext, + model.maxInputTokens, + model.autoCompactTokenLimit, + ); + } else if ( + typeof entry.context_window === "number" + && entry.context_window > 0 + && typeof model.maxInputTokens === "number" + && model.maxInputTokens > 0 + ) { + // A conservative routed fallback is not evidence for applying the optional soft policy, + // but a measured/configured input ceiling is still a hard bound. Compact before that + // ceiling even when the provider supplied no authoritative context window. + entry.auto_compact_token_limit = clampAutoCompactTokenLimit( + entry.context_window, + model.maxInputTokens, ); } if (Array.isArray(model.inputModalities) && model.inputModalities.length > 0) { diff --git a/src/codex/catalog/metadata.ts b/src/codex/catalog/metadata.ts index d071bf59bd5..5494d12ddd3 100644 --- a/src/codex/catalog/metadata.ts +++ b/src/codex/catalog/metadata.ts @@ -14,6 +14,7 @@ import { getModelMetadata, getModelMetadataCaseInsensitive, listModelMetadata, r import { enrichProviderFromRegistry, shouldCaseFoldMetadataModelId } from "../../providers/derive"; import { getProviderRegistryEntry, providerCodexAccountMode } from "../../providers/registry"; import { applyProviderContextCap, providerContextCap } from "../../providers/context-cap"; +import { clampAutoCompactTokenLimit } from "../../providers/auto-compact-budget"; import { routedSlug, slugEquals, slugsEquivalent } from "../../providers/slug-codec"; import { identifyRoutedModel } from "../../adapters/identity"; import { filterCursorConfiguredModelsByLiveDiscovery } from "../../adapters/cursor/discovery"; @@ -204,6 +205,8 @@ export interface NativeContextLimits { readonly providerWindow?: number; /** `providers.openai.modelContextWindows` — per-model, wins over `providerWindow`. */ readonly modelWindows?: Readonly>; + /** `providers.openai.modelAutoCompactTokenLimits` — soft, lowering-only budgets. */ + readonly modelAutoCompactTokenLimits?: Readonly>; } export type NativeContextLimitsInput = NativeContextLimits | number | undefined; @@ -227,12 +230,18 @@ export function nativeContextLimits( const window = positiveInt(value); if (window !== undefined) modelWindows[slug] = window; } + const modelAutoCompactTokenLimits: Record = {}; + for (const [slug, value] of Object.entries(provider?.modelAutoCompactTokenLimits ?? {})) { + const budget = positiveInt(value); + if (budget !== undefined) modelAutoCompactTokenLimits[slug] = budget; + } return { ...(positiveInt(providerContextCap(config, OPENAI_CODEX_PROVIDER_ID)) !== undefined ? { cap: providerContextCap(config, OPENAI_CODEX_PROVIDER_ID) } : {}), ...(positiveInt(provider?.contextWindow) !== undefined ? { providerWindow: provider!.contextWindow } : {}), ...(Object.keys(modelWindows).length > 0 ? { modelWindows } : {}), + ...(Object.keys(modelAutoCompactTokenLimits).length > 0 ? { modelAutoCompactTokenLimits } : {}), }; } @@ -277,6 +286,21 @@ export function nativeOpenAiMaxInputTokens(slug: string, limits?: NativeContextL return window === undefined ? narrowed : Math.min(narrowed, window); } +/** Effective native soft budget after every hard window/input limit is resolved. */ +export function nativeOpenAiAutoCompactTokenLimit( + slug: string, + limits?: NativeContextLimitsInput, +): number | undefined { + const contextWindow = nativeOpenAiContextWindow(slug, limits); + if (contextWindow === undefined) return undefined; + const configured = positiveInt(asLimits(limits).modelAutoCompactTokenLimits?.[slug]); + return clampAutoCompactTokenLimit( + contextWindow, + nativeOpenAiMaxInputTokens(slug, limits), + configured, + ); +} + export function nativeInputModalities(slug: string): string[] { const upstream = PINNED_NATIVE_CAPABILITY_ENTRIES.get(slug); if (Array.isArray(upstream?.input_modalities) && upstream!.input_modalities!.length > 0) { @@ -387,7 +411,7 @@ export function desktopVisibleNativeSlugs( ]); } -export function nativeModelRows(config: Pick): Array<{ slug: string; disabled: boolean; contextWindow?: number; maxInputTokens?: number }> { +export function nativeModelRows(config: Pick): Array<{ slug: string; disabled: boolean; contextWindow?: number; maxInputTokens?: number; autoCompactTokenLimit?: number }> { const disabled = disabledNativeSlugs(config); const shadowed = configuredNativeAliasSlugs(config); // Both user levers, not just the cap: a per-model window set from the dashboard has to show @@ -403,11 +427,13 @@ export function nativeModelRows(config: Pick !shadowed.has(slug)).map(slug => { const contextWindow = nativeOpenAiContextWindow(slug, limits); const maxInputTokens = nativeOpenAiMaxInputTokens(slug, limits); + const autoCompactTokenLimit = nativeOpenAiAutoCompactTokenLimit(slug, limits); return { slug, disabled: disabled.has(slug), ...(contextWindow !== undefined ? { contextWindow } : {}), ...(maxInputTokens !== undefined ? { maxInputTokens } : {}), + ...(autoCompactTokenLimit !== undefined ? { autoCompactTokenLimit } : {}), }; }); } diff --git a/src/codex/catalog/parsing.ts b/src/codex/catalog/parsing.ts index a47b2c6894a..7fe07aca7f8 100644 --- a/src/codex/catalog/parsing.ts +++ b/src/codex/catalog/parsing.ts @@ -31,7 +31,8 @@ import { redactSecretString } from "../../lib/redact"; import upstreamModelsSnapshot from "../data/upstream-models.json"; -import { NATIVE_OPENAI_CONTEXT_OVERRIDES, SUPPORTED_NATIVE_OPENAI_SLUGS, UPSTREAM_NATIVE_ENTRIES, isNativeOpenAiCapabilityAliasModel, nativeMultiAgentVersion, nativeOpenAiContextWindow, nativeOpenAiMaxInputTokens, type NativeContextLimitsInput } from "./metadata"; +import { NATIVE_OPENAI_CONTEXT_OVERRIDES, SUPPORTED_NATIVE_OPENAI_SLUGS, UPSTREAM_NATIVE_ENTRIES, isNativeOpenAiCapabilityAliasModel, nativeMultiAgentVersion, nativeOpenAiAutoCompactTokenLimit, nativeOpenAiContextWindow, nativeOpenAiMaxInputTokens, type NativeContextLimitsInput } from "./metadata"; +import { clampAutoCompactTokenLimit } from "../../providers/auto-compact-budget"; import { trustedAccountBoundNativeCatalogSlug } from "./account-models"; import { CODEX_NATIVE_ALIAS_CATALOG_KIND } from "./kinds"; @@ -111,6 +112,8 @@ export interface CatalogModel { defaultReasoningEffort?: string; contextWindow?: number; maxInputTokens?: number; + /** Soft client compaction threshold; hard context/input limits remain authoritative. */ + autoCompactTokenLimit?: number; contextCap?: number; contextCapped?: boolean; inputModalities?: string[]; @@ -292,22 +295,6 @@ export function isNativeOpenAiEntry(entry: RawEntry): boolean { return typeof entry.slug === "string" && !entry.slug.includes("/"); } -/** - * Auto-compaction threshold for a native row. - * - * The usual rule is 90% of the window, but a row whose input ceiling sits below that has to - * clamp to the ceiling instead — otherwise the client keeps filling until upstream answers - * `context_length_exceeded` and compaction never gets a chance to run. Native GPT-5.6 no - * longer trips this (922,000 window, 829,800 at 90%), but the routed and API-key rows carry - * the same family at a 1,050,000 window where 90% would be 945,000 — past the ceiling. - */ -function nativeAutoCompactLimit(contextWindow: number, maxInputTokens: number | undefined, contextCap?: number): number { - const ninety = Math.floor(contextWindow * 0.9); - if (typeof maxInputTokens !== "number" || maxInputTokens <= 0) return ninety; - const cappedMaxInput = applyProviderContextCap(maxInputTokens, contextCap) ?? maxInputTokens; - return Math.min(ninety, cappedMaxInput, contextWindow); -} - /** * Narrow any already-resolved native window by the user levers. * @@ -340,11 +327,6 @@ export function applyNativeOpenAiContextOverride(entry: RawEntry, limits?: Nativ if (typeof override.contextWindow === "number") { const contextWindow = nativeOpenAiContextWindow(nativeSlug, limits) ?? override.contextWindow; entry.context_window = contextWindow; - entry.auto_compact_token_limit = nativeAutoCompactLimit( - contextWindow, - nativeOpenAiMaxInputTokens(nativeSlug, limits) ?? override.maxInputTokens, - undefined, - ); } if (typeof override.maxContextWindow === "number") { const maxContextWindow = narrowNativeMaxContextWindow(nativeSlug, override.maxContextWindow, limits); @@ -359,17 +341,36 @@ export function applyNativeOpenAiContextOverride(entry: RawEntry, limits?: Nativ const cappedContext = narrowNativeMaxContextWindow(nativeSlug, currentContext, limits); if (cappedContext !== currentContext && typeof cappedContext === "number") { entry.context_window = cappedContext; - entry.auto_compact_token_limit = nativeAutoCompactLimit( - cappedContext, - nativeOpenAiMaxInputTokens(nativeSlug, limits) ?? override?.maxInputTokens, - undefined, - ); } const currentMax = typeof entry.max_context_window === "number" ? entry.max_context_window : undefined; const cappedMax = narrowNativeMaxContextWindow(nativeSlug, currentMax, limits); if (cappedMax !== currentMax) { entry.max_context_window = cappedMax; } + const effectiveContext = typeof entry.context_window === "number" && entry.context_window > 0 + ? entry.context_window + : undefined; + if (effectiveContext !== undefined) { + const derivedAutoCompactTokenLimit = nativeOpenAiAutoCompactTokenLimit(nativeSlug, limits); + const retainedAutoCompactTokenLimit = isNativeOpenAiEntry(entry) + && typeof entry.auto_compact_token_limit === "number" + && Number.isSafeInteger(entry.auto_compact_token_limit) + && entry.auto_compact_token_limit > 0 + ? entry.auto_compact_token_limit + : undefined; + // A smaller threshold retained from Codex is policy evidence too. Configuration may + // lower it further, but catalog sync must never replace it with a larger default. + const loweringAutoCompactTokenLimit = retainedAutoCompactTokenLimit === undefined + ? derivedAutoCompactTokenLimit + : derivedAutoCompactTokenLimit === undefined + ? retainedAutoCompactTokenLimit + : Math.min(retainedAutoCompactTokenLimit, derivedAutoCompactTokenLimit); + entry.auto_compact_token_limit = clampAutoCompactTokenLimit( + effectiveContext, + nativeOpenAiMaxInputTokens(nativeSlug, limits) ?? override?.maxInputTokens, + loweringAutoCompactTokenLimit, + ); + } } export function ensureStrictCatalogFields( diff --git a/src/codex/catalog/provider-fetch.ts b/src/codex/catalog/provider-fetch.ts index e0e70009fe3..798d3379206 100644 --- a/src/codex/catalog/provider-fetch.ts +++ b/src/codex/catalog/provider-fetch.ts @@ -41,6 +41,7 @@ import type { FastPolicyAuthority } from "../../providers/fastwire"; import { effectiveGoogleMode, getProviderRegistryEntry, providerMatchesRegistryTransport } from "../../providers/registry"; import { parseAntigravityAvailableModels, registerAntigravityDiscoveredWireModels } from "../../providers/antigravity-models"; import { applyProviderContextCap, providerContextCap, resolveUnknownRoutedContextWindow } from "../../providers/context-cap"; +import { clampAutoCompactTokenLimit } from "../../providers/auto-compact-budget"; import { routedSlug, slugEquals, slugEquivalenceKey, slugsEquivalent } from "../../providers/slug-codec"; import { CODEX_GPT5_IDENTITY_LINE } from "../../adapters/identity"; import { filterCursorConfiguredModelsByLiveDiscovery } from "../../adapters/cursor/discovery"; @@ -75,7 +76,7 @@ import { createAdmissionGate, ResourceAdmissionError, type AdmissionMetrics } fr import { CODEX_CUSTOM_MODEL_CATALOG_KIND, JAWCODE_CATALOG_AUGMENT_PROVIDERS, catalogModelSlug, shouldExposeRoutedModel } from "./parsing"; import type { CatalogModel } from "./parsing"; -import { disabledNativeSlugs, hasComboTargets, isNativeOpenAiCapabilityAliasModel, NATIVE_GPT56_MAX_INPUT_TOKENS, nativeContextLimits, nativeDefaultReasoningEffort, nativeInputModalities, nativeOpenAiContextWindow, nativeOpenAiMaxInputTokens, nativeOpenAiSlugs, nativeParallelToolCalls, nativeReasoningEfforts } from "./metadata"; +import { disabledNativeSlugs, hasComboTargets, isNativeOpenAiCapabilityAliasModel, NATIVE_GPT56_MAX_INPUT_TOKENS, nativeContextLimits, nativeDefaultReasoningEffort, nativeInputModalities, nativeOpenAiAutoCompactTokenLimit, nativeOpenAiContextWindow, nativeOpenAiMaxInputTokens, nativeOpenAiSlugs, nativeParallelToolCalls, nativeReasoningEfforts } from "./metadata"; import { deriveComboCatalogModel, normalizedOpenAiApiSignature, openAiApiCollisionWarnings, replaceLastComboCatalogOmissions, warnUncataloguedComboOnce } from "./aggregation"; import type { ComboCatalogOmission } from "./aggregation"; import type { CatalogGatherProviderAuthEvidence } from "./filesystem-evidence"; @@ -571,6 +572,7 @@ function providerCatalogFingerprint(name: string, prov: OcxProviderConfig): Reco ctx: prov.contextWindow ?? null, ctxW: prov.modelContextWindows ?? null, maxIn: prov.modelMaxInputTokens ?? null, + autoCompact: prov.modelAutoCompactTokenLimits ?? null, inMod: prov.modelInputModalities ?? null, re: prov.modelReasoningEfforts ?? null, defRe: prov.modelDefaultReasoningEfforts ?? null, @@ -625,6 +627,17 @@ export function configuredMaxInputTokens(prov: OcxProviderConfig, id: string): n return typeof configured === "number" && configured > 0 ? configured : undefined; } +export function configuredAutoCompactTokenLimit( + prov: OcxProviderConfig | undefined, + id: string, +): number | undefined { + if (!prov) return undefined; + const configured = modelRecordValue(prov.modelAutoCompactTokenLimits, id); + return typeof configured === "number" && Number.isSafeInteger(configured) && configured > 0 + ? configured + : undefined; +} + function configuredReasoningSummarySupport(prov: OcxProviderConfig | undefined, id: string): boolean | undefined { if (!prov) return undefined; const explicit = modelRecordValue(prov.modelSupportsReasoningSummaries, id); @@ -636,6 +649,7 @@ export function applyProviderConfigHints(name: string, prov: OcxProviderConfig, void name; const configuredCap = configuredContextWindow(prov, model.id); const configuredMaxInput = configuredMaxInputTokens(prov, model.id); + const configuredAutoCompact = configuredAutoCompactTokenLimit(prov, model.id); let inputModalities = configuredInputModalities(prov, model.id); // Vision-sidecar coverage: `noVisionModels` marks models whose images the PROXY describes // (src/vision/index.ts). The catalog must still advertise image input for them — the Codex app @@ -689,10 +703,31 @@ export function applyProviderConfigHints(name: string, prov: OcxProviderConfig, ...(prov.codexToolMode !== undefined ? { codexToolMode: prov.codexToolMode } : {}), }; const capped = applyProviderContextCap(hinted.contextWindow, providerCap); - if (providerCap !== undefined && capped !== hinted.contextWindow) { - return { ...hinted, contextWindow: capped, contextCap: providerCap, contextCapped: true }; - } - return providerCap !== undefined ? { ...hinted, contextCap: providerCap, contextCapped: false } : hinted; + const withCap = providerCap !== undefined + ? capped !== hinted.contextWindow + ? { ...hinted, contextWindow: capped, contextCap: providerCap, contextCapped: true } + : { ...hinted, contextCap: providerCap, contextCapped: false } + : hinted; + const contextWindow = typeof withCap.contextWindow === "number" && withCap.contextWindow > 0 + ? withCap.contextWindow + : undefined; + const boundedMaxInput = typeof withCap.maxInputTokens === "number" && withCap.maxInputTokens > 0 + ? (contextWindow !== undefined ? Math.min(withCap.maxInputTokens, contextWindow) : withCap.maxInputTokens) + : undefined; + const withHardBounds = boundedMaxInput !== undefined && boundedMaxInput !== withCap.maxInputTokens + ? { ...withCap, maxInputTokens: boundedMaxInput } + : withCap; + const softCandidates = [model.autoCompactTokenLimit, configuredAutoCompact] + .filter((value): value is number => typeof value === "number" && value > 0); + if (contextWindow === undefined || softCandidates.length === 0) return withHardBounds; + return { + ...withHardBounds, + autoCompactTokenLimit: clampAutoCompactTokenLimit( + contextWindow, + boundedMaxInput, + Math.min(...softCandidates), + ), + }; } export function catalogHintsFromProviderConfig(name: string, prov: OcxProviderConfig, id: string, contextCap?: number): Partial { @@ -719,6 +754,7 @@ interface ComboCatalogMemberFallback { readonly contextWindow?: number; /** Input ceiling when it is lower than the window (native GPT-5.6: 922k under 1.05M). */ readonly maxInputTokens?: number; + readonly autoCompactTokenLimit?: number; readonly inputModalities?: readonly string[]; readonly reasoningEfforts?: readonly string[]; } @@ -747,26 +783,33 @@ export function resolveComboCatalogMember( if (prov?.disabled === true) return undefined; const withFallbackMetadata = (member: CatalogModel): CatalogModel => { - if (!fallback) return member; const contextWindow = typeof member.contextWindow === "number" && member.contextWindow > 0 ? member.contextWindow : undefined; - const addMaxInput = contextWindow !== undefined + const addMaxInput = fallback !== undefined && contextWindow !== undefined && !(typeof member.maxInputTokens === "number" && member.maxInputTokens > 0); + const effectiveMaxInput = addMaxInput + ? Math.min(fallback?.maxInputTokens ?? contextWindow!, contextWindow!) + : member.maxInputTokens; + const softCandidates = [member.autoCompactTokenLimit, fallback?.autoCompactTokenLimit] + .filter((value): value is number => typeof value === "number" && value > 0); + const autoCompactTokenLimit = contextWindow !== undefined && softCandidates.length > 0 + ? clampAutoCompactTokenLimit(contextWindow, effectiveMaxInput, Math.min(...softCandidates)) + : member.autoCompactTokenLimit; + const adjustAutoCompact = autoCompactTokenLimit !== member.autoCompactTokenLimit; const addModalities = (!Array.isArray(member.inputModalities) || member.inputModalities.length === 0) - && fallback.inputModalities !== undefined; + && fallback?.inputModalities !== undefined; const addReasoning = member.reasoningEfforts === undefined - && fallback.reasoningEfforts !== undefined; - if (!addMaxInput && !addModalities && !addReasoning) return member; + && fallback?.reasoningEfforts !== undefined; + if (!addMaxInput && !adjustAutoCompact && !addModalities && !addReasoning) return member; return { ...member, // Never claim a larger input budget than the window, and prefer the model's own // measured ceiling when the fallback carries one. - ...(addMaxInput - ? { maxInputTokens: Math.min(fallback.maxInputTokens ?? contextWindow!, contextWindow!) } - : {}), - ...(addModalities ? { inputModalities: [...fallback.inputModalities!] } : {}), - ...(addReasoning ? { reasoningEfforts: [...fallback.reasoningEfforts!] } : {}), + ...(addMaxInput ? { maxInputTokens: effectiveMaxInput } : {}), + ...(adjustAutoCompact && autoCompactTokenLimit !== undefined ? { autoCompactTokenLimit } : {}), + ...(addModalities ? { inputModalities: [...fallback!.inputModalities!] } : {}), + ...(addReasoning ? { reasoningEfforts: [...fallback!.reasoningEfforts!] } : {}), }; }; @@ -847,6 +890,20 @@ export function resolveComboCatalogMember( const maxInputTokens = effectiveMaxInput !== undefined ? Math.min(effectiveMaxInput, contextWindow) : contextWindow; + const softCandidates = [ + hinted.autoCompactTokenLimit, + base.autoCompactTokenLimit, + fallback?.autoCompactTokenLimit, + configuredAutoCompactTokenLimit(prov, target.model), + ].filter((value): value is number => typeof value === "number" && value > 0); + // A generic 128k synthesis is a catalog compatibility fallback, not evidence + // that a configured soft policy has an authoritative window to clamp against. + const hasAuthoritativeAutoCompactBasis = hintedContext !== undefined + || fallbackContext !== undefined + || contextCap !== undefined; + const autoCompactTokenLimit = hasAuthoritativeAutoCompactBasis && softCandidates.length > 0 + ? clampAutoCompactTokenLimit(contextWindow, maxInputTokens, Math.min(...softCandidates)) + : undefined; return { ...hinted, @@ -854,6 +911,7 @@ export function resolveComboCatalogMember( ...(reasoningEfforts !== undefined ? { reasoningEfforts } : {}), contextWindow, maxInputTokens, + ...(autoCompactTokenLimit !== undefined ? { autoCompactTokenLimit } : {}), ...(fallbackCapped ? { contextCap, contextCapped: true as const } : {}), }; } @@ -1774,6 +1832,7 @@ async function gatherRoutedModelsUncached( // stay separate fields because routed/API rows of the same family run a wider window. // Falls back to the window for slugs with no separate ceiling. maxInputTokens: Math.min(nativeOpenAiMaxInputTokens(slug, openaiContextCap) ?? contextWindow, contextWindow), + autoCompactTokenLimit: nativeOpenAiAutoCompactTokenLimit(slug, openaiContextCap), inputModalities: nativeInputModalities(slug), reasoningEfforts: nativeReasoningEfforts(slug), ...(nativeParallelToolCalls(slug) ? { parallelToolCalls: true } : {}), @@ -1790,18 +1849,23 @@ async function gatherRoutedModelsUncached( for (const id of listComboIds(config)) { const combo = getCombo(config, id); if (!combo) continue; + const comboNativeLimits = nativeContextLimits(config); const nativeContextWindow = combo.nativeAlias && combo.alias - ? nativeOpenAiContextWindow(combo.alias, nativeContextLimits(config)) + ? nativeOpenAiContextWindow(combo.alias, comboNativeLimits) : undefined; const nativeAliasMaxInput = combo.nativeAlias && combo.alias ? (combo.alias.startsWith("gpt-5.6-") || combo.alias.includes("daybreak") ? NATIVE_GPT56_MAX_INPUT_TOKENS : nativeOpenAiMaxInputTokens(combo.alias) ?? nativeOpenAiContextWindow(combo.alias)) : undefined; + const nativeAliasAutoCompact = combo.nativeAlias && combo.alias + ? nativeOpenAiAutoCompactTokenLimit(combo.alias, comboNativeLimits) + : undefined; const nativeAliasFallback = combo.nativeAlias && combo.alias && nativeContextWindow !== undefined ? { contextWindow: nativeContextWindow, ...(nativeAliasMaxInput !== undefined ? { maxInputTokens: nativeAliasMaxInput } : {}), + ...(nativeAliasAutoCompact !== undefined ? { autoCompactTokenLimit: nativeAliasAutoCompact } : {}), inputModalities: nativeInputModalities(combo.alias), reasoningEfforts: nativeReasoningEfforts(combo.alias), } @@ -1864,9 +1928,23 @@ async function gatherRoutedModelsUncached( const nativeAliasMaxInputTokens = codexForwardNativeCapabilityAlias ? nativeOpenAiMaxInputTokens(cm.modelId, customNativeLimits) : undefined; - const customMaxInputTokens = nativeAliasMaxInputTokens !== undefined && customContextWindow !== undefined - ? Math.min(nativeAliasMaxInputTokens, customContextWindow) - : nativeAliasMaxInputTokens; + const configuredMaxInput = rawProvider + ? configuredMaxInputTokens(rawProvider, cm.modelId) + : undefined; + const hardMaxCandidates = [nativeAliasMaxInputTokens, configuredMaxInput] + .filter((value): value is number => typeof value === "number" && value > 0); + const customMaxInputTokens = hardMaxCandidates.length > 0 + ? Math.min( + ...hardMaxCandidates, + ...(customContextWindow !== undefined ? [customContextWindow] : []), + ) + : undefined; + const configuredAutoCompact = configuredAutoCompactTokenLimit(rawProvider, cm.modelId); + const customAutoCompactTokenLimit = codexForwardNativeCapabilityAlias + ? nativeOpenAiAutoCompactTokenLimit(cm.modelId, customNativeLimits) + : customContextWindow !== undefined && configuredAutoCompact !== undefined + ? clampAutoCompactTokenLimit(customContextWindow, customMaxInputTokens, configuredAutoCompact) + : undefined; const nativeAliasDefaultEffort = codexForwardNativeCapabilityAlias ? nativeDefaultReasoningEffort(cm.modelId) : undefined; @@ -1887,6 +1965,7 @@ async function gatherRoutedModelsUncached( : codexForwardNativeCapabilityAlias ? { displayName: "Daybreak Blue" } : {}), ...(customContextWindow !== undefined ? { contextWindow: customContextWindow } : {}), ...(customMaxInputTokens !== undefined ? { maxInputTokens: customMaxInputTokens } : {}), + ...(customAutoCompactTokenLimit !== undefined ? { autoCompactTokenLimit: customAutoCompactTokenLimit } : {}), ...(cm.inputModalities ? { inputModalities: cm.inputModalities } : codexForwardNativeCapabilityAlias ? { inputModalities: nativeInputModalities(cm.modelId) } : {}), @@ -1932,10 +2011,18 @@ async function gatherRoutedModelsUncached( // along when it is actually a member — otherwise a provider default like "xhigh" would // re-apply onto a narrower custom ladder and override the fallback in applyReasoningLevels. const effectiveLadder = base.reasoningEfforts ?? replaced?.reasoningEfforts; + const mergedMaxInputCandidates = [base.maxInputTokens, replaced?.maxInputTokens] + .filter((value): value is number => typeof value === "number" && value > 0); + const mergedMaxInput = mergedMaxInputCandidates.length > 0 + ? Math.min(...mergedMaxInputCandidates) + : undefined; const merged: CatalogModel = replaced ? { ...base, ...(base.contextWindow === undefined && replaced.contextWindow !== undefined ? { contextWindow: replaced.contextWindow } : {}), - ...(base.maxInputTokens === undefined && replaced.maxInputTokens !== undefined ? { maxInputTokens: replaced.maxInputTokens } : {}), + ...(mergedMaxInput !== undefined ? { maxInputTokens: mergedMaxInput } : {}), + ...(base.autoCompactTokenLimit === undefined && replaced.autoCompactTokenLimit !== undefined + ? { autoCompactTokenLimit: replaced.autoCompactTokenLimit } + : {}), ...(base.inputModalities === undefined && replaced.inputModalities !== undefined ? { inputModalities: replaced.inputModalities } : {}), ...(base.reasoningEfforts === undefined && replaced.reasoningEfforts !== undefined ? { reasoningEfforts: replaced.reasoningEfforts } : {}), ...(base.defaultReasoningEffort === undefined && replaced.defaultReasoningEffort !== undefined @@ -1952,14 +2039,36 @@ async function gatherRoutedModelsUncached( // (#349/#344). Deliberately NOT the full applyProviderConfigHints pass — custom rows are a // user override, so their explicit contextWindow / inputModalities / reasoning fields must be // preserved verbatim (the hint pass would cap context and overwrite modalities from registry). + const mergedContext = typeof merged.contextWindow === "number" && merged.contextWindow > 0 + ? merged.contextWindow + : undefined; + const boundedMergedMaxInput = typeof merged.maxInputTokens === "number" && merged.maxInputTokens > 0 + ? (mergedContext !== undefined ? Math.min(merged.maxInputTokens, mergedContext) : merged.maxInputTokens) + : undefined; + const mergedWithHardBounds = boundedMergedMaxInput !== undefined + && boundedMergedMaxInput !== merged.maxInputTokens + ? { ...merged, maxInputTokens: boundedMergedMaxInput } + : merged; + const mergedSoftCandidates = [mergedWithHardBounds.autoCompactTokenLimit, configuredAutoCompact] + .filter((value): value is number => typeof value === "number" && value > 0); + const mergedWithAutoCompact: CatalogModel = mergedContext !== undefined && mergedSoftCandidates.length > 0 + ? { + ...mergedWithHardBounds, + autoCompactTokenLimit: clampAutoCompactTokenLimit( + mergedContext, + boundedMergedMaxInput, + Math.min(...mergedSoftCandidates), + ), + } + : mergedWithHardBounds; const enrichedProvider = enrichedByName.get(cm.provider) ?? rawProvider; - if (enrichedProvider && modelInList(enrichedProvider.noVisionModels, merged.id)) { - const current = merged.inputModalities ?? ["text"]; + if (enrichedProvider && modelInList(enrichedProvider.noVisionModels, mergedWithAutoCompact.id)) { + const current = mergedWithAutoCompact.inputModalities ?? ["text"]; if (!current.includes("image")) { - return { ...merged, inputModalities: [...current, "image"] }; + return { ...mergedWithAutoCompact, inputModalities: [...current, "image"] }; } } - return merged; + return mergedWithAutoCompact; }); // Custom rows override discovered rows that encode to the same Codex-facing slug. const customKeys = new Set(customModels.map(c => routedSlug(c.provider, c.id))); @@ -2015,7 +2124,15 @@ function augmentRoutedModelsWithCapturedOpenAiApiRows( ? Math.min(officialContext, userContext ?? officialContext, providerCap ?? officialContext) : undefined; const maxInputTokens = typeof officialMaxInput === "number" - ? Math.min(officialMaxInput, userMaxInput ?? officialMaxInput) + ? Math.min( + officialMaxInput, + userMaxInput ?? officialMaxInput, + contextWindow ?? officialMaxInput, + ) + : undefined; + const configuredAutoCompact = configuredAutoCompactTokenLimit(configured, id); + const autoCompactTokenLimit = contextWindow !== undefined && configuredAutoCompact !== undefined + ? clampAutoCompactTokenLimit(contextWindow, maxInputTokens, configuredAutoCompact) : undefined; return { provider: OPENAI_API_PROVIDER_ID, @@ -2023,6 +2140,7 @@ function augmentRoutedModelsWithCapturedOpenAiApiRows( owned_by: OPENAI_API_PROVIDER_ID, ...(contextWindow ? { contextWindow } : {}), ...(maxInputTokens ? { maxInputTokens } : {}), + ...(autoCompactTokenLimit !== undefined ? { autoCompactTokenLimit } : {}), ...(policy.modelInputModalities?.[id] ? { inputModalities: [...policy.modelInputModalities[id]!] } : {}), ...(policy.modelReasoningEfforts?.[id] ? { reasoningEfforts: [...policy.modelReasoningEfforts[id]!] } : {}), }; diff --git a/src/codex/catalog/sync.ts b/src/codex/catalog/sync.ts index 258474447d1..d5a894cd0e1 100644 --- a/src/codex/catalog/sync.ts +++ b/src/codex/catalog/sync.ts @@ -1158,7 +1158,7 @@ export function mergeCatalogEntriesForSync( isNativeAliasCatalogEntry(entry) && typeof entry.slug === "string" ? [entry.slug] : [] )), ), - openaiContextCap?: number, + openaiContextCap?: NativeContextLimitsInput, keepNativeChatGptOnV1 = false, ): RawEntry[] { // Retained for source compatibility with the original helper contract. Raw provider ids must diff --git a/src/codex/convergence.ts b/src/codex/convergence.ts index d5aeb893b71..ffb8af4aa99 100644 --- a/src/codex/convergence.ts +++ b/src/codex/convergence.ts @@ -54,6 +54,7 @@ import { disabledNativeSlugs, desktopAllowlistSuppressedNativeSlugs, NATIVE_OPENAI_MODELS, + nativeContextLimits, shouldIncludeAccountBoundNativeOpenAi, shouldIncludeNativeOpenAi, } from "./catalog/metadata"; @@ -286,6 +287,7 @@ function prepareCatalog( // selector-qualified rows when a live selector is configured. const observedNativeSlugs: string[] = []; const disabledNative = disabledNativeSlugs(config); + const openaiContextCap = nativeContextLimits(config); const nativeCatalogModels = mergeCatalogModelsWithNativeRecovery( active?.models ?? catalog.models ?? [], [catalog.models ?? [], ...nativeRecoverySources], @@ -304,6 +306,7 @@ function prepareCatalog( suppressedBareNativeSlugs, disabledNativeAccountSlugs: new Set(), multiAgentV2Enabled, + openaiContextCap, }); const accountBoundEntries = accountSelectors.length === 0 ? [] @@ -320,6 +323,7 @@ function prepareCatalog( disabledNativeAccountSlugs: new Set([...disabledNative].filter(slug => suppressedBareNativeSlugs.has(slug))), multiAgentV2Enabled, keepNativeChatGptOnV1: config.keepNativeChatGptOnV1 === true, + openaiContextCap, accountNativeSlugs, accountNativeSlugsBySelector, }).filter(entry => trustedAccountBoundNativeCatalogSlug(entry) !== undefined); @@ -352,6 +356,7 @@ function prepareCatalog( includeNativeOpenAi, accountBoundEntries, suppressedBareNativeSlugs, + openaiContextCap, policy: { ...CANONICAL_NATIVE_CATALOG_CONTENT_POLICY, nativeBackfillSlugs: [...availableBareNativeSlugs, ...observedNativeSlugs], diff --git a/src/config.ts b/src/config.ts index 10032fcbcf9..dcff7487999 100644 --- a/src/config.ts +++ b/src/config.ts @@ -73,6 +73,7 @@ import { type ProviderCostOverlay, } from "./types"; import { OPENAI_CODEX_PROVIDER_ID } from "./providers/openai-tiers"; +import { modelAutoCompactTokenLimitsConfigError } from "./providers/auto-compact-budget"; import { fastWireDeclarationError, hasFastWireCapabilityConflict } from "./providers/fastwire"; import { getProviderRegistryEntry, @@ -1101,6 +1102,17 @@ const configSchema = z.object({ message: maxInputError, }); } + const autoCompactError = modelAutoCompactTokenLimitsConfigError( + (provider as { modelAutoCompactTokenLimits?: unknown }).modelAutoCompactTokenLimits, + { requireNativeIds: name === OPENAI_CODEX_PROVIDER_ID }, + ); + if (autoCompactError) { + ctx.addIssue({ + code: "custom", + path: ["providers", redactSecretString(name), "modelAutoCompactTokenLimits"], + message: autoCompactError, + }); + } const reasoningSummariesError = booleanRecordConfigError( (provider as { modelSupportsReasoningSummaries?: unknown }).modelSupportsReasoningSummaries, "modelSupportsReasoningSummaries", diff --git a/src/providers/auto-compact-budget.ts b/src/providers/auto-compact-budget.ts new file mode 100644 index 00000000000..8275d10cffb --- /dev/null +++ b/src/providers/auto-compact-budget.ts @@ -0,0 +1,65 @@ +import { SUPPORTED_NATIVE_OPENAI_SLUGS } from "../codex/catalog/native-models"; +import { redactSecretString } from "../lib/redact"; + +const RESERVED_OBJECT_KEYS = new Set(["__proto__", "constructor", "prototype"]); + +function positiveSafeInteger(value: unknown): value is number { + return typeof value === "number" && Number.isSafeInteger(value) && value > 0; +} + +/** + * Resolve a client-facing soft compaction budget without changing any hard + * model limit. Configuration and measured input ceilings may only lower the + * default 90% envelope. + */ +export function clampAutoCompactTokenLimit( + contextWindow: number, + maxInputTokens?: number, + configuredLimit?: number, +): number { + const candidates = [Math.floor(contextWindow * 0.9), contextWindow]; + if (positiveSafeInteger(maxInputTokens)) candidates.push(maxInputTokens); + if (positiveSafeInteger(configuredLimit)) candidates.push(configuredLimit); + return Math.min(...candidates); +} + +export type AutoCompactBudgetValidationOptions = Readonly<{ + /** PATCH accepts null for whole-map and per-key deletion. */ + allowTombstones?: boolean; + /** The canonical ChatGPT provider accepts only exact supported native ids. */ + requireNativeIds?: boolean; +}>; + +/** Shared config/load/management boundary for per-model soft budgets. */ +export function modelAutoCompactTokenLimitsConfigError( + value: unknown, + options: AutoCompactBudgetValidationOptions = {}, +): string | null { + const field = "modelAutoCompactTokenLimits"; + if (value === undefined || (options.allowTombstones && value === null)) return null; + if (!value || typeof value !== "object" || Array.isArray(value)) { + return `${field} must be a plain object${options.allowTombstones ? " or null" : ""}`; + } + const prototype = Object.getPrototypeOf(value); + if (prototype !== Object.prototype && prototype !== null) { + return `${field} must be a plain object with own properties`; + } + for (const [modelId, entry] of Object.entries(value as Record)) { + const safeModelId = JSON.stringify(redactSecretString(modelId)); + if (!modelId.trim()) return `${field} keys must be nonblank model ids`; + if (RESERVED_OBJECT_KEYS.has(modelId)) { + return `${field} key ${safeModelId} is reserved`; + } + if (options.requireNativeIds + && (modelId.includes("/") || !SUPPORTED_NATIVE_OPENAI_SLUGS.has(modelId))) { + return `${field} key ${safeModelId} must be an exact supported native model id`; + } + if (options.allowTombstones && entry === null) continue; + if (!positiveSafeInteger(entry)) { + return `${field}[${safeModelId}] must be a positive safe integer${ + options.allowTombstones ? " or null" : "" + }`; + } + } + return null; +} diff --git a/src/server/auth-cors.ts b/src/server/auth-cors.ts index 2257f789234..460c2ebb293 100644 --- a/src/server/auth-cors.ts +++ b/src/server/auth-cors.ts @@ -26,6 +26,7 @@ import { effectiveGoogleMode, getProviderRegistryEntry, providerCodexAccountMode import { providerConfigSeed } from "../providers/derive"; import type { OcxConfig, OcxProviderConfig } from "../types"; import { openRouterRoutingConfigError } from "../providers/openrouter-routing"; +import { modelAutoCompactTokenLimitsConfigError } from "../providers/auto-compact-budget"; import { googleVertexLocationConfigError } from "../providers/google-vertex-location"; import { xaiResponsesOptInState } from "../providers/xai-responses-opt-in"; @@ -565,6 +566,8 @@ export function providerManagementConfigError(name: unknown, provider: unknown): if (contextOverlayError) return contextOverlayError; delete canonicalCandidate.contextWindow; delete canonicalCandidate.modelContextWindows; + // User-owned soft compaction policy; it does not alter the canonical transport seed. + delete canonicalCandidate.modelAutoCompactTokenLimits; const canonical = seed && sameCanonicalProviderSeed(canonicalCandidate, seed); if (!canonical) { return `provider ${name} must equal the canonical built-in provider seed`; @@ -607,6 +610,13 @@ export function providerManagementConfigError(name: unknown, provider: unknown): if (apiKeyTransportError) return `provider ${name} ${apiKeyTransportError}`; const maxInputError = positiveIntegerRecordConfigError(raw.modelMaxInputTokens, "modelMaxInputTokens"); if (maxInputError) return `provider ${name} ${maxInputError}`; + const autoCompactError = modelAutoCompactTokenLimitsConfigError( + raw.modelAutoCompactTokenLimits, + { requireNativeIds: name === "openai" }, + ); + if (autoCompactError) { + return `provider ${JSON.stringify(redactSecretString(name))} ${autoCompactError}`; + } const reasoningSummariesError = booleanRecordConfigError(raw.modelSupportsReasoningSummaries, "modelSupportsReasoningSummaries"); if (reasoningSummariesError) return `provider ${name} ${reasoningSummariesError}`; const reasoningSummaryDeliveryError = reasoningSummaryDeliveryRecordConfigError( @@ -707,6 +717,7 @@ export function safeConfigDTO(config: OcxConfig): unknown { "models", "contextWindow", "modelContextWindows", + "modelAutoCompactTokenLimits", "defaultMaxOutputTokens", "modelMaxOutputTokens", "openRouterRouting", diff --git a/src/server/management/model-rows.ts b/src/server/management/model-rows.ts index 8fb625e79d0..c4e6ca02104 100644 --- a/src/server/management/model-rows.ts +++ b/src/server/management/model-rows.ts @@ -63,6 +63,7 @@ export async function listManagementModelRows(config: OcxConfig): Promise { @@ -82,6 +83,9 @@ export async function listManagementModelRows(config: OcxConfig): Promise { diff --git a/src/server/management/provider-routes.ts b/src/server/management/provider-routes.ts index 47909e50f0b..5c107135f85 100644 --- a/src/server/management/provider-routes.ts +++ b/src/server/management/provider-routes.ts @@ -36,7 +36,7 @@ import { ProviderOutboundPolicyError, providerOutboundGet, providerOutboundPost, import { fetchCursorUsableModels } from "../../adapters/cursor/live-models"; import { parseAntigravityAvailableModels } from "../../providers/antigravity-models"; import { enrichProviderFromCatalog, listKeyLoginProviders } from "../../oauth/key-providers"; -import { deriveProviderPresets } from "../../providers/derive"; +import { deriveProviderPresets, providerConfigSeed } from "../../providers/derive"; import { effectiveGoogleMode, providerCodexAccountMode, providerMatchesRegistryTransport } from "../../providers/registry"; import { extractModelEnvelopeRows, @@ -54,6 +54,7 @@ import { clearThreadAccountMap } from "../../codex/routing"; import { primeCodexPoolQuotas } from "../../codex/auth-api"; import { clearModelCache, getProviderDiscoveryStatus } from "../../codex/model-cache"; import { DEFAULT_PROVIDER_CONTEXT_CAP, globalContextCapValue, providerContextCap, providerContextCaps, setAllProviderContextCaps, setGlobalContextCapValue, setProviderContextCap } from "../../providers/context-cap"; +import { modelAutoCompactTokenLimitsConfigError } from "../../providers/auto-compact-budget"; import { resolveCodexHomeDir } from "../../codex/home"; import { readUsageEntries } from "../../usage/log"; import { getUsageDebugLogEntries } from "../../usage/debug"; @@ -266,6 +267,29 @@ function applyProviderPatchFields( } touched = true; } + if (Object.hasOwn(rawBody, "modelAutoCompactTokenLimits")) { + const value = rawBody.modelAutoCompactTokenLimits; + const error = modelAutoCompactTokenLimitsConfigError(value, { + allowTombstones: true, + requireNativeIds: name === "openai", + }); + if (error) return { error }; + if (value === null) { + delete next.modelAutoCompactTokenLimits; + } else { + const budgets: Record = Object.assign( + Object.create(null) as Record, + next.modelAutoCompactTokenLimits ?? {}, + ); + for (const [model, budget] of Object.entries(value as Record)) { + if (budget === null) delete budgets[model]; + else budgets[model] = budget; + } + if (Object.keys(budgets).length > 0) next.modelAutoCompactTokenLimits = budgets; + else delete next.modelAutoCompactTokenLimits; + } + touched = true; + } if (Object.hasOwn(rawBody, "modelSupportsServiceTier")) { const value = rawBody.modelSupportsServiceTier; if (value === null) { @@ -369,6 +393,28 @@ function applyProviderPatchFields( return { next, touched, editorTouched, enablingOpenAi, headersTouched }; } +/** Validate the canonical OpenAI soft-budget overlay against a fresh registry seed. */ +function canonicalOpenAiBudgetPatchError( + provider: OcxProviderConfig, + rawBody: Record, + keys: string[], + config: OcxConfig, +): string | null { + if (!isCanonicalOpenAiForwardProvider(provider)) { + return "provider openai must be the canonical built-in provider"; + } + const entry = getProviderRegistryEntry("openai"); + if (!entry) return "provider openai registry seed is unavailable"; + const seed = providerConfigSeed(entry); + if (provider.codexAccountMode !== undefined) seed.codexAccountMode = provider.codexAccountMode; + if (provider.modelAutoCompactTokenLimits !== undefined) { + seed.modelAutoCompactTokenLimits = { ...provider.modelAutoCompactTokenLimits }; + } + const applied = applyProviderPatchFields("openai", seed, rawBody, keys, config); + if ("error" in applied) return applied.error; + return providerManagementConfigError("openai", applied.next); +} + export async function handleProviderRoutes(ctx: ManagementContext): Promise { const { req, url, config, deps, principal, convergeCodexCatalog, syncClaudeAgentDefsBestEffort } = ctx; @@ -405,6 +451,7 @@ export async function handleProviderRoutes(ctx: ManagementContext): Promise key === "requestPacing"); if (applied.editorTouched && !pacingOnly) { - const providerError = providerManagementConfigError(name, next); + const providerError = canonicalBudgetOnly + ? canonicalOpenAiBudgetPatchError(next, rawBody, keys, config) + : providerManagementConfigError(name, next); if (providerError) return jsonResponse({ error: providerError }, 400); - const serviceTierError = providerServiceTierConfigError(name, next); - if (serviceTierError) return jsonResponse({ error: serviceTierError }, 400); - const resolvedError = await providerDestinationResolvedError(name, next); - if (resolvedError) return jsonResponse({ error: resolvedError }, 400); + if (!canonicalBudgetOnly) { + const serviceTierError = providerServiceTierConfigError(name, next); + if (serviceTierError) return jsonResponse({ error: serviceTierError }, 400); + const resolvedError = await providerDestinationResolvedError(name, next); + if (resolvedError) return jsonResponse({ error: resolvedError }, 400); + } } else if (applied.enablingOpenAi) { // Same DNS gate as POST: Clash fake-IP only. Never honor a persisted // allowPrivateNetwork on this path — it must not bypass the built-in guard. @@ -682,15 +742,19 @@ export async function handleProviderRoutes(ctx: ManagementContext): Promise; /** Model-specific max input token limits. Values cap auto_compact_token_limit. */ modelMaxInputTokens?: Record; + /** + * Per-model soft compaction budgets. Values may only lower the effective + * context/max-input envelope; they never raise hard admission limits. + */ + modelAutoCompactTokenLimits?: Record; /** * Provider-wide fallback for chat-completions `max_tokens` when the caller omits * Responses `max_output_tokens`. Adapters still let an explicit request win. diff --git a/tests/auto-compact-budget.test.ts b/tests/auto-compact-budget.test.ts new file mode 100644 index 00000000000..f7973d92473 --- /dev/null +++ b/tests/auto-compact-budget.test.ts @@ -0,0 +1,57 @@ +import { describe, expect, test } from "bun:test"; + +import { + clampAutoCompactTokenLimit, + modelAutoCompactTokenLimitsConfigError, +} from "../src/providers/auto-compact-budget"; + +describe("per-model auto-compaction budgets", () => { + test("configuration only lowers the effective hard-limit envelope", () => { + expect(clampAutoCompactTokenLimit(1_000)).toBe(900); + expect(clampAutoCompactTokenLimit(1_000, 800)).toBe(800); + expect(clampAutoCompactTokenLimit(1_000, 800, 700)).toBe(700); + expect(clampAutoCompactTokenLimit(1_000, 800, 5_000)).toBe(800); + }); + + test("one validation contract handles config, native ids, and PATCH tombstones", () => { + expect(modelAutoCompactTokenLimitsConfigError({ model: 64_000 })).toBeNull(); + expect(modelAutoCompactTokenLimitsConfigError( + { "gpt-5.6-sol": 64_000 }, + { requireNativeIds: true }, + )).toBeNull(); + expect(modelAutoCompactTokenLimitsConfigError( + { model: null }, + { allowTombstones: true }, + )).toBeNull(); + expect(modelAutoCompactTokenLimitsConfigError(null, { allowTombstones: true })).toBeNull(); + + for (const invalid of [ + null, + [], + { model: 0 }, + { model: 1.5 }, + { model: Number.MAX_SAFE_INTEGER + 1 }, + { model: null }, + JSON.parse('{"__proto__": 1000}'), + { constructor: 1000 }, + Object.create({ inherited: 1000 }), + ]) { + expect(modelAutoCompactTokenLimitsConfigError(invalid)).not.toBeNull(); + } + expect(modelAutoCompactTokenLimitsConfigError( + { "team/gpt-5.6-sol": 64_000 }, + { requireNativeIds: true }, + )).toContain("exact supported native model id"); + expect(modelAutoCompactTokenLimitsConfigError( + { "team/gpt-5.6-sol": null }, + { allowTombstones: true, requireNativeIds: true }, + )).toContain("exact supported native model id"); + }); + + test("validation errors redact secret-shaped model ids", () => { + const secret = "api_key=sk-secret-provider-key"; + const error = modelAutoCompactTokenLimitsConfigError({ [secret]: 0 }); + expect(error).toContain("[REDACTED]"); + expect(error).not.toContain("sk-secret-provider-key"); + }); +}); diff --git a/tests/codex-catalog.test.ts b/tests/codex-catalog.test.ts index f4e8bca8639..bf5656779b0 100644 --- a/tests/codex-catalog.test.ts +++ b/tests/codex-catalog.test.ts @@ -199,12 +199,25 @@ describe("combo catalog capability intersection", () => { owned_by: "combo", contextWindow: 128_000, maxInputTokens: 100_000, + autoCompactTokenLimit: 100_000, inputModalities: ["text"], reasoningEfforts: ["low", "medium"], defaultReasoningEffort: "medium", }); }); + test("never advertises combo max-input or compaction above its smallest final window", () => { + const derived = deriveComboCatalogModel("bounded", normalizedCombo(), [ + { provider: "a", id: "m1", contextWindow: 700_000, maxInputTokens: 922_000 }, + { provider: "b", id: "m2", contextWindow: 800_000, maxInputTokens: 900_000 }, + ]); + expect(derived).toMatchObject({ + contextWindow: 700_000, + maxInputTokens: 700_000, + autoCompactTokenLimit: 630_000, + }); + }); + test("handles vision, missing modalities, reasoning defaults, and parallel tools conservatively", () => { expect(deriveComboCatalogModel("vision", normalizedCombo({ defaultEffort: "low" }), [ memberA, @@ -975,8 +988,8 @@ describe("combo catalog capability intersection", () => { port: 10100, defaultProvider: "a", providers: { - a: { adapter: "openai-chat", baseUrl: "https://a.example/v1", liveModels: false, models: ["m1"], modelContextWindows: { m1: 200_000 } }, - b: { adapter: "openai-chat", baseUrl: "https://b.example/v1", liveModels: false, models: ["m2"], modelContextWindows: { m2: 128_000 } }, + a: { adapter: "openai-chat", baseUrl: "https://a.example/v1", liveModels: false, models: ["m1"], modelContextWindows: { m1: 200_000 }, modelAutoCompactTokenLimits: { m1: 150_000 } }, + b: { adapter: "openai-chat", baseUrl: "https://b.example/v1", liveModels: false, models: ["m2"], modelContextWindows: { m2: 128_000 }, modelAutoCompactTokenLimits: { m2: 80_000 } }, }, combos: { mixed: { targets: [{ provider: "a", model: "m1" }, { provider: "b", model: "m2" }] }, @@ -993,6 +1006,14 @@ describe("combo catalog capability intersection", () => { expect(first.map(model => `${model.provider}/${model.id}`)).toEqual([ "a/m1", "b/m2", "combo/mixed", ]); + expect(first.find(model => model.provider === "combo" && model.id === "mixed")) + .toMatchObject({ contextWindow: 128_000, maxInputTokens: 128_000, autoCompactTokenLimit: 80_000 }); + expect(buildCatalogEntries(nativeTemplate(), [], first) + .find(entry => entry.slug === "combo/mixed")).toMatchObject({ + context_window: 128_000, + max_context_window: 128_000, + auto_compact_token_limit: 80_000, + }); expect(filterCatalogVisibleModels(first, config).some(model => model.id === "mixed")).toBe(false); expect(warn).toHaveBeenCalledTimes(1); expect(String(warn.mock.calls[0]?.[0])).toContain("[REDACTED]"); @@ -1892,6 +1913,54 @@ describe("configured CatalogModel displayName -> catalog display_name", () => { } }); + test("a custom row clamps its soft budget to the provider max-input ceiling", async () => { + const models = await gatherRoutedModels({ + port: 10100, + defaultProvider: "custom-budget", + providers: { + "custom-budget": { + baseUrl: "https://custom-budget.example.test/v1", + adapter: "openai-chat", + liveModels: false, + models: [], + modelMaxInputTokens: { renamed: 60_000, contextless: 60_000 }, + modelAutoCompactTokenLimits: { renamed: 80_000, contextless: 10_000 }, + }, + }, + customModels: [{ + id: "custom-budget-row", + provider: "custom-budget", + modelId: "renamed", + contextWindow: 321_000, + }, { + id: "custom-budget-contextless", + provider: "custom-budget", + modelId: "contextless", + }], + }); + const model = models.find(row => row.provider === "custom-budget" && row.id === "renamed"); + expect(model).toMatchObject({ + contextWindow: 321_000, + maxInputTokens: 60_000, + autoCompactTokenLimit: 60_000, + }); + expect(buildCatalogEntries(nativeTemplate(), [], models) + .find(entry => entry.slug === "custom-budget/renamed")).toMatchObject({ + context_window: 321_000, + max_context_window: 321_000, + auto_compact_token_limit: 60_000, + }); + const contextless = models.find(row => row.provider === "custom-budget" && row.id === "contextless"); + expect(contextless).toMatchObject({ maxInputTokens: 60_000 }); + expect(contextless).not.toHaveProperty("autoCompactTokenLimit"); + expect(buildCatalogEntries(nativeTemplate(), [], models) + .find(entry => entry.slug === "custom-budget/contextless")).toMatchObject({ + context_window: 128_000, + max_context_window: 128_000, + auto_compact_token_limit: 60_000, + }); + }); + test("a customModel reasoning ladder overrides the inherited provider ladder end-to-end", async () => { clearModelCache("custom-provider"); const originalFetch = globalThis.fetch; @@ -3087,6 +3156,38 @@ describe("Codex catalog routed normalization", () => { } }); + test("bare and account-qualified native rows inherit one lowering-only soft budget", () => { + const entries = buildCatalogEntries( + nativeTemplate(), + NATIVE_OPENAI_MODELS, + [], + undefined, + false, + "default", + new Set(), + ["team"], + new Set(), + new Set(), + { modelAutoCompactTokenLimits: { "gpt-5.6-sol": 120_000 } }, + ); + const bare = entries.find(entry => entry.slug === "gpt-5.6-sol"); + const account = entries.find(entry => entry.slug === "team/gpt-5.6-sol"); + + expect(bare).toMatchObject({ + context_window: 272_000, + max_context_window: 272_000, + auto_compact_token_limit: 120_000, + }); + expect(account).toMatchObject({ + context_window: 272_000, + max_context_window: 272_000, + auto_compact_token_limit: 120_000, + opencodex_catalog_kind: CODEX_ACCOUNT_BOUND_CATALOG_KIND, + }); + expect(account?.context_window).toBe(bare?.context_window); + expect(account?.max_context_window).toBe(bare?.max_context_window); + }); + test("routed entries drop stale native max context with the template window (#992)", () => { const template = { ...nativeTemplate(), @@ -4655,6 +4756,7 @@ describe("Codex catalog routed normalization", () => { apiKey: "sk-test", models: ["static-model"], modelContextWindows: { "static-model": 321_000 }, + modelAutoCompactTokenLimits: { "static-model": 80_000 }, modelInputModalities: { "static-model": ["text", "image"] }, }, }, @@ -4664,10 +4766,68 @@ describe("Codex catalog routed normalization", () => { expect(routed?.context_window).toBe(321_000); expect(routed?.max_context_window).toBe(321_000); - expect(routed?.auto_compact_token_limit).toBe(288_900); + expect(routed?.auto_compact_token_limit).toBe(80_000); expect(routed?.input_modalities).toEqual(["text", "image"]); }); + test("an unknown window ignores the configured soft budget instead of treating 128k as policy evidence", async () => { + globalThis.fetch = (async () => new Response("{}", { status: 503 })) as typeof fetch; + const models = await gatherRoutedModels({ + port: 10100, + defaultProvider: "unknown-soft", + providers: { + "unknown-soft": { + adapter: "openai-chat", + baseUrl: "https://unknown-soft.test/v1", + liveModels: false, + models: ["model"], + modelAutoCompactTokenLimits: { model: 10_000 }, + }, + }, + }); + const model = models.find(row => row.provider === "unknown-soft" && row.id === "model"); + expect(model).not.toHaveProperty("autoCompactTokenLimit"); + + const emitted = buildCatalogEntries(nativeTemplate(), [], models) + .find(entry => entry.slug === "unknown-soft/model"); + expect(emitted).toMatchObject({ + context_window: 128_000, + max_context_window: 128_000, + auto_compact_token_limit: 115_200, + }); + }); + + test("a max-input-only Combo member ignores the configured soft budget", async () => { + const models = await gatherRoutedModels({ + port: 10100, + defaultProvider: "max-only", + providers: { + "max-only": { + adapter: "openai-chat", + baseUrl: "https://max-only.test/v1", + liveModels: false, + models: [], + modelMaxInputTokens: { model: 80_000 }, + modelAutoCompactTokenLimits: { model: 10_000 }, + }, + }, + combos: { + "max-only-combo": { + strategy: "failover", + targets: [{ provider: "max-only", model: "model", weight: 1 }], + }, + }, + }); + + expect(models.find(row => row.provider === "max-only" && row.id === "model")).toBeUndefined(); + expect(models.find(row => row.provider === "combo" && row.id === "max-only-combo")) + .toMatchObject({ + contextWindow: 80_000, + maxInputTokens: 80_000, + autoCompactTokenLimit: 72_000, + }); + }); + // #1073's exact reproduction: a provider whose /models returns nothing but ids. Two cases, // deliberately not one — a single test that sets `modelContextWindows` would keep passing // with the provider-wide `?? prov.contextWindow` fallback deleted, because the per-model @@ -4874,6 +5034,7 @@ describe("Codex catalog routed normalization", () => { apiKey: "sk-test", contextWindow: 128_000, modelContextWindows: { "wide-model": 100_000 }, + modelMaxInputTokens: { "wide-model": 200_000 }, modelInputModalities: { "wide-model": ["text"] }, }, }, @@ -4881,6 +5042,7 @@ describe("Codex catalog routed normalization", () => { expect(models.find(m => m.id === "wide-model")).toMatchObject({ contextWindow: 100_000, + maxInputTokens: 100_000, inputModalities: ["text"], }); expect(models.find(m => m.id === "small-model")?.contextWindow).toBe(64_000); @@ -5195,11 +5357,12 @@ describe("OpenAI API trusted catalog augmentation", () => { test("user values only lower trusted context and max-input baselines", () => { const lowered = augmentRoutedModelsWithRegistryOpenAiApiRows([], openAiApiCatalogConfig({ - modelContextWindows: { "gpt-5.6-sol": 350_000, "gpt-5.6-terra": 2_000_000 }, - modelMaxInputTokens: { "gpt-5.6-sol": 300_000, "gpt-5.6-terra": 945_000 }, + modelContextWindows: { "gpt-5.6-sol": 350_000, "gpt-5.6-terra": 2_000_000, "gpt-5.6-luna": 350_000 }, + modelMaxInputTokens: { "gpt-5.6-sol": 300_000, "gpt-5.6-terra": 945_000, "gpt-5.6-luna": 900_000 }, })); expect(lowered.find(row => row.id === "gpt-5.6-sol")).toMatchObject({ contextWindow: 350_000, maxInputTokens: 300_000 }); expect(lowered.find(row => row.id === "gpt-5.6-terra")).toMatchObject({ contextWindow: 1_050_000, maxInputTokens: 922_000 }); + expect(lowered.find(row => row.id === "gpt-5.6-luna")).toMatchObject({ contextWindow: 350_000, maxInputTokens: 350_000 }); }); test("routed auto-compaction is bounded by max-input after effective context caps", () => { diff --git a/tests/codex-convergence-account-selectors.test.ts b/tests/codex-convergence-account-selectors.test.ts index 48d15e6da78..a0e367091b9 100644 --- a/tests/codex-convergence-account-selectors.test.ts +++ b/tests/codex-convergence-account-selectors.test.ts @@ -325,6 +325,21 @@ test("convergence renders account-qualified rows and preserves only non-generate } }); +test("convergence preserves one configured soft budget on bare and account-native rows", async () => { + writeCatalog([nativeEntry()]); + const nextConfig = config(true); + nextConfig.providers.openai!.modelAutoCompactTokenLimits = { "gpt-5.6-sol": 120_000 }; + + const models = (await convergeCatalog(nextConfig)).models ?? []; + for (const slug of ["gpt-5.6-sol", "desktop/gpt-5.6-sol", "team/gpt-5.6-sol"]) { + expect(models.find(entry => entry.slug === slug)).toMatchObject({ + context_window: 272_000, + max_context_window: 272_000, + auto_compact_token_limit: 120_000, + }); + } +}); + test("disabling the picker removes generated rows, restores bare rows, and retains foreign rows", async () => { writeCatalog([ nativeEntry("hide"), diff --git a/tests/config.test.ts b/tests/config.test.ts index 108463c7491..89798607871 100644 --- a/tests/config.test.ts +++ b/tests/config.test.ts @@ -1619,6 +1619,40 @@ describe("opencodex config defaults", () => { expect(readConfigDiagnostics().error).toContain("providers.custom.modelMaxInputTokens"); }); + test("disk config validates per-model auto-compaction budgets with native exact ids", () => { + writeConfig({ + port: 10100, + providers: { + custom: { + adapter: "openai-chat", + baseUrl: "https://example.test/v1", + modelAutoCompactTokenLimits: { model: 1.5 }, + }, + }, + defaultProvider: "custom", + }); + expect(readConfigDiagnostics().source).toBe("fallback"); + expect(readConfigDiagnostics().error).toContain("providers.custom.modelAutoCompactTokenLimits"); + + rmSync(testDir, { recursive: true, force: true }); + mkdirSync(testDir, { recursive: true }); + writeConfig({ + port: 10100, + providers: { + openai: { + adapter: "openai-responses", + baseUrl: "https://chatgpt.com/backend-api/codex", + authMode: "forward", + codexAccountMode: "direct", + modelAutoCompactTokenLimits: { "team/gpt-5.6-sol": 64_000 }, + }, + }, + defaultProvider: "openai", + }); + expect(readConfigDiagnostics().source).toBe("fallback"); + expect(readConfigDiagnostics().error).toContain("exact supported native model id"); + }); + test("disk config preserves valid OpenRouter routing and rejects invalid destinations", () => { writeConfig({ port: 10100, diff --git a/tests/management-provider-validation.test.ts b/tests/management-provider-validation.test.ts index 7aaaba950ae..c9621fc79e4 100644 --- a/tests/management-provider-validation.test.ts +++ b/tests/management-provider-validation.test.ts @@ -476,6 +476,18 @@ describe("provider management validation", () => { expect(secretNameError).toContain("[REDACTED]"); }); + test("provider management redacts provider names from auto-compaction validation errors", () => { + const secretName = "sk-super-secret-9876"; + const error = providerManagementConfigError(secretName, { + adapter: "openai-chat", + baseUrl: "https://api.example.test/v1", + modelAutoCompactTokenLimits: { model: 0 }, + })!; + expect(error).toContain("modelAutoCompactTokenLimits"); + expect(error).not.toContain(secretName); + expect(error).toContain("[REDACTED]"); + }); + test("provider request pacing PATCH persists provider and model limits without catalog churn", async () => { if (existsSync(TEST_DIR)) rmSync(TEST_DIR, { recursive: true }); mkdirSync(TEST_DIR, { recursive: true }); @@ -670,13 +682,18 @@ describe("provider management validation", () => { freshHome(); const server = startServer(0); try { - expect((await seedProvider(server.url, { modelContextWindows: { "deepseek-v4-flash": 900000 } })).status).toBe(200); + expect((await seedProvider(server.url, { + modelContextWindows: { "deepseek-v4-flash": 900000 }, + modelAutoCompactTokenLimits: { "deepseek-v4-flash": 120000 }, + })).status).toBe(200); expect((await seedProvider(server.url, {})).status).toBe(200); // The user's key survives, and the registry seed is NOT persisted into user config: // router.ts fills registry values beneath user entries at resolve time, so writing // them here would be a side effect of an unrelated save. expect(loadConfig().providers["opencode-go"]?.modelContextWindows).toEqual({ "deepseek-v4-flash": 900000 }); + expect(loadConfig().providers["opencode-go"]?.modelAutoCompactTokenLimits) + .toEqual({ "deepseek-v4-flash": 120000 }); } finally { await server.stop(true); } @@ -686,11 +703,19 @@ describe("provider management validation", () => { freshHome(); const server = startServer(0); try { - expect((await seedProvider(server.url, { modelContextWindows: { "deepseek-v4-flash": 900000 } })).status).toBe(200); - expect((await seedProvider(server.url, { modelContextWindows: { "kimi-k3": 300000 } })).status).toBe(200); + expect((await seedProvider(server.url, { + modelContextWindows: { "deepseek-v4-flash": 900000 }, + modelAutoCompactTokenLimits: { "deepseek-v4-flash": 120000 }, + })).status).toBe(200); + expect((await seedProvider(server.url, { + modelContextWindows: { "kimi-k3": 300000 }, + modelAutoCompactTokenLimits: { "kimi-k3": 90000 }, + })).status).toBe(200); expect(loadConfig().providers["opencode-go"]?.modelContextWindows) .toEqual({ "deepseek-v4-flash": 900000, "kimi-k3": 300000 }); + expect(loadConfig().providers["opencode-go"]?.modelAutoCompactTokenLimits) + .toEqual({ "deepseek-v4-flash": 120000, "kimi-k3": 90000 }); } finally { await server.stop(true); } @@ -841,6 +866,9 @@ describe("provider management validation", () => { ["map-shape", { ...canonicalDirect, modelContextWindows: [] }], ["map-value", { ...canonicalDirect, modelContextWindows: { "gpt-5.6-sol": "wide" } }], ["map-key", { ...canonicalDirect, modelContextWindows: { " ": 500_000 } }], + ["soft-map-shape", { ...canonicalDirect, modelAutoCompactTokenLimits: [] }], + ["soft-map-value", { ...canonicalDirect, modelAutoCompactTokenLimits: { "gpt-5.6-sol": 1e100 } }], + ["soft-map-key", { ...canonicalDirect, modelAutoCompactTokenLimits: { "team/gpt-5.6-sol": 120_000 } }], ] as const) { const response = await fetch(new URL("/api/providers", server.url), { method: "POST", @@ -855,6 +883,7 @@ describe("provider management validation", () => { // the proxy advertises. for (const [, provider] of [ ["per-model", { ...canonicalDirect, modelContextWindows: { "gpt-5.6-sol": 500_000 } }], + ["soft-per-model", { ...canonicalDirect, modelAutoCompactTokenLimits: { "gpt-5.6-sol": 120_000 } }], ["provider-wide", { ...canonicalDirect, contextWindow: 500_000 }], ] as const) { const response = await fetch(new URL("/api/providers", server.url), { @@ -865,6 +894,19 @@ describe("provider management validation", () => { expect(response.status).toBe(200); } + // POST enriches the canonical row with registry-owned capabilities. A later budget-only + // PATCH must validate the overlay against a fresh seed instead of rejecting those fields. + const patchedBudget = await fetch(new URL("/api/providers?name=openai", server.url), { + method: "PATCH", + headers: { "content-type": "application/json" }, + body: JSON.stringify({ modelAutoCompactTokenLimits: { "gpt-5.6-terra": 90_000 } }), + }); + expect(patchedBudget.status).toBe(200); + expect(loadConfig().providers.openai.modelAutoCompactTokenLimits).toEqual({ + "gpt-5.6-sol": 120_000, + "gpt-5.6-terra": 90_000, + }); + const acceptedCustom = await fetch(new URL("/api/providers", server.url), { method: "POST", headers: { "content-type": "application/json" }, @@ -997,9 +1039,16 @@ describe("provider management validation", () => { expect(legacy.status).toBe(400); const dto = await fetch(new URL("/api/config", server.url)).then(response => response.json()) as { - providers: Record; + providers: Record; + }>; }; expect(dto.providers.openai.codexAccountMode).toBe("direct"); + expect(dto.providers.openai.modelAutoCompactTokenLimits).toEqual({ + "gpt-5.6-sol": 120_000, + "gpt-5.6-terra": 90_000, + }); expect(dto.providers["openai-multi"]).toBeUndefined(); expect(dto.providers["custom-max-input"]).not.toHaveProperty("modelMaxInputTokens"); @@ -2735,6 +2784,7 @@ describe("provider management validation", () => { models: ["wide", "narrow"], contextWindow: 256_000, modelContextWindows: { narrow: 64_000 }, + modelAutoCompactTokenLimits: { narrow: 32_000 }, modelSupportsServiceTier: { narrow: false }, }, }, @@ -2758,27 +2808,32 @@ describe("provider management validation", () => { name: string; contextWindow?: number; modelContextWindows?: Record; + modelAutoCompactTokenLimits?: Record; }>; expect(rows.find(row => row.name === "relay")).toMatchObject({ contextWindow: 256_000, modelContextWindows: { narrow: 64_000 }, + modelAutoCompactTokenLimits: { narrow: 32_000 }, modelSupportsServiceTier: { narrow: false }, }); const updated = await request("PATCH", { contextWindow: 350_000, modelContextWindows: { wide: 350_000 }, + modelAutoCompactTokenLimits: { wide: 100_000 }, modelSupportsServiceTier: { wide: true }, }); expect(updated?.status).toBe(200); expect(liveConfig.providers.relay).toMatchObject({ contextWindow: 350_000, modelContextWindows: { wide: 350_000, narrow: 64_000 }, + modelAutoCompactTokenLimits: { wide: 100_000, narrow: 32_000 }, modelSupportsServiceTier: { wide: true, narrow: false }, }); expect(loadConfig().providers.relay).toMatchObject({ contextWindow: 350_000, modelContextWindows: { wide: 350_000, narrow: 64_000 }, + modelAutoCompactTokenLimits: { wide: 100_000, narrow: 32_000 }, modelSupportsServiceTier: { wide: true, narrow: false }, }); @@ -2792,6 +2847,9 @@ describe("provider management validation", () => { { modelContextWindows: { wide: 1e100 } }, { modelContextWindows: { "": 100_000 } }, { modelContextWindows: { wide: -1 } }, + { modelAutoCompactTokenLimits: { wide: 1e100 } }, + { modelAutoCompactTokenLimits: { "": 100_000 } }, + { modelAutoCompactTokenLimits: { constructor: 100_000 } }, { modelSupportsServiceTier: { wide: "yes" } }, { modelSupportsServiceTier: { "": true } }, ]) { @@ -2800,24 +2858,32 @@ describe("provider management validation", () => { expect(liveConfig.providers.relay).toMatchObject({ contextWindow: 350_000, modelContextWindows: { wide: 350_000, narrow: 64_000 }, + modelAutoCompactTokenLimits: { wide: 100_000, narrow: 32_000 }, modelSupportsServiceTier: { wide: true, narrow: false }, }); expect((await request("PATCH", { modelContextWindows: { wide: null } }))?.status).toBe(200); expect(liveConfig.providers.relay.modelContextWindows).toEqual({ narrow: 64_000 }); + expect((await request("PATCH", { modelAutoCompactTokenLimits: { wide: null } }))?.status).toBe(200); + expect(liveConfig.providers.relay.modelAutoCompactTokenLimits).toEqual({ narrow: 32_000 }); + expect(loadConfig().providers.relay.modelAutoCompactTokenLimits).toEqual({ narrow: 32_000 }); + expect((await request("PATCH", { modelSupportsServiceTier: { wide: null } }))?.status).toBe(200); expect(liveConfig.providers.relay.modelSupportsServiceTier).toEqual({ narrow: false }); const cleared = await request("PATCH", { contextWindow: null, modelContextWindows: null, + modelAutoCompactTokenLimits: null, modelSupportsServiceTier: null, }); expect(cleared?.status).toBe(200); expect(liveConfig.providers.relay.contextWindow).toBeUndefined(); expect(liveConfig.providers.relay.modelContextWindows).toBeUndefined(); + expect(liveConfig.providers.relay.modelAutoCompactTokenLimits).toBeUndefined(); expect(liveConfig.providers.relay.modelSupportsServiceTier).toBeUndefined(); + expect(loadConfig().providers.relay.modelAutoCompactTokenLimits).toBeUndefined(); }); test("provider PATCH manages custom headers with merge and clear semantics", async () => { @@ -2999,7 +3065,7 @@ describe("provider management validation", () => { "x-opencode-client": "desktop", }); }); - test("concurrent provider PATCHes merge different headers", async () => { + test("concurrent provider PATCHes serialize mixed fields and per-model soft budgets", async () => { if (existsSync(TEST_DIR)) rmSync(TEST_DIR, { recursive: true }); mkdirSync(TEST_DIR, { recursive: true }); process.env.OPENCODEX_HOME = TEST_DIR; @@ -3034,6 +3100,19 @@ describe("provider management validation", () => { expect(first?.status).toBe(200); expect(second?.status).toBe(200); expect(liveConfig.providers.hdr.headers).toEqual({ "X-A": "a", "X-B": "b" }); + + const [third, fourth] = await Promise.all([ + patch("hdr", { + headers: { "X-C": "c" }, + modelAutoCompactTokenLimits: { m1: 80_000 }, + }), + patch("hdr", { modelAutoCompactTokenLimits: { m2: 64_000 } }), + ]); + expect(third?.status).toBe(200); + expect(fourth?.status).toBe(200); + expect(liveConfig.providers.hdr.headers).toEqual({ "X-A": "a", "X-B": "b", "X-C": "c" }); + expect(liveConfig.providers.hdr.modelAutoCompactTokenLimits).toEqual({ m1: 80_000, m2: 64_000 }); + expect(loadConfig().providers.hdr.modelAutoCompactTokenLimits).toEqual({ m1: 80_000, m2: 64_000 }); }); test("provider context-cap API persists toggles and annotates model rows", async () => { if (existsSync(TEST_DIR)) rmSync(TEST_DIR, { recursive: true }); diff --git a/tests/native-model-toggle.test.ts b/tests/native-model-toggle.test.ts index 3dabda66c65..b8d153db6ce 100644 --- a/tests/native-model-toggle.test.ts +++ b/tests/native-model-toggle.test.ts @@ -123,16 +123,55 @@ describe("native GPT model toggles (bare slugs in disabledModels)", () => { expect(nativeModelRows(both).find(r => r.slug === "gpt-5.6-sol")?.contextWindow).toBe(350_000); }); + test("a per-model soft budget lowers compaction without changing native hard limits", () => { + const configured = { + providers: { openai: { modelAutoCompactTokenLimits: { "gpt-5.6-sol": 120_000 } } }, + } as never; + const row = nativeModelRows(configured).find(item => item.slug === "gpt-5.6-sol"); + expect(row).toMatchObject({ + contextWindow: 272_000, + maxInputTokens: 272_000, + autoCompactTokenLimit: 120_000, + }); + + const oversized = { + providers: { openai: { modelAutoCompactTokenLimits: { "gpt-5.6-sol": 2_000_000 } } }, + } as never; + expect(nativeModelRows(oversized).find(item => item.slug === "gpt-5.6-sol")) + .toMatchObject({ contextWindow: 272_000, maxInputTokens: 272_000, autoCompactTokenLimit: 244_800 }); + }); + test("the on-disk catalog entry lands at the same width as the dashboard row", () => { // Regression: applyNativeOpenAiContextOverride used to re-read the static table and apply // only the cap, so a saved per-model window showed up in /api/models and was written back // at 922,000 in the Codex catalog. - const limits = { providers: { openai: { modelContextWindows: { "gpt-5.6-sol": 500_000 } } } } as never; + const limits = { providers: { openai: { + modelContextWindows: { "gpt-5.6-sol": 500_000 }, + modelAutoCompactTokenLimits: { "gpt-5.6-sol": 120_000 }, + } } } as never; const entry: Record = { slug: "gpt-5.6-sol", context_window: 922_000, max_context_window: 922_000 }; applyNativeOpenAiContextOverride(entry as never, nativeContextLimits(limits)); expect(entry.context_window).toBe(500_000); expect(entry.max_context_window).toBe(500_000); - expect(entry.auto_compact_token_limit).toBe(450_000); // 90% of the narrowed window + expect(entry.auto_compact_token_limit).toBe(120_000); + }); + + test("the on-disk catalog preserves a lower retained native compaction threshold", () => { + const retained = { + slug: "gpt-5.4-mini", + context_window: 272_000, + max_context_window: 272_000, + auto_compact_token_limit: 100_000, + }; + applyNativeOpenAiContextOverride(retained as never, nativeContextLimits({})); + expect(retained.auto_compact_token_limit).toBe(100_000); + + const configured = { + providers: { openai: { modelAutoCompactTokenLimits: { "gpt-5.4-mini": 80_000 } } }, + } as never; + const lowered = { ...retained }; + applyNativeOpenAiContextOverride(lowered as never, nativeContextLimits(configured)); + expect(lowered.auto_compact_token_limit).toBe(80_000); }); test("the advertised native window stays inside the measured ceiling after Codex spends 95% of it", () => { From b2f95e1311bffca644536796b62bf99316af601d Mon Sep 17 00:00:00 2001 From: liyongjie Date: Tue, 25 Aug 2026 11:28:25 +0800 Subject: [PATCH 11/29] fix(anthropic): apply provider default reasoning effort when caller omits it (#2494) Anthropic-compatible models that require explicit thinking now work when the caller omits reasoning.effort, as long as the provider declares modelDefaultReasoningEfforts. Before this, those defaults were visible in model metadata but the adapter did not apply them to outbound /v1/messages requests, so always-thinking gateways could reject otherwise valid calls. Explicit caller reasoning still wins over the provider default. Co-authored-by: liyongjie.103 --- src/adapters/anthropic.ts | 15 +++++++++++---- tests/anthropic-reasoning.test.ts | 22 ++++++++++++++++++++++ 2 files changed, 33 insertions(+), 4 deletions(-) diff --git a/src/adapters/anthropic.ts b/src/adapters/anthropic.ts index 865b1a9a640..d8add210514 100644 --- a/src/adapters/anthropic.ts +++ b/src/adapters/anthropic.ts @@ -27,6 +27,7 @@ import { CLAUDE_CODE_HEADERS, claudeCodeSessionId } from "./client-fingerprint"; import { buildNonOpenAIToolCatalogNudgeForTools } from "./tool-catalog-nudge"; import { decodeServerSentEvents } from "../lib/sse-decoder"; import { isTranslatorBudgetExceededError, retainTranslatedEventBatch, type TranslatorBudget } from "../lib/translator-budget"; +import { modelRecordValue } from "../reasoning-effort"; /** Map a user content part to an Anthropic content block (text or image source). */ function toAnthropicContentPart(p: OcxContentPart): unknown { @@ -518,6 +519,11 @@ function adaptiveEffort(effort: string): string { return effort === "minimal" ? "low" : effort; } +function defaultReasoningEffort(provider: OcxProviderConfig, modelId: string): string | undefined { + const value = modelRecordValue(provider.modelDefaultReasoningEfforts, modelId); + return typeof value === "string" && value.trim() ? value.trim() : undefined; +} + function usageFromAnthropic(usage: Record | undefined): OcxUsage | undefined { if (!usage) return undefined; const hasCache = usage.cache_read_input_tokens !== undefined || usage.cache_creation_input_tokens !== undefined; @@ -926,16 +932,17 @@ export function createAnthropicAdapter(provider: OcxProviderConfig, cacheRetenti // anyway, and thinking shares the caller's `max_tokens` — which truncates a small-budget // request before it can emit its stop sequence (#545). Say "disabled" out loud where the // model both defaults to thinking and accepts being told not to. - if (parsed.options.reasoning === "none" && supportsExplicitThinkingDisable(parsed.modelId)) { + const effectiveReasoning = parsed.options.reasoning ?? defaultReasoningEffort(provider, parsed.modelId); + if (effectiveReasoning === "none" && supportsExplicitThinkingDisable(parsed.modelId)) { body.thinking = { type: "disabled" }; - } else if (typeof parsed.options.reasoning === "string" && parsed.options.reasoning !== "none") { + } else if (typeof effectiveReasoning === "string" && effectiveReasoning !== "none") { if (usesAdaptiveThinking(parsed.modelId)) { // Adaptive-thinking models replace the token budget with an effort knob and reject // `thinking.type: "enabled"` outright. `max_tokens` still caps thinking plus visible // output, so high effort needs the same total-token headroom as budget thinking or a // default 8192-token request can spend everything on thought and return empty text. body.thinking = { type: "adaptive" }; - const effort = adaptiveEffort(parsed.options.reasoning); + const effort = adaptiveEffort(effectiveReasoning); body.output_config = { effort }; const explicitMaxOut = parsed.options.maxOutputTokens; const wantBudget = reasoningBudget(effort); @@ -951,7 +958,7 @@ export function createAnthropicAdapter(provider: OcxProviderConfig, cacheRetenti // 400s ("max_tokens must be greater than thinking.budget_tokens"). Size them so max_tokens // always exceeds the budget within a model-safe ceiling, reserving room for visible output. const maxOut = parsed.options.maxOutputTokens ?? DEFAULT_MAX_TOKENS; - const wantBudget = reasoningBudget(parsed.options.reasoning); + const wantBudget = reasoningBudget(effectiveReasoning); const maxTokens = Math.min(REASONING_MAX_TOKENS_CEILING, Math.max(maxOut, wantBudget + OUTPUT_HEADROOM)); const budget = Math.max(MIN_THINKING_BUDGET, Math.min(wantBudget, maxTokens - OUTPUT_FLOOR)); body.max_tokens = maxTokens; diff --git a/tests/anthropic-reasoning.test.ts b/tests/anthropic-reasoning.test.ts index 251ed2393ab..93ff87d273f 100644 --- a/tests/anthropic-reasoning.test.ts +++ b/tests/anthropic-reasoning.test.ts @@ -40,6 +40,28 @@ describe("anthropic extended-thinking gate", () => { expect(b.top_p).toBe(0.8); }); + test("modelDefaultReasoningEfforts supplies reasoning when caller omits it", async () => { + const b = await bodyOf(parsed(undefined, { temperature: 0.5, topP: 0.8 }, "always-thinking-model"), { + ...provider, + modelDefaultReasoningEfforts: { "always-thinking-model": "high" }, + }); + const thinking = b.thinking as { type: string; budget_tokens: number } | undefined; + expect(thinking?.type).toBe("enabled"); + expect(typeof thinking?.budget_tokens).toBe("number"); + expect(b.temperature).toBeUndefined(); + expect(b.top_p).toBeUndefined(); + }); + + test("explicit reasoning overrides modelDefaultReasoningEfforts", async () => { + const b = await bodyOf(parsed("low", {}, "always-thinking-model"), { + ...provider, + modelDefaultReasoningEfforts: { "always-thinking-model": "high" }, + }); + const thinking = b.thinking as { type: string; budget_tokens: number } | undefined; + expect(thinking?.type).toBe("enabled"); + expect(thinking?.budget_tokens).toBe(4096); + }); + test("reasoning 'high' enables thinking and drops sampling (extended-thinking rule)", async () => { const b = await bodyOf(parsed("high", { temperature: 0.3, topP: 0.9 })); const thinking = b.thinking as { type: string; budget_tokens: number } | undefined; From b694268e8d533a115cc55f6620e9e37fc7733c6c Mon Sep 17 00:00:00 2001 From: JUN Date: Tue, 25 Aug 2026 12:28:56 +0900 Subject: [PATCH 12/29] fix(anthropic): do not let the __omit__ sentinel become an effort value (#2523) The provider default is read from modelDefaultReasoningEfforts, which can carry the __omit__ wire sentinel meaning 'send no reasoning field'. Treated as an ordinary effort it did the opposite of what it asks: adaptive models received output_config.effort: "__omit__" on the wire, and budget models had thinking turned ON with a default budget. Verified by probe before the fix, and the new regressions fail without it. --- src/adapters/anthropic.ts | 11 +++++++-- tests/anthropic-reasoning.test.ts | 40 +++++++++++++++++++++++++++++++ 2 files changed, 49 insertions(+), 2 deletions(-) diff --git a/src/adapters/anthropic.ts b/src/adapters/anthropic.ts index d8add210514..de86753b7c6 100644 --- a/src/adapters/anthropic.ts +++ b/src/adapters/anthropic.ts @@ -27,7 +27,7 @@ import { CLAUDE_CODE_HEADERS, claudeCodeSessionId } from "./client-fingerprint"; import { buildNonOpenAIToolCatalogNudgeForTools } from "./tool-catalog-nudge"; import { decodeServerSentEvents } from "../lib/sse-decoder"; import { isTranslatorBudgetExceededError, retainTranslatedEventBatch, type TranslatorBudget } from "../lib/translator-budget"; -import { modelRecordValue } from "../reasoning-effort"; +import { isReasoningEffortOmitted, modelRecordValue } from "../reasoning-effort"; /** Map a user content part to an Anthropic content block (text or image source). */ function toAnthropicContentPart(p: OcxContentPart): unknown { @@ -521,7 +521,14 @@ function adaptiveEffort(effort: string): string { function defaultReasoningEffort(provider: OcxProviderConfig, modelId: string): string | undefined { const value = modelRecordValue(provider.modelDefaultReasoningEfforts, modelId); - return typeof value === "string" && value.trim() ? value.trim() : undefined; + if (typeof value !== "string") return undefined; + const trimmed = value.trim(); + // `__omit__` means "send no reasoning field", not "an effort literally named + // __omit__". Without this the sentinel reached the wire as + // `output_config.effort: "__omit__"` on adaptive models, and enabled budget + // thinking on the rest — the opposite of what it asks for (#2432). + if (!trimmed || isReasoningEffortOmitted(trimmed)) return undefined; + return trimmed; } function usageFromAnthropic(usage: Record | undefined): OcxUsage | undefined { diff --git a/tests/anthropic-reasoning.test.ts b/tests/anthropic-reasoning.test.ts index 93ff87d273f..c6ba14c2395 100644 --- a/tests/anthropic-reasoning.test.ts +++ b/tests/anthropic-reasoning.test.ts @@ -487,3 +487,43 @@ describe("Claude Desktop classifier round trip (#545)", () => { expect(body.stop_sequences).toEqual([""]); }); }); + +describe("provider default reasoning effort (#2494)", () => { + const withDefault = (model: string, effort: string) => ({ + ...(provider as unknown as Record), + modelDefaultReasoningEfforts: { [model]: effort }, + } as unknown as OcxProviderConfig); + + test("a configured default applies when the caller omits reasoning", async () => { + const model = "anthropic/claude-sonnet-4.5"; + const b = await bodyOf(parsed(undefined, {}, model), withDefault(model, "high")); + expect((b.thinking as { type?: string } | undefined)?.type).toBe("enabled"); + expect((b.thinking as { budget_tokens?: number }).budget_tokens).toBe(16384); + }); + + test("an explicit caller effort still wins over the configured default", async () => { + const model = "anthropic/claude-sonnet-4.5"; + const b = await bodyOf(parsed("none", {}, model), withDefault(model, "high")); + expect(b.thinking).toBeUndefined(); + }); + + // The sentinel means "send no reasoning field". Treating it as an effort put + // output_config.effort: "__omit__" on the wire for adaptive models and turned + // budget thinking ON for the rest — the opposite of the request. + test("the __omit__ sentinel never becomes an effort value", async () => { + const adaptive = "anthropic/claude-fable-5"; + const a = await bodyOf(parsed(undefined, {}, adaptive), withDefault(adaptive, "__omit__")); + expect(a.output_config).toBeUndefined(); + expect(a.thinking).toBeUndefined(); + + const budget = "anthropic/claude-sonnet-4.5"; + const b = await bodyOf(parsed(undefined, {}, budget), withDefault(budget, "__omit__")); + expect(b.thinking).toBeUndefined(); + }); + + test("a blank default is ignored rather than treated as an effort", async () => { + const model = "anthropic/claude-sonnet-4.5"; + const b = await bodyOf(parsed(undefined, {}, model), withDefault(model, " ")); + expect(b.thinking).toBeUndefined(); + }); +}); From 224f23db2d95381302207a9ceda7f6295b7b1df3 Mon Sep 17 00:00:00 2001 From: liyongjie Date: Tue, 25 Aug 2026 11:30:27 +0800 Subject: [PATCH 13/29] fix: normalize legacy exec_command/shell_command tool calls to declared exec (#2493) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Codex 0.149 declares its code-mode shell tool as `exec` (a freeform custom tool whose description mentions the nested `await tools.exec_command(...)` helper). Routed models — DeepSeek in particular — sometimes echo that helper name as the tool-call name, emitting `exec_command` instead of the declared `exec`. The undeclared-tool guard then fails the whole turn with a 502. Normalize the legacy shell bridge names to `exec` at the three guard sites (streaming bridge x2, terminal snapshot guard) only when the request catalog declares `exec` and declares no legacy shell bridge name itself, so an MCP server advertising its own `exec_command` keeps working. Namespaced calls are always matched by their full wire name and never legacy-normalized. Co-authored-by: liyongjie.103 --- src/bridge.ts | 20 ++++++---- src/server/responses-undeclared-tool-guard.ts | 12 ++++-- src/types.ts | 1 + src/types/tools.ts | 27 +++++++++++++ tests/responses-undeclared-tool-guard.test.ts | 38 +++++++++++++++++++ 5 files changed, 86 insertions(+), 12 deletions(-) diff --git a/src/bridge.ts b/src/bridge.ts index b04bbb25ec9..bd24fb78e55 100644 --- a/src/bridge.ts +++ b/src/bridge.ts @@ -19,6 +19,7 @@ import { awaitThoughtSignatureDurability, } from "./responses/thought-signature-replay"; import { resolveStallTimeoutSec } from "./stall-timeout"; +import { normalizeDeclaredToolName } from "./types"; import { usageDisplayTotalTokens } from "./usage/totals"; import { appendSafeWebSearchSource, safeWebSearchSources } from "./web-search/sources"; import { @@ -1041,13 +1042,14 @@ export function bridgeToResponsesSSE( rememberReasoningForCall(event.id, rawReasoningForNextToolCall, replayCacheScope); } if (currentToolCall) closeCurrentToolCall(); - const mapped = toolNsMap?.get(event.name); - const realName = mapped?.name ?? event.name; - if (options?.declaredToolNames && !options.declaredToolNames.has(event.name)) { + const effectiveName = normalizeDeclaredToolName(event.name, options?.declaredToolNames); + const mapped = toolNsMap?.get(effectiveName); + const realName = mapped?.name ?? effectiveName; + if (options?.declaredToolNames && !options.declaredToolNames.has(effectiveName)) { const failure = responseError( 502, "upstream_error", - `routed provider emitted undeclared client tool "${event.name}"; only request-declared tools may be called`, + `routed provider emitted undeclared client tool "${effectiveName}"; only request-declared tools may be called`, ); emit("response.failed", { response: { @@ -1783,7 +1785,7 @@ function buildResponseJSONWithBudget( )); } break; - case "tool_call_start": + case "tool_call_start": { if (currentText) flushText("commentary"); if (currentSummaryReasoning) flushSummaryReasoning(); if (currentRawReasoning) flushRawReasoning(); @@ -1791,10 +1793,11 @@ function buildResponseJSONWithBudget( rememberReasoningForCall(e.id, rawReasoningForNextToolCall, replayCacheScope); } flushToolCall(); - if (options?.declaredToolNames && !options.declaredToolNames.has(e.name)) { + const effectiveName = normalizeDeclaredToolName(e.name, options?.declaredToolNames); + if (options?.declaredToolNames && !options.declaredToolNames.has(effectiveName)) { errorEvent = { type: "error", - message: `routed provider emitted undeclared client tool "${e.name}"; only request-declared tools may be called`, + message: `routed provider emitted undeclared client tool "${effectiveName}"; only request-declared tools may be called`, status: 502, errorType: "upstream_error", }; @@ -1802,11 +1805,12 @@ function buildResponseJSONWithBudget( } currentToolCallId = e.id; budget?.openCall(e.id); - currentToolCallName = e.name; + currentToolCallName = effectiveName; currentToolCallArgs = ""; currentToolCallArgsBytes = 0; currentToolCallProviderMetadata = e.providerMetadata; break; + } case "tool_call_delta": { ({ value: currentToolCallArgs, bytes: currentToolCallArgsBytes } = appendBatchString( diff --git a/src/server/responses-undeclared-tool-guard.ts b/src/server/responses-undeclared-tool-guard.ts index 158e6585b89..12c99fd2336 100644 --- a/src/server/responses-undeclared-tool-guard.ts +++ b/src/server/responses-undeclared-tool-guard.ts @@ -1,4 +1,4 @@ -import { namespacedToolName } from "../types"; +import { namespacedToolName, normalizeDeclaredToolName } from "../types"; import { sseDataPayload, type SseBlockRewrite } from "./sse-payload-rewrite"; /** Item types the client executes through a request-declared wire name. */ @@ -277,10 +277,14 @@ function undeclaredNameInItem( if (!CLIENT_EXECUTED_CALL_TYPES.has(item.type)) return undefined; const name = item.name; if (typeof name !== "string" || name.length === 0) return undefined; - if (declared.has(name)) return undefined; - if (typeof item.namespace === "string" && declared.has(namespacedToolName(item.namespace, name))) { - return undefined; + if (typeof item.namespace === "string") { + // Namespaced calls are matched by their full wire name only — never legacy-normalize + // them, or an undeclared namespaced `exec_command` could slip through as bare `exec`. + if (declared.has(namespacedToolName(item.namespace, name))) return undefined; + return name; } + const effectiveName = normalizeDeclaredToolName(name, declared); + if (declared.has(effectiveName)) return undefined; return name; } diff --git a/src/types.ts b/src/types.ts index 559e71cbc92..22a083098cb 100644 --- a/src/types.ts +++ b/src/types.ts @@ -4,6 +4,7 @@ export type { OcxTool, OcxToolChoice } from "./types/tools"; export { namespacedToolName, + normalizeDeclaredToolName, toolChoiceAliases, createToolChoiceResolver, toolChoiceCandidates, diff --git a/src/types/tools.ts b/src/types/tools.ts index 5e4f4547a08..89ebb3acb06 100644 --- a/src/types/tools.ts +++ b/src/types/tools.ts @@ -31,6 +31,33 @@ export function namespacedToolName(namespace: string | undefined, name: string): return namespace ? `${namespace}__${name}` : name; } +/** + * Codex 0.149 unified-exec name normalization. + * + * Codex's code-mode shell tool is declared as `exec` (a freeform custom tool whose own + * description mentions the nested `await tools.exec_command(...)` helper). Routed models — + * DeepSeek in particular — sometimes echo that helper name as the tool-call name, emitting + * `exec_command` instead of the declared `exec`. Accept the legacy shell bridge names only + * when the request catalog actually declares `exec` and does not itself declare the legacy + * name (an MCP server may legitimately advertise `exec_command` under its own namespace). + */ +const LEGACY_SHELL_BRIDGE_TOOL_NAMES = ["exec_command", "shell_command"] as const; + +export function normalizeDeclaredToolName( + name: string, + declared: ReadonlySet | undefined, +): string { + if (!declared || !declared.has("exec")) return name; + if (declared.has(name)) return name; + // When the catalog explicitly declares any legacy shell bridge name, the environment + // genuinely exposes that tool — turn normalization off so a call is never mis-routed + // to `exec`. + if ((LEGACY_SHELL_BRIDGE_TOOL_NAMES as readonly string[]).some(legacy => declared.has(legacy))) { + return name; + } + return (LEGACY_SHELL_BRIDGE_TOOL_NAMES as readonly string[]).includes(name) ? "exec" : name; +} + export function toolChoiceAliases(tool: Pick): string[] { const wireName = namespacedToolName(tool.namespace, tool.name); return tool.namespace ? [wireName, `${tool.namespace}.${tool.name}`] : [wireName]; diff --git a/tests/responses-undeclared-tool-guard.test.ts b/tests/responses-undeclared-tool-guard.test.ts index 1e93fb714dd..b27fa8e30c9 100644 --- a/tests/responses-undeclared-tool-guard.test.ts +++ b/tests/responses-undeclared-tool-guard.test.ts @@ -1249,6 +1249,44 @@ describe("undeclaredToolCallNameInResponse", () => { new Set(["computer_call"]), )).toBeUndefined(); }); + + test("accepts legacy shell bridge names when the catalog declares unified exec", () => { + // Codex 0.149 declares the code-mode shell tool as `exec`; routed models (DeepSeek) + // sometimes echo the nested helper name `exec_command` instead. The guard must accept + // it when the request catalog declares `exec` and does not itself declare the legacy + // name — but must still refuse it when the legacy name is a real declared tool. + const response = { + output: [ + { type: "function_call", name: "exec_command" }, + { type: "function_call", name: "shell_command" }, + ], + }; + + expect(undeclaredToolCallNameInResponse(response, new Set(["exec"]))).toBeUndefined(); + expect(undeclaredToolCallNameInResponse(response, new Set(["exec", "exec_command"]))).toBe( + "shell_command", + ); + expect(undeclaredToolCallNameInResponse(response, new Set(["exec_command"]))).toBe( + "shell_command", + ); + expect(undeclaredToolCallNameInResponse(response, new Set())).toBe("exec_command"); + }); + + test("never legacy-normalizes a namespaced shell bridge call", () => { + // A namespaced call (e.g. an MCP server advertising its own exec_command) must be + // matched by its full wire name only — never normalized to bare `exec`. + const namespaced = { + output: [ + { type: "function_call", name: "exec_command", namespace: "mcp__server" }, + { type: "function_call", name: "exec_command" }, + ], + }; + + expect(undeclaredToolCallNameInResponse(namespaced, new Set(["exec"]))).toBe( + "exec_command", + ); + expect(undeclaredToolCallNameInResponse(namespaced, new Set(["exec", "mcp__server__exec_command"]))).toBeUndefined(); + }); }); /** From 121c1fbe285d0a60ba46aed82b2ddaa5729db010 Mon Sep 17 00:00:00 2001 From: JUN Date: Tue, 25 Aug 2026 12:31:05 +0900 Subject: [PATCH 14/29] test(bridge): pin legacy shell-name normalization on the SSE path (#2524) #2493 fixed the 502 that bridgeToResponsesSSE emitted when a routed model echoed exec_command instead of the declared exec, but tested it through the guard helper rather than the bridge path where the failure occurs. Four cases on the bridge itself: both legacy names normalize, a genuinely undeclared tool still fails the turn, and a catalog that declares exec_command itself is never rewritten. Verified load-bearing: 2 of the 4 fail against dev before #2493 landed. --- .../bridge-legacy-shell-normalization.test.ts | 63 +++++++++++++++++++ 1 file changed, 63 insertions(+) create mode 100644 tests/bridge-legacy-shell-normalization.test.ts diff --git a/tests/bridge-legacy-shell-normalization.test.ts b/tests/bridge-legacy-shell-normalization.test.ts new file mode 100644 index 00000000000..76ce42e21a9 --- /dev/null +++ b/tests/bridge-legacy-shell-normalization.test.ts @@ -0,0 +1,63 @@ +import { describe, expect, test } from "bun:test"; +import { bridgeToResponsesSSE } from "../src/bridge"; +import type { AdapterEvent } from "../src/types"; + +async function drain(stream: ReadableStream): Promise { + const reader = stream.getReader(); + const decoder = new TextDecoder(); + let out = ""; + for (;;) { + const { done, value } = await reader.read(); + if (done) break; + out += decoder.decode(value, { stream: true }); + } + return out; +} + +async function* toolTurn(name: string): AsyncGenerator { + yield { type: "tool_call_start", id: "call-1", name } as AdapterEvent; + yield { type: "tool_call_delta", id: "call-1", delta: '{"cmd":"ls"}' } as AdapterEvent; + yield { type: "tool_call_end", id: "call-1" } as AdapterEvent; + yield { type: "done" } as AdapterEvent; +} + +// #2493: Codex 0.149 declares the shell tool as `exec`, whose own description names the +// nested `tools.exec_command(...)` helper. Routed models echo the helper name back, and the +// undeclared-tool guard turned that into a 502 mid-turn. These pin the SSE path the guard +// actually runs on, which the review flagged as untested. +describe("bridge normalizes legacy shell names against the declared catalog (#2493)", () => { + test("exec_command is delivered as the declared exec instead of failing the turn", async () => { + const sse = await drain(bridgeToResponsesSSE( + toolTurn("exec_command"), "deepseek-x", undefined, undefined, undefined, undefined, 50_000, + { declaredToolNames: new Set(["exec"]) }, + )); + expect(sse).not.toContain("undeclared client tool"); + expect(sse).toContain('"name":"exec"'); + }); + + test("shell_command normalizes the same way", async () => { + const sse = await drain(bridgeToResponsesSSE( + toolTurn("shell_command"), "deepseek-x", undefined, undefined, undefined, undefined, 50_000, + { declaredToolNames: new Set(["exec"]) }, + )); + expect(sse).not.toContain("undeclared client tool"); + expect(sse).toContain('"name":"exec"'); + }); + + test("a genuinely undeclared tool still fails the turn", async () => { + const sse = await drain(bridgeToResponsesSSE( + toolTurn("apply_patch"), "deepseek-x", undefined, undefined, undefined, undefined, 50_000, + { declaredToolNames: new Set(["exec"]) }, + )); + expect(sse).toContain("undeclared client tool"); + }); + + test("a catalog that declares exec_command itself is never rewritten", async () => { + const sse = await drain(bridgeToResponsesSSE( + toolTurn("exec_command"), "deepseek-x", undefined, undefined, undefined, undefined, 50_000, + { declaredToolNames: new Set(["exec", "exec_command"]) }, + )); + expect(sse).not.toContain("undeclared client tool"); + expect(sse).toContain('"name":"exec_command"'); + }); +}); From 64bc085c445b37fa6b1c98063a0829a519fd6ea3 Mon Sep 17 00:00:00 2001 From: JUN Date: Tue, 25 Aug 2026 12:48:40 +0900 Subject: [PATCH 15/29] fix(catalog): do not carry a retained compact limit onto a corrected window (#2526) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit #1905 taught catalog sync never to raise a compaction threshold retained from Codex. The rule is right, but the retained number was trusted even when sync corrected the row's context window in the same pass. An upstream entry arriving as 128k/115_200 whose window is then widened to 272k kept the stale 115_200 — 42% of the real window — so every long turn compacted early. CI caught it on macos and test 1/4 at 121c1fbe2. A retained threshold only describes the window it arrived with. Capture the incoming window before any override or cap rewrites the row, and trust the retained value only when the window is unchanged; lower-is-policy still holds there, which is what #1905 was protecting. --- src/codex/catalog/parsing.ts | 12 +++++++++++- 1 file changed, 11 insertions(+), 1 deletion(-) diff --git a/src/codex/catalog/parsing.ts b/src/codex/catalog/parsing.ts index 7fe07aca7f8..7d150567d9b 100644 --- a/src/codex/catalog/parsing.ts +++ b/src/codex/catalog/parsing.ts @@ -320,6 +320,9 @@ export function applyNativeOpenAiContextOverride(entry: RawEntry, limits?: Nativ ?? (isNativeOpenAiEntry(entry) ? entry.slug as string : undefined); if (!nativeSlug) return; const override = NATIVE_OPENAI_CONTEXT_OVERRIDES[nativeSlug]; + // Captured before any override/cap rewrites the row: a retained compaction threshold only + // describes the window it arrived with. + const incomingContextWindow = typeof entry.context_window === "number" ? entry.context_window : undefined; if (override) { // Read the effective values through the accessors rather than re-deriving them from the // static table: this function used to apply only the provider cap, so a per-model window @@ -352,7 +355,14 @@ export function applyNativeOpenAiContextOverride(entry: RawEntry, limits?: Nativ : undefined; if (effectiveContext !== undefined) { const derivedAutoCompactTokenLimit = nativeOpenAiAutoCompactTokenLimit(nativeSlug, limits); - const retainedAutoCompactTokenLimit = isNativeOpenAiEntry(entry) + // Only trust a retained threshold that still describes THIS window. When sync corrects the + // window, the old number is an artifact of the old one: a 115_200 limit retained from a + // 128k row would pin a corrected 272k model to 42% of its real window and compact every + // long turn early. Lower-is-policy still holds whenever the window is unchanged. + const retainedDescribesCurrentContext = incomingContextWindow === undefined + || incomingContextWindow === effectiveContext; + const retainedAutoCompactTokenLimit = retainedDescribesCurrentContext + && isNativeOpenAiEntry(entry) && typeof entry.auto_compact_token_limit === "number" && Number.isSafeInteger(entry.auto_compact_token_limit) && entry.auto_compact_token_limit > 0 From 8c21b69bb5a0d7ba96e893c40898930e4764d3f2 Mon Sep 17 00:00:00 2001 From: JUN Date: Tue, 25 Aug 2026 12:48:47 +0900 Subject: [PATCH 16/29] devlog: operator visibility train roadmap unit (260825) (#2520) Docs-only roadmap for three operator-visibility defects that share one shape: OpenCodex computes the truth and does not report it. - 010 (#2457): the sidecar pair check collapses a five-member union into a two-arm ternary, so a submitted gemini backend is validated against the stored openai backend and the dashboard save 400s. - 020 (#2411): collectStatus already computes routingKind and ships it in status --json, but the human renderer never prints it, so a healthy proxy reads green while nothing routes through it. - 030 (#2412): a version-manager overwrite leaves a message-less ineligible verdict, and the CLI warns only when a message exists. 002 records the plan audit, including two amendments: WP4 must plumb messages to every reachable silent ineligible return rather than only the one the reporter hit, and WP3 must pin custom-local/unknown as intentionally silent. From 98ed186c708966f92fcfef7c178678aeee43c31a Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?=EA=B9=80=EC=9C=A4=EA=B8=B0=20=28A=ED=8C=80=20=ED=94=84?= =?UTF-8?q?=EB=A1=A0=ED=8A=B8=29?= Date: Tue, 25 Aug 2026 13:07:52 +0900 Subject: [PATCH 17/29] fix(gui): inset the Claude account-pool warning and threshold field (#2492) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * fix(gui): inset the Claude account-pool warning and threshold field `.card` carries no padding of its own — `.card-row`, `.card-sub` and `.setting-row` each supply the 16px inset. The experimental warning box and the threshold field are none of those, so in the Anthropic account-pool card they rendered flush against the card border while the title row and the rotation rows sat 16px in. Measured in the dashboard at 814px card width: title 17px from the card edge, warning box and threshold input 1px. The warning takes margin rather than padding: it draws its own border, so padding insets only its text and leaves the border on the card edge. The threshold field takes padding because the `.input` inside is full-width, so the field owns the inset for the label, the control and the help line. This is the same fix `.account-pool-strategy-card` already carries for the Codex pool card, applied to the card that still had the gap. Moving the box styles out of inline `style` is part of it: an inline `padding` outranks any stylesheet rule, so leaving it there would keep the inset unreachable. * fix(gui): keep the inset rationale in one place The stylesheet rule, the test file header, and three per-assertion comments each restated the same two facts — that `.card` has no padding and that the warning box needs margin rather than padding. CodeRabbit's slop detector flagged the repetition, and it was right: the explanation belongs on the CSS rule, where the values it justifies live. The test header now points at that rule and keeps only what is specific to the test — why the assertions read source text rather than measuring, and the browser measurements the contract is derived from. The comments that survive per assertion are the ones the CSS does not carry: why the rule regex is anchored, how a shorthand is read, and why an inline style would silently disable the fix. --- .../AnthropicAccountPoolSettings.tsx | 16 +--- gui/src/styles.css | 27 +++++- gui/tests/anthropic-pool-card-layout.test.ts | 95 +++++++++++++++++++ 3 files changed, 124 insertions(+), 14 deletions(-) create mode 100644 gui/tests/anthropic-pool-card-layout.test.ts diff --git a/gui/src/components/provider-workspace/AnthropicAccountPoolSettings.tsx b/gui/src/components/provider-workspace/AnthropicAccountPoolSettings.tsx index 735660f2acd..d0ef91fab59 100644 --- a/gui/src/components/provider-workspace/AnthropicAccountPoolSettings.tsx +++ b/gui/src/components/provider-workspace/AnthropicAccountPoolSettings.tsx @@ -145,7 +145,7 @@ export default function AnthropicAccountPoolSettings({ const toggleDisabled = loading || saving || loadError || (!enabled && accountCount < 2); return ( -
+
{t("anthropicPool.title")} @@ -179,17 +179,7 @@ export default function AnthropicAccountPoolSettings({
-
+
{t("anthropicPool.experimentalWarning")}
@@ -199,7 +189,7 @@ export default function AnthropicAccountPoolSettings({ {enabled && state && ( <> -