fix(responses): count input tokens locally for Codex OAuth - #13167
diegosouzapw merged 1 commit into
Conversation
The ChatGPT subscription backend does not serve /backend-api/codex/responses/input_tokens for the affected account. Native requests to that path are intercepted by an OpenAI Cloudflare managed challenge, while the same path over the bundled Chrome transport returns 404 Not Found. Forwarding the client preflight can therefore never return a useful count and, before the companion classifier fix, permanently disabled the healthy Codex connection on the first challenge. Add a static /v1/responses/input_tokens route that shadows the generic Responses passthrough, uses the existing offline o200k_base token counter, and returns the standard response.input_tokens contract without issuing any upstream request. Count instructions, structured input, tool definitions and config; apply a conservative five-percent margin so the failure mode is earlier client compaction rather than a context-window overflow. Preserve the API-key and model-policy boundary from the catch-all Responses path. Tests pin the public schema, prove fetch is never called, cover text, instructions, structured input, tools, non-text parts, server-held context ids, invalid JSON, OPTIONS, and the conservative lower bound.
|
Great root-cause fix — shadowing the catch-all instead of trying to patch around the Cloudflare |
|
Thanks @anhtran-ai — merging via the release merge-train. Validated in local merge-train (merge-train-20260918-120911-suite.log) on the devbox @ train tip 8c305709478b7052fb981a75852bbf88fe3a959d with the sibling PRs of this batch: typecheck:core, file-size, complexity, cognitive-complexity, changelog-integrity green; changed-area node:test 310/311 (the single red, hard-lease inventory, reproduces on the pure release tip) + vitest green (fast parity mode — full suite ran today on the tip via the base-red and 3b trains). Merged --admin per merge-gates §4/§7. |
#13167 added the local Responses token-count route with hand-rolled `typeof` checks on `request.json()`. Hard Rule #7 and the t06 gate (scripts/check/check-route-validation.mjs, mirrored by tests/unit/route-body-validation-t06.test.ts) require every route that reads request.json() to go through validateBody()/safeParse(), so the tip was red. Add `v1ResponsesInputTokensSchema` (pins the wire types the counter reads — model/instructions strings, input string-or-array, tools array — and lets unknown keys through since they are counted, never forwarded) and run the body through validateBody(); a type mismatch is now a 400 naming the field instead of a silently ignored key. Regression test added to tests/unit/responses-input-tokens-local-route.test.ts. Refs #13866
…ll, TS2677, ESLint (Refs #13866) (#14331) * fix(build): ship httpClientAbortGuard.mjs in the pack artifact; validate input_tokens with Zod Wave five of the release/v3.8.51 base-reds, part 1 — the two that matter. #14064 restored server-ws.mjs's import of ./httpClientAbortGuard.mjs and the assembleStandalone copy, but not the two pack-artifact policy entries that were lost with it. Without APP_STAGING_ALLOWED_EXACT_PATHS the prepublish prune deletes the file; without PACK_ARTIFACT_REQUIRED_PATHS nothing notices. Every boot of the published package would die with ERR_MODULE_NOT_FOUND — the 3.8.47 head-response-guard class. Both closure suites (9/9) now enforce it. #13910's /v1/responses/input_tokens read request.json() behind a hand-rolled typeof check. Hard Rule #7 wants the boundary on Zod; the t06 guard caught it. Same passthrough envelope the catch-all Responses route uses, since the counters below already walk the fields defensively. 9/9 on the route's suite. Five no-unused-vars left behind by the wave (cliRuntime execFileSync, arena test symbols and a type, compression rmSync, waitForServer req) are removed. The 'openwa routes removed without deprecation' entry from the #14101 run was an artifact of that PR trailing its base — the gate is clean on the tip. Refs #13866 * fix(compression): let anchored file-pack rules see the transformed text; align wave-5 guards Wave five of the release/v3.8.51 base-reds, part 2. One production defect. #12825 (Hungarian Caveman pack) stopped gating file-pack rules with the English keyword list and tested the rule's own regex instead — against `lowerResult`, a lower-cased copy of the ORIGINAL text that the loop never refreshed. An anchored pattern like leader_phrases' `^(?:i will|…)` therefore ran its prefilter on "sure, i will…", failed the anchor, and was skipped; the rule that strips "I will " from every English response was dead since the merge. The prefilter now sees the text as the rules so far have left it. New test fails on the tip and passes here; all Caveman suites, Hungarian included, are 95/95. A frozen no-unused-vars suppression on caveman.ts no longer had a target and is pruned. Two more TS2677 predicates of the kind #14101 fixed: #13910 (rerankProviderNodes.ts, `n is RerankProviderNodeRow` on a Record row) and #13957's mitm catalog (antigravity.ts, `c is DynamicCatalogModel` on a literal-or-null). Both narrow by NonNullable of the element's own type; the api-route typecheck was 285 against a baseline of 283 on the pristine tip. The rest are guards trailing legitimate changes: - #12663 made gemini-3.8-flash the catalog head; T28 pinned 3.7. - #13863 put mimo-v2.5 into the shared vision heuristic on purpose (the base model is multimodal, only the Pro variants are text-only). The safety test now asserts the real invariant: base and :free aliases yes, -pro no. - #12565 moved npm-prefix detection into cliRuntimeNpmPrefix.ts with a process-lifetime cache that importFresh() does not reset; the case resets it. #12565 also builds Windows candidates with path.win32 on purpose; the qodercli test compared against POSIX path.join. - #13990 (the 2 GB Docker image) copies better-sqlite3 with --chown; the guard matched the flag order literally. Now flag-order tolerant, still fails when --from=builder is removed. - #13378 reintroduced public/openference.svg under a name #11750 retired for missing provenance and swapped the Cerebras showcase cell for it. The cell is back and the asset is gone; whether the new drawing counts as provenance is the owner's call. Refs #13866 * fix(i18n): translate the 7 sidebar-pin and Claude low-priority keys into all 65 locales #7f1b4a5e (sidebar pinned items) and #1b2349de (Claude OAuth lower-priority / auto-reset) landed with their 7 new keys in en.json only, which the vi and pt-BR parity suites flag. Translated with the repo's own sync-ui-keys --translate-markers against the .113 i18n instance (codex/gpt-5.6-sol-low): +446 lines across 65 catalogs, zero __MISSING__ markers, placeholders intact. vi.json also has two keys reordered to mirror en.json; values unchanged. Refs #13866 * chore(quality): list the 8 covering tests the sixth wave added in stryker tap.testFiles 30 commits landed on release/v3.8.51 while wave five drained; eight new unit tests cover mutated modules and were not in tap.testFiles, so their mutant kills did not count and check:mutation-test-coverage --strict failed on the merged tree. Appended at the end of the list, nothing reordered. Refs #13866 * docs: document the five env vars of the 09-18 wave; regenerate the version-manager skill for the open-wa routes check:docs-all: BRIDGE_PORT, ROUTER_URL and CERT_DIR (bin/antigravity-bridge.mjs, #c74cea3d), OPENWA_SERVICE_PORT (src/lib/services/bootstrap.ts, #1e8c913c) and NEXT_PUBLIC_PORT (src/shared/hooks/useDisplayBaseUrl.ts, #d715190b) were read in code but absent from .env.example and docs/reference/ENVIRONMENT.md. Added next to their neighbours, with the defaults the code actually uses (open-wa is 8323, not the 201xx range the other services sit in). check:agent-skills-sync: the open-wa feature added eight /api/services/openwa/* routes to docs/openapi.yaml without regenerating skills/omni-version-manager/ SKILL.md. Regenerated with the repo generator; the diff is exactly those eight route sections. Refs #13866 * test: register the crash guard in the pack snapshot; inventory #13874's refresh-lane row read pack-artifact-policy pins the list of root runtime files check:pack-artifact must find in the tarball; dist/httpClientAbortGuard.mjs joined PACK_ARTIFACT_REQUIRED_PATHS in this PR and the snapshot follows. #13874 re-reads the connection row inside the Claude refresh lane so a queued health check does not POST a refresh token a Layer 2 refresh already rotated — a state read, inventoried like the family-cooldown lookup (tokenHealthCheck.ts 2 -> 3). Refs #13866 * chore(quality): list native-codex-auto-resume test in stryker tap.testFiles (#13180 landed without it) * fix(release): drain the seventh base-red wave of release/v3.8.51 (9 tests + pack-policy + dashboard-typecheck) Three production defects the tests caught: - rateLimitManager: maxWaitMs=0 (the #12902 disable sentinel) hit #12715's queue-budget gate as "0 ms left" and 503'd every protected request. - emergencyFallback: #14006 silently switched the budget-exhaustion target provider nvidia -> groq against ENVIRONMENT.md and the NIM snapshot; restored. - claudeConnectionFields.ts vs ClaudeConnectionFields.tsx (#13074) differed only by casing; helpers renamed to claudeConnectionFieldValues.ts. Guards realigned to legitimate changes: #13874 rotation map (distinct token in the error test), #13350 origin-IP denylist, #13318 shared-catalog growth (counts by invariant), comboTargetKeyPolicy import in the telegram stub, the 22 README mirrors that #13940/#14106 stamped with the retired openference.svg (translated Cerebras cells recovered from history, hashes re-stamped), bin/antigravity-bridge.mjs allowed in the pack policy, and the two dashboard typecheck regressions (typed pinned section, ComponentProps cast). Refs #13866. * test: type the #13848 Gemini pairing tests (no-explicit-any) and inventory the semantic-cache embedding picker's connection read Both arrived with the tip merge: #13848 added 13 explicit any casts to translator-openai-to-gemini.test.ts (no-explicit-any is an error under tests/), and 7a92129's embeddingOptions.ts reads provider connections once without a hard-session-lease inventory entry. Stale suppression count pruned for the test file only. Refs #13866. * test: split the #13848 turn-pairing cases out of translator-openai-to-gemini.test.ts The file sits exactly at its frozen size cap; typing the pairing tests (no-explicit-any) pushed it 14 lines over. The two cases are a coherent regression suite of their own, so they move to translator-openai-to-gemini-turn-pairing-13848.test.ts (registered in stryker tap.testFiles) instead of widening the baseline. * docs(env): document BRIDGE_PORT, ROUTER_URL, CERT_DIR, OPENWA_SERVICE_PORT and NEXT_PUBLIC_PORT (Refs #13866) check:env-doc-sync has been red on the release tip since these five vars reached code without their .env.example / ENVIRONMENT.md entries: bin/antigravity-bridge.mjs (BRIDGE_PORT, ROUTER_URL, CERT_DIR — #14006), src/lib/services/bootstrap.ts + api/services/openwa/_lib.ts (OPENWA_SERVICE_PORT) and src/shared/hooks/useDisplayBaseUrl.ts (NEXT_PUBLIC_PORT — #13533). Defaults and source files copied from the reads themselves. * chore(skills): regenerate omni-version-manager for the open-wa service routes (Refs #13866) check:agent-skills-sync (Merge integrity job) has been red on the tip since the open-wa embedded-service routes reached docs/openapi.yaml without the generated SKILL.md being refreshed. Output of scripts/skills/generate-agent-skills.mjs --apply, no hand edits: the eight /api/services/openwa/* operations. * fix(types): make the two TS2677 type predicates sound (Refs #13866) check:api-typecheck has been red on the tip with two "type predicate's type must be assignable to its parameter's type" errors: - src/app/api/v1/_shared/rerankProviderNodes.ts (#13733): the read cache hands back `Record<string, unknown> | null`, and an interface whose members are all optional is not assignable to an index-signature type. Narrow to the non-null record and assert the row shape afterwards. - src/mitm/handlers/antigravity.ts (#14006): the map callback returned `{ displayName: string }` while DynamicCatalogModel declares it optional, so the predicate could not be proven. Type the callback's return explicitly and filter on `!== null`. No runtime change; rerank-remote-provider-nodes / rerank-local-node-shapes / mitm-handler-antigravity stay green. * fix(lint): clear the 92 ESLint errors the lint gate reports on the tip (Refs #13866) - tests/unit/translator-openai-to-gemini.test.ts: #13848 / #13318 added 13 `any` casts/params on top of the 74 frozen for the file, so ESLint reported all 87. Typed them (GeminiRequestWithContents / GeminiToolPart, and the existing GeminiRequestWithConfig) and pruned the file's suppression to the new count of 71 — nothing else in eslint-suppressions.json changes. - no-unused-vars: execFileSync import (src/shared/services/cliRuntime.ts, #12565), getArenaEloSyncStatus + makeLeaderboardMap + ArenaLeaderboardMap (tests/unit/arena-elo-sync-redesign.test.ts, #13446), rmSync (compressionAnalyticsWriterFlatRate.test.ts, #13446), `req` → `_req` (waitForServer-slow-first-response.test.mjs). translator-openai-to-gemini 48/48; arena-elo-sync-redesign, compressionAnalyticsWriterFlatRate, waitForServer-slow-first-response green. * fix(compression): stop skipping anchored Caveman rules that only match after earlier rules #12825 (Hungarian pack) replaced the English keyword prefilter with a `rule.pattern.test(lowerText)` pre-check for every file-based rule, including the default `en` pack. `lowerText` is the ORIGINAL message, so anchored rules such as `leader_phrases` (`^i will …`) — which only match after `pleasantries` strips "Sure, " — were dropped before they could run. `caveman-v379` caught the regression ("I will ensure …" survived at full intensity). Tag file-based rules with their pack language in ruleLoader and let the keyword prefilter apply to `en`/built-in rules only; non-English packs (which reuse English rule names) simply run their localized regex, which is what the pre-test cost anyway. Drops the now-unused CAVEMAN_RULES import and prunes the already-stale `caveman.ts` no-unused-vars suppression (0 violations on the tip) that blocked the pre-commit hook for any change to this file. Refs #13866 * test(models): align catalog and vision-heuristic guards with the tip's intended contracts Three base-reds where the production change was deliberate and the pinned guard was simply not bumped by the PR that changed the contract: - agy-antigravity-shared-catalog-12724: #13318 added the three Gemini 3.8 Flash tiers (high/medium/low, no "-tiered" endpoint for 3.8) to the shared Antigravity/AGY base, 10 -> 13. Pin the new size in one constant and make the buildSurfaceCatalog delta assertions relative to it. - t28-model-catalog-updates: #12663 (issue #12638) registered gemini-3.8-flash at the head of the AI Studio fallback catalog as the current Flash default; assert 3.8 first and keep 3.7 present. - command-code-mimo-v2-5-safety: #13863 (issue #13847) added an explicit "mimo-v2.5" fragment to the shared vision heuristic so provider-qualified and `-free` aliases keep their vision flag. The guard's real concern (the "mimo-vl" fragment must not cover "mimo-v2.5") is asserted on the fragment itself; the bare id is now vision by heuristic on purpose, and the Pro text-only sibling stays excluded. Refs #13866 * test(cli): follow the #12565 cliRuntime module split in the npm-prefix and qodercli guards #12565 (issue #12563) moved the npm global-prefix cache out of cliRuntime.ts into cliRuntimeNpmPrefix.ts and built the Windows known-bin candidates with `path.win32` (cliRuntimeWindowsNode.ts) so they stay Windows-shaped when `process.platform` is mocked on a POSIX runner. Two pre-existing guards depended on the old layout: - cli-runtime-extended "resolves known binaries from npm global prefix": importFresh() only re-evaluates cliRuntime.ts; the prefix cache now lives in a module that stays shared across cases, so a real `npm config get prefix` from an earlier case was cached and the mocked execFileSync never ran. Reset the cache with the helper #12565 exported for exactly this in afterEach. - qodercli-windows-resolve-6263: compare against `path.win32.join` — identical to `path.join` on a real Windows host, which is the behaviour under test. Production behaviour is unchanged on both platforms. Refs #13866 * test(auto-update): write the source-mode log inside the test's own temp dir The launchAutoUpdate case pointed AUTO_UPDATE_LOG_PATH at a fixed, world-shared `/tmp/auto-update-source.log`. On the .113 runner the suite executes both as `root` and as `runner` (uid 1001): the file survives owned by whoever ran first (`-rw-r--r-- root root`), and the next `openSync(logPath, "a")` fails with EACCES for the other user. Reproduced locally by making the shared file read-only; production code is untouched (autoUpdate.ts last changed in #9354). Use a per-test mkdtemp path for the source-mode log and clean the whole temp root in the existing finally block. Refs #13866 * fix(dashboard): rename claudeConnectionFields.ts so it no longer case-collides with ClaudeConnectionFields.tsx #13074 added two modules to the provider-detail modals directory whose names differ only by casing: `ClaudeConnectionFields.tsx` (the component) and `claudeConnectionFields.ts` (the value/patch helpers). On a case-insensitive filesystem the pair breaks the webpack build (#6584 guard), and esbuild's resolver already picks the `.tsx` for the extension-less `./claudeConnectionFields` specifier, so the provider-detail client entry failed to bundle ("No matching export ... for import claudeConnectionFieldPatch"). Rename the helper module to `claudeConnectionFieldValues.ts` (the same naming the sibling `quotaScrapingFieldValues.ts` uses) and point the only importer, EditConnectionModal.tsx, at the new name. Greens tests/unit/case-collision-6584.test.ts and tests/unit/media-page-client-browser-bundle.test.ts. Refs #13866 * fix(build): allowlist dist/httpClientAbortGuard.mjs so the published tarball keeps the server-ws crash guard #14064 (re-land of #13636) made scripts/dev/standalone-server-ws.mjs import ./httpClientAbortGuard.mjs and taught assembleStandalone to copy the shared implementation next to dist/server-ws.mjs — but never registered the file in scripts/build/pack-artifact-policy.ts. The prepublish prune deletes anything outside APP_STAGING_ALLOWED_EXACT_PATHS, and check:pack-artifact only fails on PACK_ARTIFACT_REQUIRED_PATHS entries, so the next `omniroute` tarball would boot straight into ERR_MODULE_NOT_FOUND (the #7065 / tls-options class the closure tests exist to catch). Add the bare and dist/ entries to both lists and extend the required-paths snapshot in tests/unit/pack-artifact-policy.test.ts. Greens tests/unit/pack-artifact-entrypoint-closures.test.ts and tests/unit/pack-artifact-server-ws-closure.test.ts. Refs #13866 * test(docker): accept --chown=node:node on the better-sqlite3 runner COPY #14010 deliberately changed the runner-stage COPYs to `COPY --chown=node:node --from=builder ...` (ownership at copy time instead of a second ~2 GB `chown -R` overlay layer). The Dockerfile contract test still matched the old `COPY --from=builder /app/node_modules/better-sqlite3` prefix and went red on the tip even though the native-addon guard it protects is intact. Tolerate the optional --chown flag; every other assertion (node-gyp rebuild, both `test -f .../better_sqlite3.node` checks) is unchanged. Refs #13866 * fix(api): validate /v1/responses/input_tokens bodies with Zod (t06) #13167 added the local Responses token-count route with hand-rolled `typeof` checks on `request.json()`. Hard Rule #7 and the t06 gate (scripts/check/check-route-validation.mjs, mirrored by tests/unit/route-body-validation-t06.test.ts) require every route that reads request.json() to go through validateBody()/safeParse(), so the tip was red. Add `v1ResponsesInputTokensSchema` (pins the wire types the counter reads — model/instructions strings, input string-or-array, tools array — and lets unknown keys through since they are counted, never forwarded) and run the body through validateBody(); a type mismatch is now a 400 naming the field instead of a silently ignored key. Regression test added to tests/unit/responses-input-tokens-local-route.test.ts. Refs #13866 * fix(docs): drop the retired openference.svg asset reintroduced by #13378 `openference.svg` is one of the 78 provider assets retired for missing provenance (tests/unit/provider-assets-generic-fallback.test.mjs freezes that list and forbids any tracked surface from referencing a retired name). #13378 added a new hand-drawn `public/openference.svg` outside the manifest-audited public/providers/ tree and pointed the README free-tier table (plus the 22 i18n mirrors that carry the row) at it, which put the retired name back on a tracked surface and left an unaudited asset in the package. Use the generic fallback icon (`public/providers/cli-generic.svg`) the other provenance-less providers already use, delete the unaudited file, and adopt the mechanical README edit into .i18n-state.json (`i18n:run -- --adopt --files=README.md`, no API calls) so the i18n drift gate does not flag README.md as source-changed. Refs #13866 * test(lease): classify the two connection-query sites added by #14159 and #13874 The hard-lease bypass inventory froze every getProviderConnections / getProviderConnectionById site with a class; two landed on the tip without a golden update: - src/app/api/settings/cache-config/embeddingOptions.ts (#14159, re-land of #12630): read-only listing that feeds the semantic-cache embedding dropdown, same shape as the qdrant embedding-models route — class C. - src/lib/tokenHealthCheck.ts 2 -> 3 (#13874): re-reads the row by id after an unrecoverable refresh error to detect credentials rotated by a concurrent Layer 2 refresh before deactivating — a state read, not dispatch; stays C. Refs #13866 * fix(sse): restore nvidia as the emergency budget-fallback provider #14006 (Antigravity MITM catalog injection) flipped EMERGENCY_FALLBACK_CONFIG.provider from "nvidia" to "groq" in one line, without touching ENVIRONMENT.md, .env.example, the chat.ts comment or the NVIDIA hosted-model snapshot, all of which still promise nvidia/openai/gpt-oss-120b. Operators without a Groq connection got the original 402 back instead of the free reroute, and chat-route-coverage ("uses the emergency fallback model on budget exhaustion" / "returns the primary budget error when emergency fallback also fails") went red on the tip. Put the documented default back; the #14006 bridge tests exercise bin/antigravity-bridge.mjs and do not read this config. Refs #13866 * fix(resilience): keep maxWaitMs=0 a "no queue deadline" sentinel #12902 released requestQueue.maxWaitMs=0 as the sentinel that disables the queue-wait deadline, but the #12715 queue-budget gate in withRateLimit() (`if (queueRemainingMs <= 0) throw`) read 0 as "budget spent" and rejected every request on a protected connection with an immediate 503 queue-budget error — the exact opposite of what the setting promises. rate-limit-maxwaitms-disable-execution ("400ms job completes without 504") was red on the tip. When no caller budget is passed and the configured queue budget is 0, skip the gate, never arm the queue-wait timer and hand awaitProviderDefaultSlot no budget (it falls back to the window). Execution stays bounded by executionMaxWaitMs and the upstream fetch-start timeout, as before. Refs #13866 * test: align three fixtures with the #13874, #13861 and #13350 contracts Three base-reds that are deliberate contract changes, not defects: - executor-default-base "refreshCredentials swallows refresh errors": #13874 records rotations on the Layer 2 (no connectionId) refresh path too, so the "refresh-me" token the previous case already rotated was served from the rotation map without the network POST the test wanted to fail. Use a token nobody rotated. - telegram-keycache-bounded-13165: #13861 made comboTargetKeyPolicy import isModelBlockedByPatterns from db/apiKeys; the loader-stubbed module lacked it and the suite died at module load. Export an honest "not blocked" stub (the test has no blocked models). - upstream-headers-proxy-auth "ordinary headers are still allowed": #13350 forbids the whole origin-IP forwarding set upstream (covered by upstream-headers-sanitize). Swap x-forwarded-for for x-request-id. Refs #13866
…apw#13167) The ChatGPT subscription backend does not serve /backend-api/codex/responses/input_tokens for the affected account. Native requests to that path are intercepted by an OpenAI Cloudflare managed challenge, while the same path over the bundled Chrome transport returns 404 Not Found. Forwarding the client preflight can therefore never return a useful count and, before the companion classifier fix, permanently disabled the healthy Codex connection on the first challenge. Add a static /v1/responses/input_tokens route that shadows the generic Responses passthrough, uses the existing offline o200k_base token counter, and returns the standard response.input_tokens contract without issuing any upstream request. Count instructions, structured input, tool definitions and config; apply a conservative five-percent margin so the failure mode is earlier client compaction rather than a context-window overflow. Preserve the API-key and model-policy boundary from the catch-all Responses path. Tests pin the public schema, prove fetch is never called, cover text, instructions, structured input, tools, non-text parts, server-held context ids, invalid JSON, OPTIONS, and the conservative lower bound. Co-authored-by: anhth2 <anhth2@vng.com.vn>
Summary
Add a static
POST /v1/responses/input_tokensroute that counts locally and returns the standard OpenAI Responses contract without issuing an upstream request.This complements #13161 / #13157. #13161 makes a Cloudflare managed challenge non-terminal; this PR removes the request that triggers that challenge in the first place.
Problem
The generic
/v1/responses/[...path]handler currently forwardsinput_tokensas a native Responses suffix. For a Codex OAuth / ChatGPT-subscription connection that produces:Measured behavior for the same live account and token:
403,cf-mitigated: challenge;wreq-jstransport: HTTP404 {"detail":"Not Found"};/backend-api/codex/responsesinference: succeeds.The subscription backend does not serve this OpenAI API subresource. Forwarding the preflight can therefore never produce a useful token count. It adds latency and fallback on every call and, before #13161, the managed-challenge 403 was persisted as terminal
banned/isActive:false, taking the healthy Codex provider offline.Fix
A static App Router segment shadows the catch-all:
It:
{ "object": "response.input_tokens", "input_tokens": <int> };countTextTokens()path (o200k_basefor Codex /cxmodels);tool_choice, text and reasoning configuration;fetch, so the Cloudflare challenge cannot occur on this path.The estimate is intentionally conservative. Live A/B showed a stable 9–10 token Responses envelope overhead for no-tool requests, while tool definitions already carry their own framing. The route adds that fixed overhead for no-tool requests and keeps a 5% margin for tool requests. Over-counting means slightly earlier client compaction; under-counting risks a context-window overflow.
Verification
Local
node --import tsx/esm --test tests/unit/responses-input-tokens-local-route.test.ts: 9/9 pass.npm run typecheck:core: clean.Tests pin the exact two-field response schema, prove
fetchis never called, and cover instructions, structured input, tools, non-text parts, server-held context IDs, invalid JSON, OPTIONS and the conservative lower bound.Production A/B
Deployed to the affected gateway as
llm-gateway/omniroute:3.8.51-113ddf9a-azox1and compared the local result to the subsequent real Codexusage.input_tokensfor eight request shapes:Results: 8/8 valid, 0 under-counts, median absolute error 0%. Both over-counts are in the safe direction (earlier compaction).
Additional production evidence:
proxy_logsby exactly 0 rows.is_active=1,test_status=activewith clear error fields.