Skip to content

fix(google): clamp max output tokens per model, without inventing a ceiling for unknown ids (rebase of #2512) - #2576

Merged
lidge-jun merged 3 commits into
devfrom
codex/google-output-clamp-2512
Aug 25, 2026
Merged

lidge-jun merged 3 commits into
devfrom
codex/google-output-clamp-2512

Conversation

@lidge-jun

@lidge-jun lidge-jun commented Aug 25, 2026 •

Copy link
Copy Markdown
Owner

Summary

Rebase of #2512 (by @Hsia97) onto current dev, with the review finding fixed.

The original change clamps Google-surface output tokens per model. The review objected to
two things, both real:

  1. An invented ceiling for unknown models. Anything unmatched fell back to 16,384, so
    an alias, a gateway-prefixed id, or any model newer than the table was silently
    truncated no matter what the operator requested. structure/02_config-and-codex-home.md
    is explicit that an explicit request value wins. Unknown ids now return undefined and
    the request passes through untouched — the upstream stays the authority on its own limit.
  2. Substring matching. includes("pro") matched my-prototype-model and
    includes("oss") matched crossover-v2. Matching is now prefix/family based.

The original test asserted the 16,384 fallback, which locked the defect in; it is replaced
with assertions for the passthrough contract and for the substring cases.

Closes #2512.

Verification

$ bun test tests/google-output-clamp.test.ts
 6 pass, 0 fail

$ bun test tests/google-errors.test.ts tests/google-output-clamp.test.ts
 16 pass, 0 fail, 109 expect() calls

$ bun x tsc --noEmit
(clean)

clampGoogleMaxOutputTokens has a single call site (src/adapters/google.ts:690), which
is unaffected by the widened return type.

Checklist

  • Targets dev
  • Rebased onto the current dev head (2/2, no conflicts)
  • Typecheck clean after the return-type change
  • Original authorship preserved in the commit history
  • No secrets or account identifiers in the diff

Summary by CodeRabbit

  • Bug Fixes
    • Google model requests now automatically respect documented output-token limits.
    • Valid token requests are preserved, while excessive values are reduced to the supported maximum.
    • Invalid or non-positive token values continue to be omitted.
    • Unknown model types retain their requested limits.

Hsia97 and others added 3 commits August 26, 2026 00:49
The clamp matched by substring and fell back to 16,384 for anything
unmatched, so an alias, a gateway id, or any model newer than the table
was silently truncated to 16,384 regardless of what the operator asked
for. structure/02_config-and-codex-home.md is explicit that an explicit
request value wins.

Unknown ids now return undefined and pass the request through untouched;
the upstream stays the authority on its own limit. Matching is also
prefix/family based, because includes(pro) matched my-prototype-model and
includes(oss) matched crossover-v2.
@lidge-jun
lidge-jun requested a review from Ingwannu as a code owner August 25, 2026 15:50
@lidge-jun
lidge-jun merged commit dd0f4af into dev Aug 25, 2026
7 of 8 checks passed
@lidge-jun
lidge-jun deleted the codex/google-output-clamp-2512 branch August 25, 2026 15:50
@github-actions

Copy link
Copy Markdown
Contributor

✅ Deterministic PR hygiene checks passed.

@coderabbitai

coderabbitai Bot commented Aug 25, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

Caution

Review failed

The pull request is closed.

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 359763af-ee76-4f70-84af-2f0ee48b6b37

📥 Commits

Reviewing files that changed from the base of the PR and between bfe2cb5 and c6d8edf.

📒 Files selected for processing (2)
  • src/adapters/google.ts
  • tests/google-output-clamp.test.ts

📝 Walkthrough

Walkthrough

Google model output-token requests now apply documented ceilings for supported model families. Unknown models preserve valid requested values. Invalid or non-positive values remain omitted. Tests cover model matching, clamping, passthrough, and invalid inputs.

Changes

Google output token clamping

Layer / File(s) Summary
Token ceiling helpers
src/adapters/google.ts, tests/google-output-clamp.test.ts
The adapter adds model ceiling lookup and positive-value clamping for Gemini, Claude, and GPT-OSS models. Tests cover exact family matching, supported ceilings, unknown models, valid requests, excessive requests, and invalid values.
Request clamping integration
src/adapters/google.ts
buildRequest uses clampGoogleMaxOutputTokens before assigning generationConfig.maxOutputTokens.

Estimated code review effort: 3 (Moderate) | ~20 minutes

Suggested reviewers: ingwannu

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch codex/google-output-clamp-2512

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added the bug Something isn't working label Aug 25, 2026

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: c6d8edf6fb

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/adapters/google.ts
// Pro tops out one token below the flash/other Gemini ceiling; both are documented values.
return /(^|[-.])pro([-.]|$)/.test(lower) ? 65535 : 65536;
}
if (lower.startsWith("claude")) return 64000;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Preserve per-model Claude output ceilings

When Google Antigravity live discovery selects claude-sonnet-4-6-thinking, the repository metadata at scripts/model-metadata.source.json:12511-12529 records a 128,000-token output limit, but this family-wide branch makes clampGoogleMaxOutputTokens(..., 100000) return 64,000. Before this change the requested 100,000 tokens reached upstream; now it is silently truncated, so read exact ceilings from the canonical provider metadata and pass through unmatched Claude IDs rather than treating every Claude model as 64,000.

AGENTS.md reference: src/AGENTS.md:L18-L18

Useful? React with 👍 / 👎.

tarunravi pushed a commit to tarunravi/opencodex that referenced this pull request Sep 14, 2026
…eiling for unknown ids (rebase of lidge-jun#2512) (lidge-jun#2576)

* fix(google): clamp max output tokens per model

* fix(google): clamp max output tokens per model

* fix(google): do not invent an output ceiling for unrecognized models

The clamp matched by substring and fell back to 16,384 for anything
unmatched, so an alias, a gateway id, or any model newer than the table
was silently truncated to 16,384 regardless of what the operator asked
for. structure/02_config-and-codex-home.md is explicit that an explicit
request value wins.

Unknown ids now return undefined and pass the request through untouched;
the upstream stays the authority on its own limit. Matching is also
prefix/family based, because includes(pro) matched my-prototype-model and
includes(oss) matched crossover-v2.

---------

Co-authored-by: Hsia97 <xjxj1997@163.com>
agentHits pushed a commit to agentHits/opencodex that referenced this pull request Sep 17, 2026
…eiling for unknown ids (rebase of lidge-jun#2512) (lidge-jun#2576)

* fix(google): clamp max output tokens per model

* fix(google): clamp max output tokens per model

* fix(google): do not invent an output ceiling for unrecognized models

The clamp matched by substring and fell back to 16,384 for anything
unmatched, so an alias, a gateway id, or any model newer than the table
was silently truncated to 16,384 regardless of what the operator asked
for. structure/02_config-and-codex-home.md is explicit that an explicit
request value wins.

Unknown ids now return undefined and pass the request through untouched;
the upstream stays the authority on its own limit. Matching is also
prefix/family based, because includes(pro) matched my-prototype-model and
includes(oss) matched crossover-v2.

---------

Co-authored-by: Hsia97 <xjxj1997@163.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants