Skip to content

feat(glm): add GLM-5.2 with effort-tier routing (high/max) - #3885

Merged
diegosouzapw merged 5 commits into
diegosouzapw:release/v3.8.26from
dhaern:feat/glm-5.2-reasoning-support
Jun 15, 2026
Merged

diegosouzapw merged 5 commits into
diegosouzapw:release/v3.8.26from
dhaern:feat/glm-5.2-reasoning-support

Conversation

@dhaern

@dhaern dhaern commented Jun 15, 2026

Copy link
Copy Markdown
Contributor

Summary

GLM-5.2 (released 2026-06-13) brings a 1M context window, 131072 max output tokens, and always-on thinking. While Zhipu's OpenAI-compatible endpoint only supports binary thinking (enabled/disabled), the Anthropic-compatible endpoint graduates reasoning intensity through Claude Code's effort selector (high/max), carried by the effort-2025-11-24 beta header already present in GLM_ANTHROPIC_BETA.

This PR adds three catalog entries that expose GLM-5.2 at full capacity:

Model Transport Reasoning Context Max Output
glm-5.2 OpenAI (coding/paas/v4) thinking always-on (binary) 1M 131072
glm-5.2-high Anthropic (direct, no fallback) effort: high 1M 131072
glm-5.2-max Anthropic (direct, no fallback) effort: max 1M 131072

The three entries are shared across glm, glm-cn, and glmt via GLM_SHARED_MODELS.

How it works

Effort tiers route directly through the Anthropic transport (no OpenAI fallback) because effort graduation is only supported on Zhipu's Anthropic endpoint:

glm-5.2-high/max
  → transformForTransport strips the suffix (upstream always sees "glm-5.2")
  → injects effort: "high"|"max" into the Anthropic body
  → injects thinking: { type: "enabled" } so the upstream emits thinking_delta blocks
  → GLM_ANTHROPIC_BETA carries the effort-2025-11-24 header

The thinking.type=enabled injection is required: without it, Zhipu's Anthropic endpoint does not surface thinking_delta blocks in the SSE stream, and clients (OpenCode, Claude Code, Cursor) see no reasoning content. The existing claude-to-openai translator maps thinking_delta → reasoning_content, so thinking text renders correctly in the client UI.

The base model (glm-5.2) retains the existing transport behavior (OpenAI primary + Anthropic fallback on retryable errors).

Files changed

File Change
open-sse/config/glmProvider.ts Add 3 models to GLM_SHARED_MODELS (1M ctx, 131K output)
open-sse/executors/glm.ts parseGlm52Effort() helper, force Anthropic transport for tiers, strip suffix, inject effort + thinking
src/shared/constants/modelSpecs.ts Model specs (1M context, 131K output, thinking + tools)
src/shared/constants/pricing.ts Pricing entries (mirror glm-5.1, same Coding Plan quota)
open-sse/mcp-server/__tests__/glmCodingProviderConfig.test.ts Catalog inventory, specs, pricing, tool-calling assertions

Design decisions

  • GLM_REQUEST_DEFAULTS untouched (maxTokens: 16384) — bumping it would break 4.x/5.x models with smaller output caps (e.g. glm-4.6v max 32768, glm-4.5 max 98304).
  • No changes to glmt — its GLMT_REQUEST_DEFAULTS are a separate preset; fixing the invalid thinkingType: "adaptive" (Zhipu only supports enabled/disabled) is out of scope for this PR.
  • Pricing mirrors glm-5.1 — GLM-5.2 consumes the same Coding Plan quota; public per-token pricing is not yet published (API opens "next week").
  • Suffix stripping happens in the executor, not in the global model resolver — keeps the effort tier visible for routing/quota while sending a clean upstream ID.

Validation

  • 9/9 focused tests pass (glmCodingProviderConfig.test.ts)
  • Runtime-verified all three tiers against the live Zhipu upstream:
    • glm-5.2 → OpenAI transport, reasoning_content in response ✓
    • glm-5.2-high → Anthropic transport, effort high, reasoning_content visible ✓
    • glm-5.2-max → Anthropic transport, effort max, reasoning_content visible ✓

Compatibility

  • No changes to existing GLM models or transports
  • No changes to global request defaults
  • New models are additive — clients not using them are unaffected

Refs

…pic transport

GLM-5.2 (released 2026-06-13) brings a 1M context window, 131072 max output
tokens and always-on thinking. While Zhipu's OpenAI-compatible endpoint only
supports binary thinking (enabled/disabled), the Anthropic-compatible endpoint
graduates reasoning intensity through Claude Code's effort selector
(high/max), carried by the effort-2025-11-24 beta header already present in
GLM_ANTHROPIC_BETA.

This change introduces three catalog entries:

- glm-5.2       → OpenAI transport (coding/paas/v4), thinking always-on
- glm-5.2-high  → Anthropic transport, effort: high, no fallback
- glm-5.2-max   → Anthropic transport, effort: max, no fallback

The effort tiers route directly through the Anthropic transport (no fallback)
because effort graduation is only supported there. The model suffix is
stripped before the upstream call so Zhipu always receives the base id
'glm-5.2'. thinking.type=enabled is injected into the Anthropic body so the
upstream emits thinking_delta blocks; these are translated back to
reasoning_content by the existing claude-to-openai translator, surfacing
thinking content to clients (OpenCode, Claude Code, Cursor, etc.).

GLM_REQUEST_DEFAULTS (maxTokens 16384) is left untouched to avoid breaking
existing 4.x/5.x models with smaller output caps. Pricing entries mirror
glm-5.1 (same Coding Plan quota). Model specs (modelSpecs.ts) declare the
1M context / 131K output for the discovery surface.

Refs: https://docs.z.ai/devpack/latest-model
@dhaern
dhaern requested a review from diegosouzapw as a code owner June 15, 2026 10:22

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request adds support for the GLM-5.2 model and its effort tiers (high and max), routing them through the Anthropic transport to map Claude Code effort selectors to Zhipu's reasoning intensity. It also updates model specifications, pricing, and provider configurations, alongside adding unit tests. Feedback highlights a critical issue where the current check fails to override the 'adaptive' thinking type sent by some clients, which is unsupported by Zhipu's Anthropic endpoint and would result in a 400 Bad Request; a fix is suggested to explicitly force the thinking type to 'enabled'.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment thread open-sse/executors/glm.ts Outdated
Comment on lines +284 to +286
if (!translatedRecord.thinking) {
translatedRecord.thinking = { type: "enabled" };
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Zhipu's Anthropic endpoint only supports enabled and disabled for the thinking configuration (it does not support adaptive). Since Claude Code and other clients often default to sending thinking: { type: "adaptive" } for reasoning models, the current check if (!translatedRecord.thinking) will evaluate to false and forward the unsupported adaptive type upstream, causing a 400 Bad Request or failing to return thinking. We should explicitly force the thinking.type to "enabled" whenever an effort tier is active, while preserving any other fields (like budget_tokens).

        if (!translatedRecord.thinking || asRecord(translatedRecord.thinking)?.type !== "enabled") {
          translatedRecord.thinking = {
            ...(asRecord(translatedRecord.thinking) ?? {}),
            type: "enabled",
          };
        }

Clients like Claude Code default to thinking.type=adaptive for reasoning
models, but Zhipu only supports enabled/disabled. The previous check
(!translatedRecord.thinking) passed adaptive through, causing 400s.
Force enabled while preserving other thinking fields (e.g. budget_tokens).

Fixes review feedback from @gemini-code-assist
@dhaern

dhaern commented Jun 15, 2026

Copy link
Copy Markdown
Contributor Author

Good catch, @gemini-code-assist.

Fixed in af8414d9. The effort-tier branch now forces thinking.type = "enabled" unconditionally (while preserving other fields like budget_tokens), instead of only injecting it when thinking was absent. This covers the Claude Code / client default of thinking: { type: "adaptive" }, which would otherwise be forwarded as-is and rejected by Zhipu with HTTP 400.

if (!translatedRecord.thinking || asRecord(translatedRecord.thinking)?.type !== "enabled") {
  translatedRecord.thinking = {
    ...(asRecord(translatedRecord.thinking) ?? {}),
    type: "enabled",
  };
}

9/9 tests still passing.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

The updated logic correctly handles the edge case where thinking is present but set to a non-enabled state. For a slight improvement in readability and to avoid redundant asRecord calls, you could consider this minor refactor:

const existingThinking = asRecord(translatedRecord.thinking);
if (!existingThinking || existingThinking.type !== "enabled") {
  translatedRecord.thinking = {
    ...existingThinking,
    type: "enabled",
  };
}

This maintains the same behavior while being slightly more efficient.

@dhaern

dhaern commented Jun 15, 2026

Copy link
Copy Markdown
Contributor Author

Applied in b8066cea — thanks @gemini-code-assist. existingThinking is now extracted once, keeping the same behavior while avoiding the redundant asRecord calls.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Glad to hear the refactor is working as expected. Let me know if you encounter any other issues during testing.

diegosouzapw and others added 2 commits June 15, 2026 12:24
PR diegosouzapw#3885 own growth: pricing.ts 1508->1529 (+21 = glm-5.2 + effort
aliases). Updates the frozen file-size baseline so Fast Quality Gates
pass on release/v3.8.26.

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>
@diegosouzapw

Copy link
Copy Markdown
Owner

Thanks, @dhaern! 🙏 Clean GLM-5.2 addition with the effort-tier routing done right — glm-5.2-high/-max route exclusively through the Anthropic transport, set effort, and force thinking.type=enabled (since Zhipu rejects adaptive). I verified GLM_ANTHROPIC_BETA already carries the effort-2025-11-24 beta header, so the intensity selector is wired end-to-end. Registry/pricing/specs tests green (9/9 vitest), lint clean. Added a file-size baseline bump for the +21 pricing rows. Merging into release/v3.8.26. 🚀

@diegosouzapw
diegosouzapw merged commit 624c846 into diegosouzapw:release/v3.8.26 Jun 15, 2026
1 check passed
@diegosouzapw diegosouzapw mentioned this pull request Jun 16, 2026
diegosouzapw added a commit that referenced this pull request Jun 16, 2026
OmniRoute v3.8.26 — see CHANGELOG.md [3.8.26] for the full notes.

Highlights: Vertex AI media generation (#3929), GLM-5.2 effort-tier routing (#3885),
sticky round-robin combos (#3846), OpenRouter connection presets (#3878), compression
prompt-cache fix (#3936/#3890), and a security pass (form-data/vite + workflow hardening, #3949).

Co-authored-by: artickc <artickc@users.noreply.github.com>
Co-authored-by: rdself <rdself@users.noreply.github.com>
Co-authored-by: herjarsa <herjarsa@users.noreply.github.com>
Co-authored-by: Jack Smith <16862258+YunyunZhai@users.noreply.github.com>
Co-authored-by: dhaern <dhaern@users.noreply.github.com>
Co-authored-by: adivekar-utexas <adivekar-utexas@users.noreply.github.com>
Co-authored-by: megamen32 <megamen32@users.noreply.github.com>
Co-authored-by: zhiru <zhiru@users.noreply.github.com>
Co-authored-by: insoln <insoln@users.noreply.github.com>
Co-authored-by: diego-anselmo <diego-anselmo@users.noreply.github.com>
HouMinXi pushed a commit to HouMinXi/OmniRoute that referenced this pull request Aug 2, 2026
OmniRoute v3.8.26 — see CHANGELOG.md [3.8.26] for the full notes.

Highlights: Vertex AI media generation (diegosouzapw#3929), GLM-5.2 effort-tier routing (diegosouzapw#3885),
sticky round-robin combos (diegosouzapw#3846), OpenRouter connection presets (diegosouzapw#3878), compression
prompt-cache fix (diegosouzapw#3936/diegosouzapw#3890), and a security pass (form-data/vite + workflow hardening, diegosouzapw#3949).

Co-authored-by: artickc <artickc@users.noreply.github.com>
Co-authored-by: rdself <rdself@users.noreply.github.com>
Co-authored-by: herjarsa <herjarsa@users.noreply.github.com>
Co-authored-by: Jack Smith <16862258+YunyunZhai@users.noreply.github.com>
Co-authored-by: dhaern <dhaern@users.noreply.github.com>
Co-authored-by: adivekar-utexas <adivekar-utexas@users.noreply.github.com>
Co-authored-by: megamen32 <megamen32@users.noreply.github.com>
Co-authored-by: zhiru <zhiru@users.noreply.github.com>
Co-authored-by: insoln <insoln@users.noreply.github.com>
Co-authored-by: diego-anselmo <diego-anselmo@users.noreply.github.com>
Poid-ZA pushed a commit to Poid-ZA/OmniRoute that referenced this pull request Aug 5, 2026
OmniRoute v3.8.26 — see CHANGELOG.md [3.8.26] for the full notes.

Highlights: Vertex AI media generation (diegosouzapw#3929), GLM-5.2 effort-tier routing (diegosouzapw#3885),
sticky round-robin combos (diegosouzapw#3846), OpenRouter connection presets (diegosouzapw#3878), compression
prompt-cache fix (diegosouzapw#3936/diegosouzapw#3890), and a security pass (form-data/vite + workflow hardening, diegosouzapw#3949).

Co-authored-by: artickc <artickc@users.noreply.github.com>
Co-authored-by: rdself <rdself@users.noreply.github.com>
Co-authored-by: herjarsa <herjarsa@users.noreply.github.com>
Co-authored-by: Jack Smith <16862258+YunyunZhai@users.noreply.github.com>
Co-authored-by: dhaern <dhaern@users.noreply.github.com>
Co-authored-by: adivekar-utexas <adivekar-utexas@users.noreply.github.com>
Co-authored-by: megamen32 <megamen32@users.noreply.github.com>
Co-authored-by: zhiru <zhiru@users.noreply.github.com>
Co-authored-by: insoln <insoln@users.noreply.github.com>
Co-authored-by: diego-anselmo <diego-anselmo@users.noreply.github.com>
tkgo11 pushed a commit to tkgo11/OmniRoute that referenced this pull request Sep 23, 2026
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
OmniRoute v3.8.26 — see CHANGELOG.md [3.8.26] for the full notes.

Highlights: Vertex AI media generation (diegosouzapw#3929), GLM-5.2 effort-tier routing (diegosouzapw#3885),
sticky round-robin combos (diegosouzapw#3846), OpenRouter connection presets (diegosouzapw#3878), compression
prompt-cache fix (diegosouzapw#3936/diegosouzapw#3890), and a security pass (form-data/vite + workflow hardening, diegosouzapw#3949).

Co-authored-by: artickc <artickc@users.noreply.github.com>
Co-authored-by: rdself <rdself@users.noreply.github.com>
Co-authored-by: herjarsa <herjarsa@users.noreply.github.com>
Co-authored-by: Jack Smith <16862258+YunyunZhai@users.noreply.github.com>
Co-authored-by: dhaern <dhaern@users.noreply.github.com>
Co-authored-by: adivekar-utexas <adivekar-utexas@users.noreply.github.com>
Co-authored-by: megamen32 <megamen32@users.noreply.github.com>
Co-authored-by: zhiru <zhiru@users.noreply.github.com>
Co-authored-by: insoln <insoln@users.noreply.github.com>
Co-authored-by: diego-anselmo <diego-anselmo@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants