feat(glm): add GLM-5.2 with effort-tier routing (high/max) - #3885
diegosouzapw merged 5 commits into
Conversation
…pic transport GLM-5.2 (released 2026-06-13) brings a 1M context window, 131072 max output tokens and always-on thinking. While Zhipu's OpenAI-compatible endpoint only supports binary thinking (enabled/disabled), the Anthropic-compatible endpoint graduates reasoning intensity through Claude Code's effort selector (high/max), carried by the effort-2025-11-24 beta header already present in GLM_ANTHROPIC_BETA. This change introduces three catalog entries: - glm-5.2 → OpenAI transport (coding/paas/v4), thinking always-on - glm-5.2-high → Anthropic transport, effort: high, no fallback - glm-5.2-max → Anthropic transport, effort: max, no fallback The effort tiers route directly through the Anthropic transport (no fallback) because effort graduation is only supported there. The model suffix is stripped before the upstream call so Zhipu always receives the base id 'glm-5.2'. thinking.type=enabled is injected into the Anthropic body so the upstream emits thinking_delta blocks; these are translated back to reasoning_content by the existing claude-to-openai translator, surfacing thinking content to clients (OpenCode, Claude Code, Cursor, etc.). GLM_REQUEST_DEFAULTS (maxTokens 16384) is left untouched to avoid breaking existing 4.x/5.x models with smaller output caps. Pricing entries mirror glm-5.1 (same Coding Plan quota). Model specs (modelSpecs.ts) declare the 1M context / 131K output for the discovery surface. Refs: https://docs.z.ai/devpack/latest-model
There was a problem hiding this comment.
Code Review
This pull request adds support for the GLM-5.2 model and its effort tiers (high and max), routing them through the Anthropic transport to map Claude Code effort selectors to Zhipu's reasoning intensity. It also updates model specifications, pricing, and provider configurations, alongside adding unit tests. Feedback highlights a critical issue where the current check fails to override the 'adaptive' thinking type sent by some clients, which is unsupported by Zhipu's Anthropic endpoint and would result in a 400 Bad Request; a fix is suggested to explicitly force the thinking type to 'enabled'.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
| if (!translatedRecord.thinking) { | ||
| translatedRecord.thinking = { type: "enabled" }; | ||
| } |
There was a problem hiding this comment.
Zhipu's Anthropic endpoint only supports enabled and disabled for the thinking configuration (it does not support adaptive). Since Claude Code and other clients often default to sending thinking: { type: "adaptive" } for reasoning models, the current check if (!translatedRecord.thinking) will evaluate to false and forward the unsupported adaptive type upstream, causing a 400 Bad Request or failing to return thinking. We should explicitly force the thinking.type to "enabled" whenever an effort tier is active, while preserving any other fields (like budget_tokens).
if (!translatedRecord.thinking || asRecord(translatedRecord.thinking)?.type !== "enabled") {
translatedRecord.thinking = {
...(asRecord(translatedRecord.thinking) ?? {}),
type: "enabled",
};
}Clients like Claude Code default to thinking.type=adaptive for reasoning models, but Zhipu only supports enabled/disabled. The previous check (!translatedRecord.thinking) passed adaptive through, causing 400s. Force enabled while preserving other thinking fields (e.g. budget_tokens). Fixes review feedback from @gemini-code-assist
|
Good catch, @gemini-code-assist. Fixed in if (!translatedRecord.thinking || asRecord(translatedRecord.thinking)?.type !== "enabled") {
translatedRecord.thinking = {
...(asRecord(translatedRecord.thinking) ?? {}),
type: "enabled",
};
}9/9 tests still passing. |
|
The updated logic correctly handles the edge case where const existingThinking = asRecord(translatedRecord.thinking);
if (!existingThinking || existingThinking.type !== "enabled") {
translatedRecord.thinking = {
...existingThinking,
type: "enabled",
};
}This maintains the same behavior while being slightly more efficient. |
…alls Addresses readability feedback from @gemini-code-assist
|
Applied in |
|
Glad to hear the refactor is working as expected. Let me know if you encounter any other issues during testing. |
….2-reasoning-support
PR diegosouzapw#3885 own growth: pricing.ts 1508->1529 (+21 = glm-5.2 + effort aliases). Updates the frozen file-size baseline so Fast Quality Gates pass on release/v3.8.26. Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>
|
Thanks, @dhaern! 🙏 Clean GLM-5.2 addition with the effort-tier routing done right — |
OmniRoute v3.8.26 — see CHANGELOG.md [3.8.26] for the full notes. Highlights: Vertex AI media generation (#3929), GLM-5.2 effort-tier routing (#3885), sticky round-robin combos (#3846), OpenRouter connection presets (#3878), compression prompt-cache fix (#3936/#3890), and a security pass (form-data/vite + workflow hardening, #3949). Co-authored-by: artickc <artickc@users.noreply.github.com> Co-authored-by: rdself <rdself@users.noreply.github.com> Co-authored-by: herjarsa <herjarsa@users.noreply.github.com> Co-authored-by: Jack Smith <16862258+YunyunZhai@users.noreply.github.com> Co-authored-by: dhaern <dhaern@users.noreply.github.com> Co-authored-by: adivekar-utexas <adivekar-utexas@users.noreply.github.com> Co-authored-by: megamen32 <megamen32@users.noreply.github.com> Co-authored-by: zhiru <zhiru@users.noreply.github.com> Co-authored-by: insoln <insoln@users.noreply.github.com> Co-authored-by: diego-anselmo <diego-anselmo@users.noreply.github.com>
OmniRoute v3.8.26 — see CHANGELOG.md [3.8.26] for the full notes. Highlights: Vertex AI media generation (diegosouzapw#3929), GLM-5.2 effort-tier routing (diegosouzapw#3885), sticky round-robin combos (diegosouzapw#3846), OpenRouter connection presets (diegosouzapw#3878), compression prompt-cache fix (diegosouzapw#3936/diegosouzapw#3890), and a security pass (form-data/vite + workflow hardening, diegosouzapw#3949). Co-authored-by: artickc <artickc@users.noreply.github.com> Co-authored-by: rdself <rdself@users.noreply.github.com> Co-authored-by: herjarsa <herjarsa@users.noreply.github.com> Co-authored-by: Jack Smith <16862258+YunyunZhai@users.noreply.github.com> Co-authored-by: dhaern <dhaern@users.noreply.github.com> Co-authored-by: adivekar-utexas <adivekar-utexas@users.noreply.github.com> Co-authored-by: megamen32 <megamen32@users.noreply.github.com> Co-authored-by: zhiru <zhiru@users.noreply.github.com> Co-authored-by: insoln <insoln@users.noreply.github.com> Co-authored-by: diego-anselmo <diego-anselmo@users.noreply.github.com>
OmniRoute v3.8.26 — see CHANGELOG.md [3.8.26] for the full notes. Highlights: Vertex AI media generation (diegosouzapw#3929), GLM-5.2 effort-tier routing (diegosouzapw#3885), sticky round-robin combos (diegosouzapw#3846), OpenRouter connection presets (diegosouzapw#3878), compression prompt-cache fix (diegosouzapw#3936/diegosouzapw#3890), and a security pass (form-data/vite + workflow hardening, diegosouzapw#3949). Co-authored-by: artickc <artickc@users.noreply.github.com> Co-authored-by: rdself <rdself@users.noreply.github.com> Co-authored-by: herjarsa <herjarsa@users.noreply.github.com> Co-authored-by: Jack Smith <16862258+YunyunZhai@users.noreply.github.com> Co-authored-by: dhaern <dhaern@users.noreply.github.com> Co-authored-by: adivekar-utexas <adivekar-utexas@users.noreply.github.com> Co-authored-by: megamen32 <megamen32@users.noreply.github.com> Co-authored-by: zhiru <zhiru@users.noreply.github.com> Co-authored-by: insoln <insoln@users.noreply.github.com> Co-authored-by: diego-anselmo <diego-anselmo@users.noreply.github.com>
…apw#3885) Integrated into release/v3.8.26
OmniRoute v3.8.26 — see CHANGELOG.md [3.8.26] for the full notes. Highlights: Vertex AI media generation (diegosouzapw#3929), GLM-5.2 effort-tier routing (diegosouzapw#3885), sticky round-robin combos (diegosouzapw#3846), OpenRouter connection presets (diegosouzapw#3878), compression prompt-cache fix (diegosouzapw#3936/diegosouzapw#3890), and a security pass (form-data/vite + workflow hardening, diegosouzapw#3949). Co-authored-by: artickc <artickc@users.noreply.github.com> Co-authored-by: rdself <rdself@users.noreply.github.com> Co-authored-by: herjarsa <herjarsa@users.noreply.github.com> Co-authored-by: Jack Smith <16862258+YunyunZhai@users.noreply.github.com> Co-authored-by: dhaern <dhaern@users.noreply.github.com> Co-authored-by: adivekar-utexas <adivekar-utexas@users.noreply.github.com> Co-authored-by: megamen32 <megamen32@users.noreply.github.com> Co-authored-by: zhiru <zhiru@users.noreply.github.com> Co-authored-by: insoln <insoln@users.noreply.github.com> Co-authored-by: diego-anselmo <diego-anselmo@users.noreply.github.com>
Summary
GLM-5.2 (released 2026-06-13) brings a 1M context window, 131072 max output tokens, and always-on thinking. While Zhipu's OpenAI-compatible endpoint only supports binary thinking (
enabled/disabled), the Anthropic-compatible endpoint graduates reasoning intensity through Claude Code's effort selector (high/max), carried by theeffort-2025-11-24beta header already present inGLM_ANTHROPIC_BETA.This PR adds three catalog entries that expose GLM-5.2 at full capacity:
glm-5.2coding/paas/v4)glm-5.2-highglm-5.2-maxThe three entries are shared across
glm,glm-cn, andglmtviaGLM_SHARED_MODELS.How it works
Effort tiers route directly through the Anthropic transport (no OpenAI fallback) because effort graduation is only supported on Zhipu's Anthropic endpoint:
The
thinking.type=enabledinjection is required: without it, Zhipu's Anthropic endpoint does not surfacethinking_deltablocks in the SSE stream, and clients (OpenCode, Claude Code, Cursor) see no reasoning content. The existingclaude-to-openaitranslator mapsthinking_delta→reasoning_content, so thinking text renders correctly in the client UI.The base model (
glm-5.2) retains the existing transport behavior (OpenAI primary + Anthropic fallback on retryable errors).Files changed
open-sse/config/glmProvider.tsGLM_SHARED_MODELS(1M ctx, 131K output)open-sse/executors/glm.tsparseGlm52Effort()helper, force Anthropic transport for tiers, strip suffix, inject effort + thinkingsrc/shared/constants/modelSpecs.tssrc/shared/constants/pricing.tsopen-sse/mcp-server/__tests__/glmCodingProviderConfig.test.tsDesign decisions
GLM_REQUEST_DEFAULTSuntouched (maxTokens: 16384) — bumping it would break 4.x/5.x models with smaller output caps (e.g.glm-4.6vmax 32768,glm-4.5max 98304).glmt— itsGLMT_REQUEST_DEFAULTSare a separate preset; fixing the invalidthinkingType: "adaptive"(Zhipu only supportsenabled/disabled) is out of scope for this PR.glm-5.1— GLM-5.2 consumes the same Coding Plan quota; public per-token pricing is not yet published (API opens "next week").Validation
glmCodingProviderConfig.test.ts)glm-5.2→ OpenAI transport,reasoning_contentin response ✓glm-5.2-high→ Anthropic transport, effort high,reasoning_contentvisible ✓glm-5.2-max→ Anthropic transport, effort max,reasoning_contentvisible ✓Compatibility
Refs