Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -192,7 +192,12 @@ ANTHROPIC_API_KEY=sk-ant-your-key-here
# For Z.AI GLM Coding Plan, set:
# OPENAI_BASE_URL=https://api.z.ai/api/coding/paas/v4
# OPENAI_MODEL=glm-5.2
# Optional: OPENAI_MODEL=glm-5.3 (the default remains glm-5.2)
# Optional: OPENAI_MODEL=GLM-5.1, GLM-5-Turbo, GLM-4.7, or GLM-4.5-Air
# Optional GLM-5.3 thinking controls:
# OPENAI_MODEL='glm-5.3?reasoning=low' # requests Z.AI reasoning_effort=low
# OPENAI_MODEL='glm-5.3?reasoning=high' # requests Z.AI reasoning_effort=high
# OPENAI_MODEL='glm-5.3?reasoning=xhigh' # maps to Z.AI reasoning_effort=max
# Optional GLM-5.2 thinking controls:
# OPENAI_MODEL='glm-5.2?reasoning=high' # enhanced reasoning
# OPENAI_MODEL='glm-5.2?reasoning=xhigh' # maps to Z.AI reasoning_effort=max
Expand Down
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -343,7 +343,7 @@ OpenClaude supports multiple providers, but behavior is not identical across all
- Some providers impose lower output caps than the CLI defaults, and OpenClaude adapts where possible
- AI/ML API uses the OpenAI-compatible route, defaults to `gpt-4o`, and only surfaces chat-capable models from its public catalog
- Gitlawb Opengateway is the fresh-install startup default and requires an API key from https://gitlawb.com/opengateway/keys. It uses one OpenAI-compatible base URL; switch between `mimo-*` and `google/gemini-3.1-flash-lite-preview` with `/model`, and do not pin the base URL to `/v1/xiaomi-mimo`.
- Z.AI GLM Coding Plan uses `https://api.z.ai/api/coding/paas/v4` with `glm-5.2` by default. Use `glm-5.2?reasoning=high` for enhanced reasoning, `glm-5.2?reasoning=xhigh` to request Z.AI `reasoning_effort=max`, or `glm-5.2?thinking=disabled` for faster direct answers.
- Z.AI GLM Coding Plan uses `https://api.z.ai/api/coding/paas/v4` with `glm-5.2` by default. GLM-5.3 is selectable as `glm-5.3`; use `glm-5.3?reasoning=low`, `glm-5.3?reasoning=high`, or `glm-5.3?reasoning=xhigh` to request its documented low, high, or maximum effort. The existing GLM-5.2 query controls remain supported.
- Xiaomi MiMo uses `api-key` header auth on the direct OpenAI-compatible route and currently does not support `/usage` reporting in OpenClaude
- GitHub Copilot serializes sub-agent execution by default to reduce Premium Request consumption — see [Agent Routing and Step Limits](docs/agent-routing.md#github-copilot-sub-agent-optimization) for tuning

Expand Down
15 changes: 14 additions & 1 deletion docs/integrations/how-to/add-model.md
Original file line number Diff line number Diff line change
Expand Up @@ -31,7 +31,10 @@ still the source of truth for where a model is offered.
capabilities.
4. Add optional shared metadata.
Include `brandId`, `contextWindow`, `maxOutputTokens`, and `cacheConfig`
when the data is stable enough to be reused.
when the data is stable enough to be reused. Set
`runtimeMetadataScope: 'catalog'` when verified limits and capabilities
should apply only on route catalogs that explicitly reference the
descriptor.
5. Add `providerModelMap` only when the same model needs route-specific API
names across multiple catalogs.
6. Update route-owned catalogs only if the model should be offered by those
Expand All @@ -47,6 +50,11 @@ Model descriptor files should:
- avoid encoding gateway availability as if every route automatically exposes
the shared model.

Shared runtime metadata uses the legacy global model-name fallback by default.
Use `runtimeMetadataScope: 'catalog'` for a model whose verified limits and
capabilities belong to specific routes; that metadata then applies only when a
route catalog entry names the descriptor through `modelDescriptorId`.

Normal contributor-facing examples should not call `registerModel(...)`
directly.

Expand Down Expand Up @@ -190,6 +198,11 @@ Model lookup should prefer:
second built-in model table. Built-in model limits belong in model descriptor
files.

A descriptor with `runtimeMetadataScope: 'catalog'` is intentionally excluded
from global name-only lookups. Its limits and capabilities are available only
through an explicit route catalog entry, preventing one vendor's verified
contract from leaking onto an uncataloged gateway model with the same API name.

## What not to do

Avoid these patterns:
Expand Down
1 change: 1 addition & 0 deletions src/integrations/brands/glm.ts
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,7 @@ export default defineBrand({
supportsPreciseTokenCount: false,
},
modelIds: [
'glm-5.3',
'glm-5.2',
'GLM-5.1',
'GLM-5-Turbo',
Expand Down
5 changes: 5 additions & 0 deletions src/integrations/descriptors.ts
Original file line number Diff line number Diff line change
Expand Up @@ -355,6 +355,11 @@ export interface ModelDescriptor {
reasoning?: ReasoningControlMetadata
contextWindow?: number
maxOutputTokens?: number
/**
* Restrict shared runtime metadata to catalog entries that explicitly
* reference this descriptor. Omit for the legacy global model-name fallback.
*/
runtimeMetadataScope?: 'global' | 'catalog'
cacheConfig?: CacheConfig
}

Expand Down
4 changes: 4 additions & 0 deletions src/integrations/models/glm.ts
Original file line number Diff line number Diff line change
@@ -1,4 +1,5 @@
import { defineModel } from '../define.js'
import type { ModelDescriptor } from '../descriptors.js'

const glmCapabilities = {
supportsVision: false,
Expand All @@ -14,6 +15,7 @@ function glmModel(
label: string,
contextWindow: number,
maxOutputTokens: number,
runtimeMetadataScope?: ModelDescriptor['runtimeMetadataScope'],
) {
return defineModel({
id,
Expand All @@ -25,10 +27,12 @@ function glmModel(
capabilities: glmCapabilities,
contextWindow,
maxOutputTokens,
...(runtimeMetadataScope ? { runtimeMetadataScope } : {}),
})
}

export default [
glmModel('glm-5.3', 'GLM 5.3', 1_000_000, 131_072, 'catalog'),
defineModel({
id: 'glm-5v-turbo',
label: 'GLM 5V Turbo',
Expand Down
106 changes: 105 additions & 1 deletion src/integrations/runtimeMetadata.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -17,9 +17,55 @@ import {
getRouteDiscoveryHeaders,
} from './discoveryService'
import { setClaudeConfigHomeDirForTesting } from '../utils/envUtils.js'
import glmBrand from './brands/glm.js'
import glmModels from './models/glm.js'
import zaiVendor from './vendors/zai.js'

const originalConfigDir = process.env.CLAUDE_CONFIG_DIR

describe('Z.AI GLM-5.3 descriptor contract', () => {
it('wires the verified shared model, brand, and direct catalog entry without changing the default', () => {
const model = glmModels.find(candidate => candidate.id === 'glm-5.3')
expect(model).toMatchObject({
id: 'glm-5.3',
label: 'GLM 5.3',
vendorId: 'zai',
brandId: 'glm',
classification: ['chat', 'reasoning', 'coding'],
defaultModel: 'glm-5.3',
contextWindow: 1_000_000,
maxOutputTokens: 131_072,
runtimeMetadataScope: 'catalog',
capabilities: {
supportsVision: false,
supportsStreaming: true,
supportsFunctionCalling: true,
supportsJsonMode: true,
supportsReasoning: true,
supportsPreciseTokenCount: false,
},
})
expect(glmBrand.modelIds?.[0]).toBe('glm-5.3')

const catalogEntry = zaiVendor.catalog?.models?.[0]
expect(catalogEntry).toMatchObject({
id: 'glm-5.3',
apiName: 'glm-5.3',
label: 'GLM-5.3',
modelDescriptorId: 'glm-5.3',
reasoning: {
mode: 'levels',
levels: ['low', 'high', 'xhigh'],
wireFormat: 'zai_compatible',
},
transportOverrides: {
openaiShim: { enableToolStreaming: true },
},
})
expect(zaiVendor.defaultModel).toBe('glm-5.2')
})
})

async function withTempConfigDir<T>(fn: () => Promise<T>): Promise<T> {
await acquireSharedMutationLock('integrations/runtimeMetadata.test.ts')
let tempDir: string | null = null
Expand Down Expand Up @@ -120,7 +166,25 @@ describe('resolveModelRuntimeLimits', () => {
}
})
})
it('uses built-in Z.AI GLM-5.2 runtime limits', () => {
it.each([
'glm-5.3',
'glm-5.3?reasoning=low',
'glm-5.3?reasoning=xhigh',
'glm-5.3?thinking=disabled',
])('uses verified Z.AI GLM-5.3 runtime limits for %s', model => {
const limits = resolveModelRuntimeLimits({
model,
processEnv: {
CLAUDE_CODE_USE_OPENAI: '1',
OPENAI_BASE_URL: 'https://api.z.ai/api/coding/paas/v4',
},
})

expect(limits.contextWindow).toBe(1_000_000)
expect(limits.maxOutputTokens).toBe(131_072)
})

it('keeps the built-in Z.AI GLM-5.2 runtime limits', () => {
const limits = resolveModelRuntimeLimits({
model: 'glm-5.2',
processEnv: {
Expand All @@ -131,6 +195,23 @@ describe('resolveModelRuntimeLimits', () => {
expect(limits.contextWindow).toBe(1_000_000)
expect(limits.maxOutputTokens).toBe(131_072)
})

it.each([
['NVIDIA NIM', 'https://integrate.api.nvidia.com/v1', { NVIDIA_NIM: '1' }],
['OpenRouter', 'https://openrouter.ai/api/v1', { CLAUDE_CODE_USE_OPENAI: '1' }],
['custom endpoint', 'https://proxy.example.test/v1', { CLAUDE_CODE_USE_OPENAI: '1' }],
] as const)('does not leak direct Z.AI GLM-5.3 limits onto %s', (_name, baseUrl, routeEnv) => {
expect(resolveModelRuntimeLimits({
model: 'glm-5.3',
processEnv: {
...routeEnv,
OPENAI_BASE_URL: baseUrl,
},
})).toEqual({
contextWindow: undefined,
maxOutputTokens: undefined,
})
})
it('uses the applied provider profile route before generic custom base URL fallback', () => {
expect(
resolveModelRuntimeLimits({
Expand Down Expand Up @@ -309,6 +390,29 @@ describe('resolveOpenAIShimRuntimeContext - Z.A.I GLM-5.2', () => {
})
})

describe('resolveOpenAIShimRuntimeContext - Z.A.I GLM-5.3', () => {
it.each([
'glm-5.3',
'glm-5.3?reasoning=xhigh',
'glm-5.3?thinking=disabled',
])('uses the explicit direct-route GLM-5.3 contract for %s', model => {
const result = resolveOpenAIShimRuntimeContext({
model,
baseUrl: 'https://api.z.ai/api/coding/paas/v4',
processEnv: {},
})

expect(result.routeId).toBe('zai')
expect(result.catalogEntry?.id).toBe('glm-5.3')
expect(result.openaiShimConfig.thinkingRequestFormat).toBe('zai-compatible')
expect(result.openaiShimConfig.preserveReasoningContent).toBe(true)
expect(result.openaiShimConfig.requireReasoningContentOnAssistantMessages).toBe(true)
expect(result.openaiShimConfig.maxTokensField).toBe('max_tokens')
expect(result.openaiShimConfig.removeBodyFields).toContain('store')
expect(result.openaiShimConfig.enableToolStreaming).toBe(true)
})
})

describe('resolveOpenAIShimRuntimeContext - GLM on a non-Z.AI gateway (#1896)', () => {
it('infers the GLM reasoning shim but not tool streaming for a third-party gateway', () => {
const result = resolveOpenAIShimRuntimeContext({
Expand Down
10 changes: 8 additions & 2 deletions src/integrations/runtimeMetadata.ts
Original file line number Diff line number Diff line change
Expand Up @@ -509,10 +509,16 @@ export function resolveModelRuntimeLimits(options: {
modelApiName,
runtimeEnv,
)
const modelDescriptor =
const catalogModelDescriptor =
getModelDescriptorForCatalogEntry(catalogEntry) ??
getModelDescriptorForCatalogEntry(cachedCatalogEntry) ??
getModelDescriptorForCatalogEntry(cachedCatalogEntry)
const inferredModelDescriptor =
findModelDescriptorForApiName(routeId, modelApiName)
const modelDescriptor =
catalogModelDescriptor ??
(inferredModelDescriptor?.runtimeMetadataScope === 'catalog'
? null
: inferredModelDescriptor)
const externalContextWindow = getOpenAIContextWindowMatches(
modelApiName,
runtimeEnv,
Expand Down
16 changes: 16 additions & 0 deletions src/integrations/vendors/zai.ts
Original file line number Diff line number Diff line change
Expand Up @@ -44,6 +44,22 @@ export default defineVendor({
catalog: {
source: 'static',
models: [
{
id: 'glm-5.3',
apiName: 'glm-5.3',
label: 'GLM-5.3',
modelDescriptorId: 'glm-5.3',
reasoning: {
mode: 'levels',
levels: ['low', 'high', 'xhigh'],
wireFormat: 'zai_compatible',
},
transportOverrides: {
openaiShim: {
enableToolStreaming: true,
},
},
},
{
id: 'glm-5.2',
apiName: 'glm-5.2',
Expand Down
95 changes: 95 additions & 0 deletions src/services/api/openaiShim.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -5437,6 +5437,100 @@ test.each([
expect(requestBody?.reasoning_effort).toBe(effort)
})

test.each([
['glm-5.3', undefined, undefined],
['glm-5.3?reasoning=low', 'enabled', 'low'],
['glm-5.3?reasoning=high', 'enabled', 'high'],
['glm-5.3?reasoning=xhigh', 'enabled', 'max'],
['glm-5.3?thinking=disabled', 'enabled', 'low'],
['glm-5.3?thinking=disabled&reasoning=high', 'enabled', 'high'],
] as const)('Z.AI GLM-5.3 serializes the verified request contract for %s', async (
model,
thinkingType,
reasoningEffort,
) => {
process.env.OPENAI_BASE_URL = 'https://api.z.ai/api/coding/paas/v4'
process.env.OPENAI_API_KEY = 'sk-zai-test'

let requestBody: Record<string, unknown> | undefined
globalThis.fetch = (async (_input, init) => {
requestBody = JSON.parse(String(init?.body))
return new Response(
JSON.stringify({
id: 'chatcmpl-1',
model: 'glm-5.3',
choices: [
{ message: { role: 'assistant', content: 'ok' }, finish_reason: 'stop' },
],
}),
{ headers: { 'Content-Type': 'application/json' } },
)
}) as unknown as FetchType

const client = createOpenAIShimClient({}) as OpenAIShimClient
await client.beta.messages.create({
model,
messages: [{ role: 'user', content: 'hi' }],
max_tokens: 64,
stream: false,
})

expect(requestBody?.model).toBe('glm-5.3')
expect(requestBody?.max_tokens).toBe(64)
expect(requestBody?.max_completion_tokens).toBeUndefined()
expect(requestBody?.store).toBeUndefined()
expect(requestBody?.thinking).toEqual(
thinkingType ? { type: thinkingType } : undefined,
)
expect(requestBody?.reasoning_effort).toBe(reasoningEffort)
})

test('streaming direct Z.AI GLM-5.3 tool requests opt into tool_stream', async () => {
process.env.OPENAI_BASE_URL = 'https://api.z.ai/api/coding/paas/v4'
process.env.OPENAI_API_KEY = 'sk-zai-test'

let requestBody: Record<string, unknown> | undefined
globalThis.fetch = (async (_input, init) => {
requestBody = JSON.parse(String(init?.body))
return makeSseResponse(makeStreamChunks([
{
id: 'chatcmpl-1',
object: 'chat.completion.chunk',
model: 'glm-5.3',
choices: [{ index: 0, delta: { content: 'ok' }, finish_reason: null }],
},
{
id: 'chatcmpl-1',
object: 'chat.completion.chunk',
model: 'glm-5.3',
choices: [{ index: 0, delta: {}, finish_reason: 'stop' }],
},
]))
}) as unknown as FetchType

const client = createOpenAIShimClient({}) as OpenAIShimClient
const stream = await client.beta.messages.create({
model: 'glm-5.3',
messages: [{ role: 'user', content: 'add two numbers' }],
max_tokens: 64,
stream: true,
tools: [{
name: 'add_numbers',
description: 'Add two numbers',
input_schema: {
type: 'object',
properties: { a: { type: 'number' }, b: { type: 'number' } },
required: ['a', 'b'],
},
}],
})
for await (const _event of stream as AsyncIterable<unknown>) {
// Drain the mocked response so request execution completes.
}

expect(requestBody?.tool_stream).toBe(true)
})

test.each([
'GLM-5.1?reasoning=high',
'GLM-4.5-Air?reasoning=high',
Expand Down Expand Up @@ -5477,6 +5571,7 @@ test.each([
test.each([
['non-streaming Z.AI request with tools', 'https://api.z.ai/api/coding/paas/v4', false, true, 'glm-5.2'],
['streaming Z.AI request without tools', 'https://api.z.ai/api/coding/paas/v4', true, false, 'glm-5.2'],
['streaming NVIDIA GLM-5.3 request with tools', 'https://integrate.api.nvidia.com/v1', true, true, 'glm-5.3'],
['streaming non-Z.AI request with tools', 'https://api.openai.com/v1', true, true, 'gpt-4o'],
] as const)('does not send tool_stream for %s', async (_name, baseUrl, stream, includeTools, model) => {
process.env.OPENAI_BASE_URL = baseUrl
Expand Down
Loading
Loading