Skip to content

fix(shim): don't infer Z.AI tool_stream for non-catalog GLM gateways - #1908

Merged
kevincodex1 merged 1 commit into
Twigpine:mainfrom
0xghost42:fix/glm-tool-stream-nim
Jul 10, 2026
Merged

kevincodex1 merged 1 commit into
Twigpine:mainfrom
0xghost42:fix/glm-tool-stream-nim

Conversation

@0xghost42

@0xghost42 0xghost42 commented Jul 8, 2026 •

Copy link
Copy Markdown
Contributor

Closes #1896.

Problem

Serving a GLM model (z-ai/glm-5.2) through a third-party OpenAI-compatible gateway such as NVIDIA NIM (https://integrate.api.nvidia.com/v1) fails on the first turn:

API Error: 400 Validation: Unsupported parameter(s): `tool_stream`

inferRemoteModelOpenAIShimConfig (in runtimeMetadata.ts) infers the full ZAI_GLM_OPENAI_SHIM — including enableToolStreaming: true — for any glm-<n> model that has no catalog entry. tool_stream is a Z.AI-proprietary streaming extension, but the matcher applied it from the model name alone, independent of which gateway is actually serving the model. Any non-Z.AI gateway then rejects the request.

Fix

Keep inferring the GLM reasoning-shaping fields (preserved/echoed reasoning content, zai-compatible thinking format, max_tokens, stripping store) — those help GLM on any endpoint — but stop inferring enableToolStreaming. Only a catalog entry may opt into it: the Z.AI-contract gateways (zai, opencode-go, atlas-cloud, hicap) already set it explicitly via transportOverrides.openaiShim, so they are unaffected (their requests still send tool_stream).

This is host-agnostic on purpose — it fixes hosted NIM, self-hosted NIM, and any other custom OpenAI-compatible GLM endpoint, matching the report’s "breaks any non-Z.AI-official gateway". tool_stream is only a streaming optimization; without it tool calls still work, just not streamed (the shim already has a delete body.tool_stream fallback for tool-incompatible retries).

Test

Added a resolveOpenAIShimRuntimeContext case for z-ai/glm-5.2 on the NIM base URL: asserts no catalog entry (inference path), the reasoning fields still apply, and enableToolStreaming is not true. Verified fail-on-main (the case fails without the change, enableToolStreaming === true). Full runtimeMetadata suite (36) and the openaiShim tool_stream cases (Z.AI path still sends it) pass; tsc --noEmit clean for the touched files.

Distinct from #1892 (that covers chat_template_kwargs / reasoning on NIM); this is only the tool_stream rejection.

Summary by CodeRabbit

  • Bug Fixes
    • Improved compatibility for GLM models when routed through an OpenAI-compatible third-party endpoint.
    • Ensures GLM reasoning-shaping settings are still applied, while tool streaming is explicitly disabled when no matching catalog entry is found.
    • Added automated coverage to prevent regressions for this scenario.

@coderabbitai

coderabbitai Bot commented Jul 8, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: d7512dfb-9932-46aa-a177-e3a8411ef159

📥 Commits

Reviewing files that changed from the base of the PR and between 8b6250f and b956637.

📒 Files selected for processing (2)
  • src/integrations/runtimeMetadata.test.ts
  • src/integrations/runtimeMetadata.ts
📜 Recent review details
⏰ Context from checks skipped due to timeout. (3)
  • GitHub Check: smoke-and-tests (24.11.x)
  • GitHub Check: smoke-and-tests (22)
  • GitHub Check: typecheck
🧰 Additional context used
📓 Path-based instructions (5)
**/*.{ts,tsx}

📄 CodeRabbit inference engine (AGENTS.md)

TypeScript code in this repository must use strict mode and ESM imports.

Files:

  • src/integrations/runtimeMetadata.test.ts
  • src/integrations/runtimeMetadata.ts
**

⚙️ CodeRabbit configuration file

**: # AGENTS.md - AI Agent Coding Guide

This guide is for AI coding agents working in the OpenClaude repository. Read it before changing code, and also follow CONTRIBUTING.md for contributor policy, PR expectations, review follow-up, and project scope.

Project Snapshot

OpenClaude is a coding-agent CLI for cloud and local model providers. It supports OpenAI-compatible APIs, Anthropic, Gemini, DeepSeek, Ollama, MCP, local backends, slash commands, tools, agents, and a React/Ink terminal UI.

The installed CLI runs on Node.js >=22.0.0. Bun is used for source builds, scripts, dependency management, and tests.

Work Style

  • Keep changes focused on one problem.
  • Prefer existing patterns in the file or nearby module.
  • Avoid unrelated formatting, renames, dependency changes, or broad rewrites.
  • Add or update tests when behavior changes.
  • Update docs when setup, commands, provider behavior, or user-facing behavior changes.
  • For new features, larger refactors, dependencies, or runtime changes, follow the issue-first guidance in CONTRIBUTING.md.

Stack And Conventions

  • TypeScript with strict mode and ESM imports.
  • React + Ink for terminal UI.
  • Bun lockfile and Bun scripts for development workflows.
  • Node runtime for the built CLI.

Common libraries and patterns:

  • chalk for terminal color.
  • commander for CLI argument parsing.
  • execa for child processes.
  • Existing service, provider, settings, permission, and UI patterns over new abstractions.

Repository Map

  • src/commands/ - slash and CLI command implementations.
  • src/components/ - React/Ink UI components.
  • src/services/ - API, MCP, OAuth, wiki, voice, and other service integrations.
  • src/tools/ - tool implementations.
  • src/utils/ - shared utilities.
  • src/integrations/ - provider and model integration metadata.
  • src/entrypoints/ - CLI, MCP, SDK, and generated public types.
  • src/tasks/ - local, remote, workflow, and monitor tas...

Files:

  • src/integrations/runtimeMetadata.test.ts
  • src/integrations/runtimeMetadata.ts
**/*

⚙️ CodeRabbit configuration file

**/*: Apply the OpenClaude maintainer review rubric from AGENTS.md. Review the current diff, not stale discussion context. Separate real blockers from suggestions. Do not request changes for vague style churn. Treat approval as merge-ready from CodeRabbit's side, pending required human review and GitHub Checks. If checks are failing or unavailable, say so clearly instead of implying the PR is fully ready.

Files:

  • src/integrations/runtimeMetadata.test.ts
  • src/integrations/runtimeMetadata.ts
{src/services/api/**,src/integrations/**,src/utils/model/**,src/utils/provider*.ts,src/commands/provider/**}

⚙️ CodeRabbit configuration file

{src/services/api/**,src/integrations/**,src/utils/model/**,src/utils/provider*.ts,src/commands/provider/**}: Review provider routing, model selection, env precedence, auth/token handling, OpenAI-compatible shims, retries, proxy behavior, and outbound HTTP behavior with high scrutiny. Block on silent default changes, hidden fallback expansion, credential reuse mistakes, hardcoded provider assumptions, or new network reach that is not intentional and documented.

Files:

  • src/integrations/runtimeMetadata.test.ts
  • src/integrations/runtimeMetadata.ts
{src/**/*.test.ts,src/**/*.test.tsx,tests/**,scripts/**/*.test.ts,vscode-extension/**/*.test.js}

⚙️ CodeRabbit configuration file

{src/**/*.test.ts,src/**/*.test.tsx,tests/**,scripts/**/*.test.ts,vscode-extension/**/*.test.js}: Review tests for meaningful coverage of the changed behavior, isolation of global/env/config state, async cleanup, fake timers, provider profile leaks, and Windows-compatible assumptions. Block when risky runtime changes lack focused regression coverage or tests assert implementation details while missing the user-visible behavior.

Files:

  • src/integrations/runtimeMetadata.test.ts
🔇 Additional comments (3)
src/integrations/runtimeMetadata.ts (2)

221-229: LGTM!


487-510: 📐 Maintainability & Code Quality | 💤 Low value

resolveModelRuntimeLimits precedence change appears unrelated to the GLM tool_stream fix.

The PR objective is narrowly scoped to the GLM/OpenAI-shim tool_stream inference bug. This resolveModelRuntimeLimits precedence adjustment (moving settings below prefix in the nullish coalescing chain) is a separate concern. The logic itself is correct and well-documented, but per AGENTS.md, changes should be kept focused on one problem. Consider splitting this into its own PR with a linked issue, or confirm with maintainers that bundling is intentional.

As per coding guidelines: "Keep changes focused on one problem" and "Avoid unrelated formatting, renames, dependency changes, or broad rewrites."

Sources: Coding guidelines, Path instructions

src/integrations/runtimeMetadata.test.ts (1)

186-205: LGTM!


📝 Walkthrough

Walkthrough

The GLM openai-shim inference path now forces enableToolStreaming: false for inferred GLM configs without a catalog entry, and a new test covers the non-Z.A.I gateway case to confirm the flag stays off while GLM reasoning fields remain set.

Changes

GLM tool streaming inference fix

Layer / File(s) Summary
Disable enableToolStreaming for inferred GLM shim config
src/integrations/runtimeMetadata.ts, src/integrations/runtimeMetadata.test.ts
inferRemoteModelOpenAIShimConfig now returns ZAI_GLM_OPENAI_SHIM with enableToolStreaming: false for inferred GLM models, and the new test verifies the non-catalog, non-Z.A.I gateway path keeps reasoning-shaping config while not enabling tool streaming.

Estimated code review effort: 1 (Simple) | ~10 minutes

Possibly related PRs

Suggested labels: bug

Suggested reviewers: jatmn, kevincodex1

🚥 Pre-merge checks | ✅ 6 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Risk Surface Disclosed ⚠️ Warning The PR changes outgoing OpenAI request behavior (tool_stream inference), but the recorded review only mentions an unrelated flaky test and never flags this risk surface or blocker. Add a review note explicitly calling out the non-Z.AI gateway risk (unsupported tool_stream) and state whether it is blocking or acceptable.
✅ Passed checks (6 passed)
Check name Status Explanation
Title check ✅ Passed The title is concise and accurately describes the GLM tool_stream inference fix.
Description check ✅ Passed The description covers the problem, fix, verification, and context, so it is mostly complete despite using custom headings.
Linked Issues check ✅ Passed The runtimeMetadata change and added test align with #1896 by disabling inferred tool streaming for uncataloged GLM gateways.
Out of Scope Changes check ✅ Passed The diff is limited to the runtimeMetadata fix and its test, with no obvious unrelated changes.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
No Hidden Policy Change ✅ Passed The PR makes one explicit GLM-routing change in runtimeMetadata.ts and adds a focused test; there’s no unrelated cleanup hiding product/trust/network/permission policy shifts.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@kevincodex1

Copy link
Copy Markdown
Member

please fix failing tests

@jatmn jatmn left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I found one repository-state blocker that needs to be addressed before this is ready.

Findings

  • [P1] Fix the failing smoke-and-tests check before this is ready
    Smoke / smoke-and-tests
    GitHub currently reports smoke-and-tests (24.11.x) failing in the Smoke and full unit test suite step, and the smoke-and-tests (22) matrix leg was cancelled after that failure. Since this PR changes the runtime metadata test surface and a reviewer has already asked for the failing tests to be fixed, please get the smoke-and-tests job green or update the PR with evidence that the failure is unrelated to this change before this can be treated as ready.

The name-based shim matcher inferred the full Z.AI GLM contract —
including enableToolStreaming — for any glm-<n> model without a catalog
entry. tool_stream is a Z.AI-proprietary streaming extension, so serving
GLM through an arbitrary OpenAI-compatible gateway (e.g. NVIDIA NIM,
integrate.api.nvidia.com) made every request fail immediately with
400 Unsupported parameter(s): tool_stream.

Only a catalog entry may opt into tool_stream (Z.AI-contract gateways set
it explicitly via transportOverrides.openaiShim). Inferred GLM routes keep
the reasoning-shaping fields, which any GLM endpoint benefits from, but no
longer send tool_stream; without it tool calls are simply not streamed.
@0xghost42
0xghost42 force-pushed the fix/glm-tool-stream-nim branch from 8b6250f to b956637 Compare July 9, 2026 07:30
@0xghost42

Copy link
Copy Markdown
Contributor Author

Thanks @jatmn. Dug into the smoke-and-tests failure — it's braveProvider search > rejects when the provider-level timeout elapses (1 fail / 6572), a timing-based test in the Brave search provider that timed out at ~5.7s under CI load. This PR only touches src/integrations/runtimeMetadata.ts and its test (GLM tool_stream inference) — nothing in WebSearchTool/brave — so they don't intersect. Locally that test passes consistently (3/3), and the runtimeMetadata suite this PR changes is 36/36 green. I've rebased onto latest main to re-run CI and clear the flake; will confirm once it's green.

@opassoca

opassoca commented Jul 9, 2026

Copy link
Copy Markdown

Tested this against the real-world bug (#1896).

Environment: Samsung, Exynos 9820, custom ROM, Android 15 (API 35), Termux/aarch64, Node v26.3.1, npm 11.18.0, @gitlawb/openclaude 0.23.0

Instead of gating on "no catalog entry", I scoped the disable directly at the tool_stream assignment site:

const isNvidiaNimEndpoint = request.baseUrl.includes('nvidia')
if (effectiveTransport === 'chat_completions' && params.stream && shimConfig.enableToolStreaming === true && !isNvidiaNimEndpoint) {
body.tool_stream = true
}

Confirmed working end-to-end against NVIDIA NIM with z-ai/glm-5.2 (multi-step tool-call task, no tool_stream in the outgoing body).

One thing worth considering: this PR's approach (disable when there's no catalog entry) is scoped to inference-path GLM models. A future catalog entry for GLM on a non-Z.AI gateway with enableToolStreaming: true would still slip through the "no catalog entry" check - the baseUrl-based gate covers that case too. Might be worth combining both.

@jatmn jatmn left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the update. I rechecked the previously discussed paths and do not see any remaining actionable issues from my side.

@kevincodex1 LGTM

@kevincodex1
kevincodex1 merged commit 2259c80 into Twigpine:main Jul 10, 2026
5 checks passed
hotmanxp pushed a commit to hotmanxp/openclaude that referenced this pull request Jul 12, 2026
…wigpine#1908)

The name-based shim matcher inferred the full Z.AI GLM contract —
including enableToolStreaming — for any glm-<n> model without a catalog
entry. tool_stream is a Z.AI-proprietary streaming extension, so serving
GLM through an arbitrary OpenAI-compatible gateway (e.g. NVIDIA NIM,
integrate.api.nvidia.com) made every request fail immediately with
400 Unsupported parameter(s): tool_stream.

Only a catalog entry may opt into tool_stream (Z.AI-contract gateways set
it explicitly via transportOverrides.openaiShim). Inferred GLM routes keep
the reasoning-shaping fields, which any GLM endpoint benefits from, but no
longer send tool_stream; without it tool calls are simply not streamed.
hotmanxp added a commit to hotmanxp/openclaude that referenced this pull request Jul 12, 2026
T1 cherry-pick landed 8 upstream commits (Twigpine#1908 shim, Twigpine#1916 diff, Twigpine#1917 hunks,
Twigpine#1913 powershell, Twigpine#1905 plan mode, Twigpine#1906 API cleanup, Twigpine#1901 content, Twigpine#1932
gitdiff cap). The following typecheck fixes were needed because OpenCC has
stricter types than upstream (per `cherry-pick-test-localization-not-shipped`):

1. descriptors.ts:64 -- `enableToolStreaming?: true` widened to `boolean`
   so upstream's `enableToolStreaming: false` (in shim 2259c80) typechecks
2. apiTransform.ts:33 -- `message.uuid ?? ''` (UserMessage.uuid is optional
   in OpenCC Message type; upstream's UserMessage.uuid is required)
3. apiTransform.ts:225-261 -- `stripCallerFieldFromAssistantMessage`
   early-returns for string content; Upstream assumed `content` was array
4. content.ts:77-119 -- guard for OpenCC's optional `Message.message` and
   widened param types to accept `UserMessage` (uuid is `string | undefined`)
5. compact.test.ts:363 -- preserves OpenCC's existing `getAssistantMessageText`
   mock instead of upstream's "_realMessagesModule" (per user directive)

Verification:
- bun run typecheck: 0 errors (was 8 with cherry-picks+before-fix; 0 at baseline pre-T1)
- bun test (T1 areas): 272 pass / 21 fail (21 fails are pre-existing GLM-5.2/estimateMessageTokens)
- bun test (full): 4858 pass / 131 fail (same as baseline pre-T1; no regression)

No rebrand or provider-policy changes needed; cherry-picks are 3-Provider-clean.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

GLM catalog entries hardcode enableToolStreaming: true, sending unsupported tool_stream param to any non-Z.AI-official gateway (breaks NVIDIA NIM)

4 participants