feat: add zai-coding provider for Z.AI GLM Coding Plan - #32174
Conversation
Add a new zai-coding provider that uses the Anthropic-compatible endpoint (https://api.z.ai/api/anthropic) instead of the OpenAI-compatible /api/paas/v4 endpoint. The Z.AI Coding Plan quota is only accessible via the Anthropic endpoint — the standard endpoint returns 1113 (insufficient balance). Changes: - New plugin: plugins/model-providers/zai-coding/ with anthropic_messages api_mode - Register zai-coding in PROVIDER_REGISTRY (hermes_cli/auth.py) with inference_base_url pointing to the Anthropic endpoint - Add provider aliases: zai-coding, zai-coding-plan, glm-coding, z-ai-coding Co-Authored-By: Oz <oz-agent@warp.dev>
👋 Superseded — please see the new PR@RedClaus thanks for the proposal. #62080 is the new rebased PR that closes PR #61575. Your approach adds a separate Both approaches are valid. If maintainers prefer the 4-variant split approach from this PR (#32174) over the unified probe approach in #62080, that's a maintainer design decision — both achieve the same routing correctness. Live audit (9 production keys): both approaches would route all 9 correctly. |
End-to-end fix for the Z.AI credential pool cascade bug discovered
during a live audit of 8 production keys. Five coordinated layers
ensure the right endpoint is selected, the wrong error is not treated
as a payment error, and the user-configured overrides work end-to-end.
== Problem statement ==
Z.AI Coding Plan subscriptions (id.secret.* format) authenticate on
/api/coding/paas/v4 but the metered /api/paas/v4 endpoint also returns
HTTP 200 for glm-5. Without explicit routing, the metered endpoint
gets cached in auth.json and Coding Plan traffic silently draws from
the metered billing pool instead of the subscription.
A second symptom: Z.AI Coding Plan error codes 1113 (Insufficient
balance or no resource package) and 1308 (Usage limit reached for 5
hour) are per-key rolling quotas. They were misclassified as payment
errors via substring traps intended for Vertex AI (resource exhausted)
and Nous Portal (reached your session usage limit). A single exhausted
key cascade-marked every other key in the pool, taking the whole
provider offline.
A third gap: keys on the Anthropic Messages wire (/api/anthropic) were
not in ZAI_ENDPOINTS at all, so detect_zai_endpoint() returned the
wrong URL for any Anthropic-wire Coding Plan key.
== Solution overview ==
Five layers, all coordinated, all tested end-to-end:
Layer 1 - _is_payment_error() exemption for Z.AI Coding Plan
Detects api.z.ai / z.ai/api/coding / zhipuai+coding context plus
codes 1113, 1308 and message patterns. Returns False so the
credential pool rotation layer handles per-key 5h rolling quotas
correctly. Real Z.AI payment errors (e.g. code 1311 plan-block)
continue to flow through the normal payment-fallback path.
Layer 2 - Vision auto-detect zai_openai_urls ordering
The vision helper list now probes Coding Plan endpoints (global +
China) before the metered paas/v4 fallbacks. Coding Plan keys
authenticate on first vision call. Corrects an indentation bug in
the original PR NousResearch#55116.
Layer 3 - Runtime pool base_url re-resolution
_resolve_api_key_provider() now re-invokes _resolve_zai_base_url()
for provider_id == zai so manual-pool entries added via
her mes auth add honor the cached detected_endpoint state in
auth.json and the GLM_BASE_URL env override. Brings manual-pool
parity with the env-seeded path fixed in commit 9e84416.
Probe failures are caught and logged so they never break pool
selection.
Layer 4 - config.yaml model.base_url precedence
New _configured_zai_base_url() helper reads model.base_url from
config.yaml when model.provider is a Z.AI alias (zai / glm / z-ai /
z.ai / zhipu). The precedence chain in _resolve_zai_base_url() is:
1. GLM_BASE_URL env var (highest)
2. model.base_url from config.yaml (when provider is Z.AI)
3. cached detected_endpoint in auth.json
4. live probe of all candidate endpoints
5. registry default
The provider alias guard prevents leakage from non-Z.AI configs.
Layer 5 - Anthropic-wire endpoints in ZAI_ENDPOINTS
Extended from 4 to 6 candidates with anthropic-global and
anthropic-cn. coding-global is probed FIRST (99% of keys that
accept coding endpoint also accept metered, so probing coding
first caches the right URL on first try). anthropic-global is
position 2 (fallthrough for pure Anthropic-wire keys).
New _zai_probe_path() and _zai_probe_body() helpers dispatch the
right HTTP body shape: Anthropic Messages for anthropic-* ids
(/v1/messages, no stream field, anthropic-version header) and
OpenAI chat completions for everything else.
== Audit results ==
Validated against 8 real production keys (see
tests/agent/test_zai_8keys_audit.py):
- 6/8 keys: pure Anthropic-wire subscribers (Anthropic Messages)
- 1/8 keys: pure OpenAI-wire (standard paas/v4)
- 1/8 keys: full-stack (works on all 3 endpoints)
Live test audit (tests/agent/test_zai_live_audit.py) covers every
Z.AI error category against the real api.z.ai API: 200, 1305, 1308,
1113, 401, 402 - all classified correctly.
Before fix: detect_zai_endpoint() returned the wrong endpoint for
6/8 keys. After fix: 9/9 keys correctly routed and verified working.
== Test coverage ==
249 unit + integration tests, 21 live tests (opt-in via env var):
tests/hermes_cli/test_zai_5gateway_extension.py (29 tests)
ZAI_ENDPOINTS structure, dispatch helpers, signature stability,
backward compat, probe order pinning
tests/hermes_cli/test_zai_config_yaml_precedence.py (8 tests)
GLM_BASE_URL > model.base_url > cached > probe > default
tests/agent/test_zai_manual_pool_routing.py (7 tests)
Runtime re-resolution, GLM_BASE_URL forwarding, probe failure
tolerance, non-zai provider leak guard
tests/agent/test_auxiliary_client_zai_payment_classification.py
(14 tests)
_is_payment_error Z.AI Coding Plan exemption, verbatim Z.AI
response bodies, negative cases for Vertex/Bedrock/OpenRouter
tests/agent/test_zai_e2e_pool.py + test_zai_e2e_pool_rotation.py
(8 tests)
Real HTTP round-trip via local mock Z.AI server, 5-key
round_robin, revoked key, probe failure recovery
tests/agent/test_zai_live.py + test_zai_live_audit.py +
test_zai_8keys_audit.py (21 tests)
Live against api.z.ai, opt-in via HERMES_RUN_LIVE=1 or
GLM_AUDIT_KEYS env var. Keys masked to 8-char prefix in all
output. File is safe to commit.
== Related work ==
PR NousResearch#61492 - _is_payment_error() Z.AI exemption (cloned in Layer 1)
PR NousResearch#55116 - vision helper coding endpoint (Layer 2)
PR NousResearch#58088 - config.yaml base_url precedence (Layer 4)
PR NousResearch#24915 - 4-variant provider split (intentionally NOT adopted:
orthogonal refactor, deferred)
PR NousResearch#54643 - /api/anthropic to /api/coding/paas/v4 rewrite
(orthogonal; this PR probes the right wire, NousResearch#54643 rewrites at
runtime if a user forces the OpenAI wire via config)
PR NousResearch#55007 - stream probes without reading bodies (orthogonal;
this PR can refactor _zai_probe_path to use httpx.stream() once
NousResearch#55007 lands)
PR NousResearch#61333 - skip /anthropic to /v1 rewrite (complementary)
PR NousResearch#60753 - preserve /anthropic for custom vision (complementary)
PR NousResearch#32174 - add zai-coding provider (different design choice)
PR NousResearch#60034 - Z.AI Coding overload adaptive backoff (already merged)
== Related issues ==
NousResearch#61487 - cascade _is_payment_error (Layer 1 closes)
NousResearch#61563 - manual-pool routing audit (this PR implements)
NousResearch#47970 - GLM-5.2 context_length fallback (Layer 5 helps)
NousResearch#55112 - auxiliary vision hardcoded zai (Layer 5 helps)
NousResearch#47685 - Hermes Agent prompt block on Z.ai (orthogonal)
== Test commands ==
Unit + integration (no network):
pytest tests/hermes_cli/test_api_key_providers.py \
tests/hermes_cli/test_zai_config_yaml_precedence.py \
tests/hermes_cli/test_zai_5gateway_extension.py \
tests/agent/test_auxiliary_client.py::TestIsPaymentError \
tests/agent/test_auxiliary_client_zai_payment_classification.py \
tests/agent/test_zai_manual_pool_routing.py \
tests/agent/test_zai_e2e_pool.py \
tests/agent/test_zai_e2e_pool_rotation.py -v
Live (consumes Z.AI quota, requires HERMES_RUN_LIVE=1):
pytest tests/agent/test_zai_live.py \
tests/agent/test_zai_live_audit.py \
tests/agent/test_zai_8keys_audit.py -v --runlive
Cross-platform check:
scripts/check-windows-footguns.py --diff upstream/main
== Security ==
- All live test keys read from env var (GLM_TEST_KEYS, GLM_AUDIT_KEYS,
GLM_WORKING_KEY) NEVER from files. Files are safe to commit
publicly.
- All keys masked to 8-char prefix in any printed output.
- No hardcoded credentials, no path leaks, no LAN IPs.
- Privacy scan: 0 secrets, 0 private paths in diff.
Co-authored-by: Hermes triage bot <bot@nousresearch.com>
Refs: NousResearch#61487, NousResearch#61563, PR NousResearch#61492, PR NousResearch#55116, PR NousResearch#58088, PR NousResearch#24915,
PR NousResearch#54643, PR NousResearch#55007, PR NousResearch#61333, PR NousResearch#60753, PR NousResearch#32174,
PR NousResearch#60034, NousResearch#47970, NousResearch#55112, NousResearch#47685
teknium1
left a comment
There was a problem hiding this comment.
Thanks for isolating the Coding Plan route as a small provider change. The premise remains current: hermes_cli/auth.py:623-667 only probes OpenAI-compatible Z.AI endpoints, while hermes_cli/runtime_provider.py:101-139 already supports /anthropic URLs as anthropic_messages.
Problems
- The PR diff contains no regression tests. Please cover
resolve_runtime_provider("zai-coding")through the API-key path and assert the Anthropic URL andanthropic_messagesmode (hermes_cli/runtime_provider.py:2033-2062). - The provider guide still documents automatic probing only (
website/docs/integrations/providers.md:279-281); it needs an entry for this explicitly selected Coding Plan route and aliases.
Suggested changes
- Add runtime/provider-alias tests and update the provider guide before salvage.
This is an automated hermes-sweeper review.
| from providers.base import ProviderProfile | ||
|
|
||
|
|
||
| zai_coding = ProviderProfile( |
There was a problem hiding this comment.
Please add regression coverage for the end-to-end runtime resolution of this profile: with zai-coding selected, assert the resolved base URL remains https://api.z.ai/api/anthropic and the API mode is anthropic_messages. The runtime mode is derived from the URL at hermes_cli/runtime_provider.py:2057-2062, so profile registration alone does not exercise the routing contract.
GottZ
left a comment
There was a problem hiding this comment.
This was generated by AI during triage.
Summary
Three PRs address Z.AI Coding Plan routing, but their diffs split into two mechanisms: #32174 adds the reported Anthropic-compatible /api/anthropic route, while #34323 and #47140 add a broader dedicated provider around the OpenAI-compatible /api/coding/paas/v4 family.
Related pull requests
- #32174
related— (+47/-0) — keep open for salvage: the narrow diff directly targets the reported cause by registeringzai-codingwith the Anthropic endpoint andanthropic_messagesmode. Consistent with the keep_open review on #32174, it still needs runtime/alias regression tests and provider documentation before merge. - #34323 [closed]
related— (+740/-82) — closed alternative/reference implementation: it adds extensive provider, setup, runtime, model, credential-pool, and test integration, but routes Coding Plan traffic through OpenAI-compatible/api/coding/paas/v4, not the Anthropic-only route asserted by #32174. It remains relevant because it demonstrates the broader dedicated-provider approach later repeated and refined by #47140. - #47140 [closed]
related— (+1171/-160) — closed as redundant with current main: it expands the same OpenAI-compatible dedicated-provider approach with endpoint-family guards, shared credentials, setup integration, tests, andglm-5.2. The contributor closure verified that main already probes Coding Plan endpoints and supportsglm-5.2, but that finding does not resolve #32174's distinct/api/anthropicpremise because #47140 neither implements nor tests that transport.
Duplicates
#34323 and #47140 are substantially duplicate implementations of a dedicated zai-coding provider using the OpenAI-compatible /api/coding/paas/v4 endpoint family; #47140 is the more comprehensive iteration. #32174 overlaps in user-facing goal but is not a code duplicate because it uses the Anthropic-compatible endpoint and transport.
Suggested consolidation
Consolidate on #32174, but do not merge it yet: preserve the keep_open review's required runtime/alias tests and provider-guide update, and verify the claimed Anthropic-only quota behavior against the current main implementation. Keep #34323 and #47140 closed as redundant OpenAI-compatible alternatives; if the live verification confirms that current main serves all relevant Coding Plan keys, close #32174 too, otherwise merge the tested and documented narrow #32174 route.
Cross-PR triage: Reviewed 3 pull requests and 0 issues in this complex. Each diff was read against this issue; Assessment working set: 155 kB of PR diffs, 10 kB of issue/PR text, 5 kB of discussion (8 comments), 1 verify verdict. verdicts reflect diff content, not PR titles. Part of an automated triage batch.
Summary
Adds a new
zai-codingprovider that routes through the Anthropic-compatible endpoint (https://api.z.ai/api/anthropic) instead of the OpenAI-compatible/api/paas/v4endpoint.The Z.AI Coding Plan quota is only accessible via the Anthropic endpoint — the standard
/api/paas/v4endpoint returns error 1113 (insufficient balance) even with a valid coding plan subscription.Changes
plugins/model-providers/zai-coding/withapi_mode="anthropic_messages"and fallback models (glm-5.1, glm-5-turbo, glm-4.7, glm-4.5-air)zai-codingtoPROVIDER_REGISTRYinhermes_cli/auth.pywithinference_base_urlpointing to the Anthropic endpointzai-coding,zai-coding-plan,glm-coding,z-ai-codingUsage
How it works
The existing Anthropic transport (
anthropic_messagesapi_mode) already handles third-party providers — the code inagent_init.pyexplicitly supports non-native Anthropic providers like GLM, sending the API key viax-api-keyheader and skipping OAuth/Claude identity injection.The
_detect_api_mode_for_url()inruntime_provider.pyauto-detectsanthropic_messagesmode for URLs ending in/anthropic, so the transport routes correctly.Testing
hermes -zworks with--provider zai-coding -m glm-5.1hermes config setCo-Authored-By: Oz oz-agent@warp.dev