Skip to content

feat: add zai-coding provider for Z.AI GLM Coding Plan - #32174

Open
RedClaus wants to merge 1 commit into
NousResearch:mainfrom
RedClaus:feat/zai-coding-plan-provider
Open

feat: add zai-coding provider for Z.AI GLM Coding Plan#32174
RedClaus wants to merge 1 commit into
NousResearch:mainfrom
RedClaus:feat/zai-coding-plan-provider

Conversation

@RedClaus

Copy link
Copy Markdown
Contributor

Summary

Adds a new zai-coding provider that routes through the Anthropic-compatible endpoint (https://api.z.ai/api/anthropic) instead of the OpenAI-compatible /api/paas/v4 endpoint.

The Z.AI Coding Plan quota is only accessible via the Anthropic endpoint — the standard /api/paas/v4 endpoint returns error 1113 (insufficient balance) even with a valid coding plan subscription.

Changes

  • New plugin: plugins/model-providers/zai-coding/ with api_mode="anthropic_messages" and fallback models (glm-5.1, glm-5-turbo, glm-4.7, glm-4.5-air)
  • Registry: Added zai-coding to PROVIDER_REGISTRY in hermes_cli/auth.py with inference_base_url pointing to the Anthropic endpoint
  • Aliases: zai-coding, zai-coding-plan, glm-coding, z-ai-coding

Usage

hermes config set model.provider zai-coding
hermes config set model.base_url https://api.z.ai/api/anthropic
hermes config set model.default glm-5.1

How it works

The existing Anthropic transport (anthropic_messages api_mode) already handles third-party providers — the code in agent_init.py explicitly supports non-native Anthropic providers like GLM, sending the API key via x-api-key header and skipping OAuth/Claude identity injection.

The _detect_api_mode_for_url() in runtime_provider.py auto-detects anthropic_messages mode for URLs ending in /anthropic, so the transport routes correctly.

Testing

  • Verified the Anthropic endpoint accepts the coding plan key (HTTP 200)
  • Confirmed hermes -z works with --provider zai-coding -m glm-5.1
  • Confirmed runtime config persists via hermes config set

Co-Authored-By: Oz oz-agent@warp.dev

Add a new zai-coding provider that uses the Anthropic-compatible endpoint
(https://api.z.ai/api/anthropic) instead of the OpenAI-compatible /api/paas/v4
endpoint. The Z.AI Coding Plan quota is only accessible via the Anthropic
endpoint — the standard endpoint returns 1113 (insufficient balance).

Changes:
- New plugin: plugins/model-providers/zai-coding/ with anthropic_messages api_mode
- Register zai-coding in PROVIDER_REGISTRY (hermes_cli/auth.py) with
  inference_base_url pointing to the Anthropic endpoint
- Add provider aliases: zai-coding, zai-coding-plan, glm-coding, z-ai-coding

Co-Authored-By: Oz <oz-agent@warp.dev>
@alt-glitch alt-glitch added type/feature New feature or request P3 Low — cosmetic, nice to have comp/plugins Plugin system and bundled plugins comp/cli CLI entry point, hermes_cli/, setup wizard provider/zai ZAI provider labels May 25, 2026
@alt-glitch

Copy link
Copy Markdown
Collaborator

Competing with open PRs #13500, #13911, and #24915 which all address Z.AI provider splitting (Global/China x direct/Coding Plan). This PR takes a narrower approach (Anthropic-endpoint-only Coding Plan provider) vs the broader 4-way split in the competing PRs.

@DeamonDev888

DeamonDev888 commented Jul 9, 2026

Copy link
Copy Markdown

👋 Superseded — please see the new PR

@RedClaus thanks for the proposal. #62080 is the new rebased PR that closes PR #61575.

Your approach adds a separate zai-coding provider. #62080 takes a different approach: instead of adding provider entries to the registry, it extends ZAI_ENDPOINTS with 6 candidates (added anthropic-global and anthropic-cn to the existing 4) and lets detect_zai_endpoint() pick the right wire format automatically.

Both approaches are valid. If maintainers prefer the 4-variant split approach from this PR (#32174) over the unified probe approach in #62080, that's a maintainer design decision — both achieve the same routing correctness. Live audit (9 production keys): both approaches would route all 9 correctly.

DeamonDev888 pushed a commit to DeamonDev888/hermes-agent that referenced this pull request Jul 11, 2026
End-to-end fix for the Z.AI credential pool cascade bug discovered
during a live audit of 8 production keys. Five coordinated layers
ensure the right endpoint is selected, the wrong error is not treated
as a payment error, and the user-configured overrides work end-to-end.

== Problem statement ==

Z.AI Coding Plan subscriptions (id.secret.* format) authenticate on
/api/coding/paas/v4 but the metered /api/paas/v4 endpoint also returns
HTTP 200 for glm-5. Without explicit routing, the metered endpoint
gets cached in auth.json and Coding Plan traffic silently draws from
the metered billing pool instead of the subscription.

A second symptom: Z.AI Coding Plan error codes 1113 (Insufficient
balance or no resource package) and 1308 (Usage limit reached for 5
hour) are per-key rolling quotas. They were misclassified as payment
errors via substring traps intended for Vertex AI (resource exhausted)
and Nous Portal (reached your session usage limit). A single exhausted
key cascade-marked every other key in the pool, taking the whole
provider offline.

A third gap: keys on the Anthropic Messages wire (/api/anthropic) were
not in ZAI_ENDPOINTS at all, so detect_zai_endpoint() returned the
wrong URL for any Anthropic-wire Coding Plan key.

== Solution overview ==

Five layers, all coordinated, all tested end-to-end:

  Layer 1 - _is_payment_error() exemption for Z.AI Coding Plan
    Detects api.z.ai / z.ai/api/coding / zhipuai+coding context plus
    codes 1113, 1308 and message patterns. Returns False so the
    credential pool rotation layer handles per-key 5h rolling quotas
    correctly. Real Z.AI payment errors (e.g. code 1311 plan-block)
    continue to flow through the normal payment-fallback path.

  Layer 2 - Vision auto-detect zai_openai_urls ordering
    The vision helper list now probes Coding Plan endpoints (global +
    China) before the metered paas/v4 fallbacks. Coding Plan keys
    authenticate on first vision call. Corrects an indentation bug in
    the original PR NousResearch#55116.

  Layer 3 - Runtime pool base_url re-resolution
    _resolve_api_key_provider() now re-invokes _resolve_zai_base_url()
    for provider_id == zai so manual-pool entries added via
    her mes auth add honor the cached detected_endpoint state in
    auth.json and the GLM_BASE_URL env override. Brings manual-pool
    parity with the env-seeded path fixed in commit 9e84416.
    Probe failures are caught and logged so they never break pool
    selection.

  Layer 4 - config.yaml model.base_url precedence
    New _configured_zai_base_url() helper reads model.base_url from
    config.yaml when model.provider is a Z.AI alias (zai / glm / z-ai /
    z.ai / zhipu). The precedence chain in _resolve_zai_base_url() is:
        1. GLM_BASE_URL env var (highest)
        2. model.base_url from config.yaml (when provider is Z.AI)
        3. cached detected_endpoint in auth.json
        4. live probe of all candidate endpoints
        5. registry default
    The provider alias guard prevents leakage from non-Z.AI configs.

  Layer 5 - Anthropic-wire endpoints in ZAI_ENDPOINTS
    Extended from 4 to 6 candidates with anthropic-global and
    anthropic-cn. coding-global is probed FIRST (99% of keys that
    accept coding endpoint also accept metered, so probing coding
    first caches the right URL on first try). anthropic-global is
    position 2 (fallthrough for pure Anthropic-wire keys).
    New _zai_probe_path() and _zai_probe_body() helpers dispatch the
    right HTTP body shape: Anthropic Messages for anthropic-* ids
    (/v1/messages, no stream field, anthropic-version header) and
    OpenAI chat completions for everything else.

== Audit results ==

Validated against 8 real production keys (see
tests/agent/test_zai_8keys_audit.py):
  - 6/8 keys: pure Anthropic-wire subscribers (Anthropic Messages)
  - 1/8 keys: pure OpenAI-wire (standard paas/v4)
  - 1/8 keys: full-stack (works on all 3 endpoints)

Live test audit (tests/agent/test_zai_live_audit.py) covers every
Z.AI error category against the real api.z.ai API: 200, 1305, 1308,
1113, 401, 402 - all classified correctly.

Before fix: detect_zai_endpoint() returned the wrong endpoint for
6/8 keys. After fix: 9/9 keys correctly routed and verified working.

== Test coverage ==

249 unit + integration tests, 21 live tests (opt-in via env var):

  tests/hermes_cli/test_zai_5gateway_extension.py      (29 tests)
    ZAI_ENDPOINTS structure, dispatch helpers, signature stability,
    backward compat, probe order pinning

  tests/hermes_cli/test_zai_config_yaml_precedence.py (8 tests)
    GLM_BASE_URL > model.base_url > cached > probe > default

  tests/agent/test_zai_manual_pool_routing.py         (7 tests)
    Runtime re-resolution, GLM_BASE_URL forwarding, probe failure
    tolerance, non-zai provider leak guard

  tests/agent/test_auxiliary_client_zai_payment_classification.py
                                                      (14 tests)
    _is_payment_error Z.AI Coding Plan exemption, verbatim Z.AI
    response bodies, negative cases for Vertex/Bedrock/OpenRouter

  tests/agent/test_zai_e2e_pool.py + test_zai_e2e_pool_rotation.py
                                                     (8 tests)
    Real HTTP round-trip via local mock Z.AI server, 5-key
    round_robin, revoked key, probe failure recovery

  tests/agent/test_zai_live.py + test_zai_live_audit.py +
  test_zai_8keys_audit.py                           (21 tests)
    Live against api.z.ai, opt-in via HERMES_RUN_LIVE=1 or
    GLM_AUDIT_KEYS env var. Keys masked to 8-char prefix in all
    output. File is safe to commit.

== Related work ==

  PR NousResearch#61492 - _is_payment_error() Z.AI exemption (cloned in Layer 1)
  PR NousResearch#55116 - vision helper coding endpoint (Layer 2)
  PR NousResearch#58088 - config.yaml base_url precedence (Layer 4)
  PR NousResearch#24915 - 4-variant provider split (intentionally NOT adopted:
    orthogonal refactor, deferred)
  PR NousResearch#54643 - /api/anthropic to /api/coding/paas/v4 rewrite
    (orthogonal; this PR probes the right wire, NousResearch#54643 rewrites at
    runtime if a user forces the OpenAI wire via config)
  PR NousResearch#55007 - stream probes without reading bodies (orthogonal;
    this PR can refactor _zai_probe_path to use httpx.stream() once
    NousResearch#55007 lands)
  PR NousResearch#61333 - skip /anthropic to /v1 rewrite (complementary)
  PR NousResearch#60753 - preserve /anthropic for custom vision (complementary)
  PR NousResearch#32174 - add zai-coding provider (different design choice)
  PR NousResearch#60034 - Z.AI Coding overload adaptive backoff (already merged)

== Related issues ==

  NousResearch#61487 - cascade _is_payment_error (Layer 1 closes)
  NousResearch#61563 - manual-pool routing audit (this PR implements)
  NousResearch#47970 - GLM-5.2 context_length fallback (Layer 5 helps)
  NousResearch#55112 - auxiliary vision hardcoded zai (Layer 5 helps)
  NousResearch#47685 - Hermes Agent prompt block on Z.ai (orthogonal)

== Test commands ==

  Unit + integration (no network):
    pytest tests/hermes_cli/test_api_key_providers.py \
           tests/hermes_cli/test_zai_config_yaml_precedence.py \
           tests/hermes_cli/test_zai_5gateway_extension.py \
           tests/agent/test_auxiliary_client.py::TestIsPaymentError \
           tests/agent/test_auxiliary_client_zai_payment_classification.py \
           tests/agent/test_zai_manual_pool_routing.py \
           tests/agent/test_zai_e2e_pool.py \
           tests/agent/test_zai_e2e_pool_rotation.py -v

  Live (consumes Z.AI quota, requires HERMES_RUN_LIVE=1):
    pytest tests/agent/test_zai_live.py \
           tests/agent/test_zai_live_audit.py \
           tests/agent/test_zai_8keys_audit.py -v --runlive

  Cross-platform check:
    scripts/check-windows-footguns.py --diff upstream/main

== Security ==

  - All live test keys read from env var (GLM_TEST_KEYS, GLM_AUDIT_KEYS,
    GLM_WORKING_KEY) NEVER from files. Files are safe to commit
    publicly.
  - All keys masked to 8-char prefix in any printed output.
  - No hardcoded credentials, no path leaks, no LAN IPs.
  - Privacy scan: 0 secrets, 0 private paths in diff.

Co-authored-by: Hermes triage bot <bot@nousresearch.com>
Refs: NousResearch#61487, NousResearch#61563, PR NousResearch#61492, PR NousResearch#55116, PR NousResearch#58088, PR NousResearch#24915,
      PR NousResearch#54643, PR NousResearch#55007, PR NousResearch#61333, PR NousResearch#60753, PR NousResearch#32174,
      PR NousResearch#60034, NousResearch#47970, NousResearch#55112, NousResearch#47685

@teknium1 teknium1 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for isolating the Coding Plan route as a small provider change. The premise remains current: hermes_cli/auth.py:623-667 only probes OpenAI-compatible Z.AI endpoints, while hermes_cli/runtime_provider.py:101-139 already supports /anthropic URLs as anthropic_messages.

Problems

  • The PR diff contains no regression tests. Please cover resolve_runtime_provider("zai-coding") through the API-key path and assert the Anthropic URL and anthropic_messages mode (hermes_cli/runtime_provider.py:2033-2062).
  • The provider guide still documents automatic probing only (website/docs/integrations/providers.md:279-281); it needs an entry for this explicitly selected Coding Plan route and aliases.

Suggested changes

  • Add runtime/provider-alias tests and update the provider guide before salvage.

This is an automated hermes-sweeper review.

from providers.base import ProviderProfile


zai_coding = ProviderProfile(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please add regression coverage for the end-to-end runtime resolution of this profile: with zai-coding selected, assert the resolved base URL remains https://api.z.ai/api/anthropic and the API mode is anthropic_messages. The runtime mode is derived from the URL at hermes_cli/runtime_provider.py:2057-2062, so profile registration alone does not exercise the routing contract.

@teknium1 teknium1 added sweeper:risk-compatibility Sweeper risk: may break existing users, config, migrations, defaults, or upgrades sweeper:blast-contained Sweeper blast radius: contained — one narrow path / opt-in / few users labels Jul 13, 2026

@GottZ GottZ left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This was generated by AI during triage.

Summary

Three PRs address Z.AI Coding Plan routing, but their diffs split into two mechanisms: #32174 adds the reported Anthropic-compatible /api/anthropic route, while #34323 and #47140 add a broader dedicated provider around the OpenAI-compatible /api/coding/paas/v4 family.

Related pull requests

  • #32174 related — (+47/-0) — keep open for salvage: the narrow diff directly targets the reported cause by registering zai-coding with the Anthropic endpoint and anthropic_messages mode. Consistent with the keep_open review on #32174, it still needs runtime/alias regression tests and provider documentation before merge.
  • #34323 [closed] related — (+740/-82) — closed alternative/reference implementation: it adds extensive provider, setup, runtime, model, credential-pool, and test integration, but routes Coding Plan traffic through OpenAI-compatible /api/coding/paas/v4, not the Anthropic-only route asserted by #32174. It remains relevant because it demonstrates the broader dedicated-provider approach later repeated and refined by #47140.
  • #47140 [closed] related — (+1171/-160) — closed as redundant with current main: it expands the same OpenAI-compatible dedicated-provider approach with endpoint-family guards, shared credentials, setup integration, tests, and glm-5.2. The contributor closure verified that main already probes Coding Plan endpoints and supports glm-5.2, but that finding does not resolve #32174's distinct /api/anthropic premise because #47140 neither implements nor tests that transport.

Duplicates

#34323 and #47140 are substantially duplicate implementations of a dedicated zai-coding provider using the OpenAI-compatible /api/coding/paas/v4 endpoint family; #47140 is the more comprehensive iteration. #32174 overlaps in user-facing goal but is not a code duplicate because it uses the Anthropic-compatible endpoint and transport.

Suggested consolidation

Consolidate on #32174, but do not merge it yet: preserve the keep_open review's required runtime/alias tests and provider-guide update, and verify the claimed Anthropic-only quota behavior against the current main implementation. Keep #34323 and #47140 closed as redundant OpenAI-compatible alternatives; if the live verification confirms that current main serves all relevant Coding Plan keys, close #32174 too, otherwise merge the tested and documented narrow #32174 route.

Cross-PR triage: Reviewed 3 pull requests and 0 issues in this complex. Each diff was read against this issue; Assessment working set: 155 kB of PR diffs, 10 kB of issue/PR text, 5 kB of discussion (8 comments), 1 verify verdict. verdicts reflect diff content, not PR titles. Part of an automated triage batch.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/cli CLI entry point, hermes_cli/, setup wizard comp/plugins Plugin system and bundled plugins P3 Low — cosmetic, nice to have provider/zai ZAI provider sweeper:blast-contained Sweeper blast radius: contained — one narrow path / opt-in / few users sweeper:risk-compatibility Sweeper risk: may break existing users, config, migrations, defaults, or upgrades type/feature New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants