Skip to content

fix(catalog): declare GLM reasoning effort tiers - #10963

Merged
diegosouzapw merged 123 commits into
diegosouzapw:release/v3.8.50from
xz-dev:fix/glm-effort-metadata
Aug 23, 2026
Merged

diegosouzapw merged 123 commits into
diegosouzapw:release/v3.8.50from
xz-dev:fix/glm-effort-metadata

Conversation

@xz-dev

@xz-dev xz-dev commented Aug 21, 2026

Copy link
Copy Markdown
Contributor

Summary

  • Declare exact, routable GLM 5.2/5.3 reasoning-effort tiers for the native GLM providers while keeping earlier thinking-capable GLM families tierless.
  • Stop the unified catalog from inventing generic OpenAI effort tiers for GLM-family models unless their provider registry declares a verified contract; existing Ollama Cloud, Crof, and Command Code contracts remain authoritative.
  • Align ZCode's catalog and runtime allowlist by removing GlmExecutor-only effort aliases and advertising no selectable effort tiers on the local app-server transport.

Related Issues

Validation

Choose the change type and focused loop from the
Contribution Golden Path. The full unit suite,
Vitest, the 60% coverage gate, and the production build all run in CI on this PR (#8329):

  • Change type: provider / routing
  • Focused tests and category gates from the golden path
  • npm run lint
  • Reconciled with the current active release base; focused checks rerun afterward
  • Production-code changes include a new or updated automated test in this PR
  • SonarQube is temporarily opt-in while the private project has no quota; it is not a PR gate.

Focused validation:

  • npx vitest run open-sse/mcp-server/__tests__/glmCodingProviderConfig.test.ts
  • node --import tsx/esm --test tests/unit/glm-5.3-catalog-and-effort-tiers.test.ts
  • node --import tsx/esm --test tests/unit/effort-thinking-standardization-6241.test.ts
  • node --import tsx/esm --test tests/unit/sync-reasoning-supported-efforts-7694.test.ts
  • node --import tsx/esm --test tests/unit/deepseek-thinking-efforts.test.ts
  • node --import tsx/esm --test tests/unit/catalog-helpers-extraction.test.ts
  • node --import tsx/esm --test tests/unit/zcode-provider.test.ts
  • node --import tsx/esm --test tests/unit/zcode-executor.test.ts
  • npm run typecheck:core
  • bun scripts/check/check-provider-consistency.ts
  • npm run check:cycles
  • npm run check:changelog-integrity
  • git diff --check origin/release/v3.8.50

Tests Added Or Updated

  • open-sse/mcp-server/__tests__/glmCodingProviderConfig.test.ts
  • tests/unit/glm-5.3-catalog-and-effort-tiers.test.ts
  • tests/unit/zcode-executor.test.ts
  • tests/unit/zcode-provider.test.ts

Coverage Notes

The GLM-specific tests cover exact native tiers, authoritative empty declarations, registry-wide GLM fallback suppression, catalog enrichment, alias de-duplication, and GLM executor transport mapping. ZCode tests cover the provider model list, rejected effort aliases, shared runtime allowlist, JSON responses, and SSE responses. CI owns the full coverage gate.

Reviewer Notes

  • The registry-wide regression dynamically audits every GLM-family registry entry. Provider-declared effort arrays win; undeclared GLM entries receive an empty tier list rather than generic OpenAI tiers.
  • src/lib/modelMetadataRegistry.ts overlaps the effort fallback area touched by fix(catalog): preserve provider effort tiers #10953. If that PR lands first, reconciliation must preserve source-declared Kimi tiers, explicit registry arrays (including []), GLM undeclared-tier suppression, then generic fallback. No Kimi test-file changes are included here.
  • ZCode's app-server executor never consumed reasoning_effort; this removes four previously listed pseudo-model aliases that were not routable through that transport.
  • npm run check:provider-assets still reports the unchanged upstream freebuff.png as 512x512 over the 256px budget. This PR changes no provider assets.
  • No migrations or feature flags.

@xz-dev
xz-dev requested a review from diegosouzapw as a code owner August 21, 2026 08:30
@xz-dev

xz-dev commented Aug 21, 2026

Copy link
Copy Markdown
Contributor Author

CI triage for f63b4313594576d42b5f082b1945c15575400a59:

The GLM/ZCode checks added or changed by this PR pass in CI:

  • Vitest fast-path: success (including glmCodingProviderConfig.test.ts)
  • GLM family detection covers numeric, Z1, and bare provider model ids: pass
  • registry-wide GLM undeclared-tier suppression: pass
  • provider-routable GLM catalog tiers: pass
  • ZCode provider registry/allowlist tests: pass

The remaining red checks are release-line drift outside this PR's files:

  • Build + DAST: src/app/api/usage/utilization/route.ts imports missing @/lib/db/connections (introduced on the current base by fix(dashboard): show account email/name in Utilization Account Split cards #10939).
  • Docs: provider reference/SVGs still say 346 while the current base has 347 providers.
  • Fast Quality: stale public-credential allowlist, missing mutation coverage for src/sse/services/auth.ts, and typecheck-ratchet errors in freebuff.ts, providerHealthMatrix.ts, and auth.ts.
  • ESLint ratchet: OpenAPI coverage is 38.6% vs the 39.2% baseline.
  • Unit shards report numerous unrelated current-base regressions (env-doc drift, ONNX dependency duplication, Kimi URL/golden drift, TLS golden drift, etc.). The two legacy glm-executor.test.ts failures concern missing x-api-key headers; this PR changes only a documentation URL in glm.ts, not header logic.

Recent merged PRs against this release line show the same broad Quality Gates failure pattern. No branch-specific failing assertion was found.

diegosouzapw and others added 27 commits August 21, 2026 08:02
…souzapw#10888)

Validado no worktree combinado: typecheck:core, changelog-integrity, file-size, lint todos verdes. Correção real dos 3 alertas CodeQL (HMAC em vez de hash bruto, URL parsing em vez de substring, dismiss documentado). CI vermelho é o base-red já rastreado em diegosouzapw#9985.
…ouzapw#10727) (diegosouzapw#10916)

Validado no worktree combinado: typecheck:core, changelog-integrity, file-size, lint e 2/2 testes focados passando. Diagnóstico bem investigado do timeout WS do Meta AI (readyState exposto no erro). CI vermelho é o base-red já rastreado em diegosouzapw#9985.
diegosouzapw#7592) (diegosouzapw#10921)

Validado no worktree combinado: typecheck:core, changelog-integrity, file-size, lint e 7/7 testes focados passando. Investigação completa com verificação de ancestralidade via merge-base antes de fechar a issue original. CI vermelho é o base-red já rastreado em diegosouzapw#9985.
…ode_modules (diegosouzapw#7346) (diegosouzapw#10924)

Validado no worktree combinado: typecheck:core, changelog-integrity, file-size, lint e teste focado passando. Root cause bem documentado (distDir customizado gera dois node_modules externalizados). CI vermelho é o base-red já rastreado em diegosouzapw#9985.
…osouzapw#5483 regression) (diegosouzapw#10946)

Validado no worktree combinado: typecheck:core, changelog-integrity, complexity, cognitive-complexity, file-size, lint e testes focados todos verdes. Regressão real corrigida (apiType=chat agora é honrado em vez de forçado para /responses). CI vermelho é o base-red já rastreado em diegosouzapw#9985.
… on (diegosouzapw#10945) (diegosouzapw#10951)

Validado no worktree combinado: mesmos gates + testes focados verdes. Bug real e bem reproduzido (least-used nunca gravava lastUsedAt, sempre a mesma conexão escolhida). CI vermelho é o base-red já rastreado em diegosouzapw#9985.
Validado no worktree combinado: mesmos gates + teste focado verde. Preserva effort_tiers declarados pelo provider (Kimi k3) em vez de substituir pela lista canônica genérica. CI vermelho é o base-red já rastreado em diegosouzapw#9985.
… effort immediately (diegosouzapw#10957)

Validado no worktree combinado: mesmos gates + testes focados verdes. Fix bem medido (context window real vs anunciado divergindo por até 24h para modelos sincronizados fora do ciclo). CI vermelho é o base-red já rastreado em diegosouzapw#9985.
… 404ing (diegosouzapw#10947) (diegosouzapw#10958)

Validado no worktree combinado: mesmos gates + teste focado verde. Root cause medido na release publicada v3.8.49 (nome de artefato NSIS com espaço vs. hífen no manifest). CI vermelho é o base-red já rastreado em diegosouzapw#9985.
…zapw#10926)

Validado no worktree combinado: mesmos gates + testes focados verdes. Extensão opt-in bem desenhada sobre diegosouzapw#10909 (dimensão de uso real via call_logs). CI vermelho é o base-red já rastreado em diegosouzapw#9985.
…iegosouzapw#10855)

Tirado de Draft e validado no worktree combinado: mesmos gates verdes (mudança de UI/i18n sem cobertura automatizada dedicada, mas de baixo risco — só warnings e ocultação condicional de UI). Fix de UX real (diegosouzapw#10794 — 401 confuso ao pular senha no onboarding). CI vermelho é o base-red já rastreado em diegosouzapw#9985.
diegosouzapw#10948)

Validado no worktree combinado: mesmos gates + 36 testes focados verdes. Feature bem documentada e testada (tool calling completo para copilot-m365-web via SignalR, incluindo keepalives e detecção de erro silencioso). CI vermelho é o base-red já rastreado em diegosouzapw#9985.
Validado no worktree combinado: typecheck:core, changelog-integrity, complexity, cognitive-complexity, file-size, lint e teste focado (vps-compose) todos verdes. Bundle Docker aditivo, seguro-por-padrão (loopback, secrets obrigatórios, imagem pinada), bem documentado. CI vermelho é o base-red já rastreado em diegosouzapw#9985.
…est) (diegosouzapw#9985) (diegosouzapw#10778)

Reconciliado com a release e revalidado: typecheck:core, check:dead-code (410 real vs 416 na baseline resolvida — a PR mede corretamente sua própria melhoria), lint (adicionei 1 entrada de suppression para GrokBuildToolCard.tsx, arquivo mergeado depois que esta branch nasceu, 2 violações novas de react-hooks/set-state-in-effect não capturadas pela contagem original), complexity, cognitive-complexity, file-size, changelog-integrity e 11/11 testes do tieredRotation todos verdes. Drena 3 dos 8 hard failures do diegosouzapw#9985. Obrigado!
…d empty-turn errors (diegosouzapw#9909)

⭐5 — Cursor PKCE login com Bearer quota, auto router e empty-turn errors. Feature completa e testada (11 arquivos de teste, 133 testes focados, todos verdes).

**Validação (worktree combinado `.claude/worktrees/fix-9909`, board sobre `origin/release/v3.8.50`):**
- 3 conflitos reais resolvidos: `config/quality/eslint-suppressions.json` (aditivo), `open-sse/config/providers/registry/cursor/index.ts` (dedup de 208 entradas de catálogo, 0 IDs duplicados verificado), `open-sse/executors/cursor.ts` (imports aditivos).
- `npm run typecheck:core`: limpo.
- `check-changelog-integrity`, `check-file-size`, `check-complexity` (2615/2774), `check-cognitive-complexity` (1175/1223), `check-dead-code` (410/416): todos OK.
- `check-public-creds`: 1 entrada obsoleta pré-existente na allowlist (`copilot-m365-web.ts:330`), já presente no tip da release — não é desta PR.
- `npm run lint`: 0 errors (5 warnings pré-existentes).
- Testes focados (`cursor-agent-cli-version`, `cursor-available-models`, `cursor-catalog-combo-compat`, `cursor-errors-classify`, `cursor-login-pkce`, `cursor-model-effort-suffix-7289`, `cursor-streaming`, `cursor-token-extractor`, `cursor-token-refresh-wiring`, `cursor-usage-fetcher`, `empty-stream-no-content-8649`): 133/133 verdes.
- Corrigido durante a validação: 1 teste novo da própria PR (`cursor-model-effort-suffix-7289.test.ts`, "splits effort off legacy grok- ids") colidia com `CURSOR_MODEL_ALIASES` já mesclado na release (mapeia `grok-4.5-high` → `cursor-grok-4.5-high` antes do fallback legado rodar); ajustado para usar um id não-aliasado (`grok-3-high`) que de fato exercita o fallback — commit `68b58ed`.

Obrigado pela contribuição, @yansigit — feature robusta com boa cobertura de testes.
…g, db-backups tier, uppercase authz bypass, spawn-veto drift) (diegosouzapw#11028)

⭐5 — 4 achados STILL-REAL de advisories de segurança, cada um com TDD (RED→GREEN) e crédito ao reporter original: ACP RCE hardening (resolveVersionProbe), db-backups Tier-2 allowlist, uppercase authz bypass (matcher case-insensitive), spawn-veto drift (chatgpt-web-codex-doctor). typecheck/lint limpos, suíte authz/acp/cors verde. UNSTABLE é o base-red inherited diegosouzapw#9985, já documentado no corpo da PR.
…ams + requestBody (diegosouzapw#10955) (diegosouzapw#11013)

⭐5 — Fix(diegosouzapw#10955): generator de CLI não resolvia $ref em parâmetros do OpenAPI (bug em 20 lugares do spec), PATCH combo sem requestBody. TDD RED→GREEN, gates completos (file-size/complexity/cognitive/changelog/typecheck/lint/docs-all) todos OK. UNSTABLE é o base-red inherited diegosouzapw#9985.
…recovery hint (diegosouzapw#10967, diegosouzapw#10966) (diegosouzapw#11012)

⭐5 — Fix(diegosouzapw#10967,diegosouzapw#10966): combo diag exhausted_connection truncava o UUID por hardcode de provider="unknown"; recovery hint de quota caía em "retry" genérico. TDD RED→GREEN, 149 testes-irmãos verdes. UNSTABLE é o base-red inherited diegosouzapw#9985.
…egosouzapw#10954) (diegosouzapw#11011)

⭐5 — Fix(diegosouzapw#10954): `combo create` via CLI sempre criava combos vazios (models: [] hardcoded, sem flag). Adiciona --models/--model com parser próprio (CLI .mjs sem alias @/). TDD RED→GREEN, 22/22 testes verdes. UNSTABLE é o base-red inherited diegosouzapw#9985.
…0940) (diegosouzapw#11010)

⭐5 — Fix(diegosouzapw#10940): OpenCode config rejeitava modelos sem metadata de catálogo por faltar limit.output (campo obrigatório no schema v1). Agora sempre emite limit com fallback (catálogo → override → 8192). TDD RED→GREEN; teste pré-existente que codificava o bug corrigido. UNSTABLE é o base-red inherited diegosouzapw#9985.
…gin-aware passage (diegosouzapw#11009)

⭐5 — Health-check failures compartilhavam o mesmo caminho de escrita de terminalStatus que requests reais, banindo conexão por health-check falho (403) por um ano. Novo helper origin-aware centraliza toda escrita terminal; health-check só loga, request real desativa como antes. TDD, 6/6 testes, lint/typecheck/cycles OK.
… terminal evicts (diegosouzapw#11008)

⭐5 — Rotação de conta não distinguia falha transitória (quota) de terminal (credencial morta) — ambas só esfriavam e eram retentadas para sempre. Agora markCooldown aceita kind transient/terminal; 3 terminais consecutivos evictam a conta (com fallback para não travar se todas evictadas). TDD, 17/17 testes, lint/typecheck/cycles OK.
…rs use (diegosouzapw#10978)

⭐5 — Cache de reasoning-replay escrevia em toda resposta com reasoning_content, mesmo quando nenhum read-path jamais consumiria (install sem provider de replay). Guard com requiresReasoningReplay() nos dois write-sites, superset seguro do que os readers checam. Testes cobrindo o predicate isoladamente e o wiring real via handleChatCore.
…g that governs nothing (diegosouzapw#10974)

⭐5 — Remove ALLOW_MULTI_CONNECTIONS_PER_COMPAT_NODE, flag morta desde 5b5e21a (que removeu o guard que a lia, resolvendo diegosouzapw#1566). Zero mudança de comportamento; EXPECTED_FEATURE_FLAG_COUNT é o regression guard. Sibling de diegosouzapw#10973 (mesma raiz).
… creation (diegosouzapw#10973)

⭐5 — Remove leitura morta de existingConnections em route.ts, resíduo do mesmo commit 5b5e21a que já removeu o guard que a usava (diegosouzapw#1566). Zero comportamento alterado, 14/14 testes-irmãos verdes. Sibling de diegosouzapw#10974.
…ouzapw#11017) (diegosouzapw#11022)

⭐5 — ENVIRONMENT.md dizia que DEFAULT_RATE_LIMIT_PER_DAY unset = 1000/dia (legado); código e testes desde diegosouzapw#2289 tratam unset/vazio como sem cap implícito. Doc-only, guardado por teste de asserção da tabela. Fecha diegosouzapw#11017.
MeRezaRezaei and others added 28 commits August 22, 2026 18:55
…11045)

Validated on the combined board over tip 80d931a: kimi suites green (executor-kimi-web, kimi-partner-aff-links, token-health-check-kimi 33 assertions), vitest providerPageHeaderKimiPartnerLink green, typecheck:core clean. Audited the diff: only kimi-web switches to the international www.kimi.ai (Connect-RPC base, website, auth hints); kimi-coding / kimi-coding-apikey affiliate links stay on kimi.com as intended — asserted by the updated tests. Owner approved merge without the VPS smoke. Thank you @MeRezaRezaei!
Validated on the combined board over tip 80d931a: quota-redis-store (incl. the KEY_PREFIX derivation test), local-redis-status and rate-limiter-redis-optional green, typecheck:core clean. One pre-merge fix pushed to the branch: docs/reference/ENVIRONMENT.md gained the REDIS_KEY_PREFIX row (env-doc-sync gate requires every .env.example var documented). Board note: the redis tests leave an ioredis retry handle open and hang the runner exit locally — assertions all pass; pre-existing pattern, not from this PR. Thank you @MeRezaRezaei!
Merged with the ENVIRONMENT.md hunk dropped: the tip already documents unset=unlimited for DEFAULT_RATE_LIMIT_PER_DAY via diegosouzapw#11022 (eba58cc), so the docs conflict resolved to the tip text. What lands is the .env.example comment correction, verified against src/shared/utils/apiKeyPolicy.ts::buildDefaultRateLimits — unset/empty → [] (unlimited), malformed → legacy 1000/day windows, explicit 0 → unlimited. Conflict resolution validated on the combined board (env-doc-sync gate green). Thank you @Prajeeth-12!
… calls (diegosouzapw#11085)

Merged after conflict resolution validated on the combined board (50/50 casing tests green, typecheck:core clean). Two pre-merge adjustments on the branch: (1) the utilization route conflict resolved to the tip shape — its asNullableString/displayName version is newer than the branch's; (2) dropped the newly-added src/lib/db/connections.ts, orphaned once the route kept the tip shape (tip already uses getProviderConnectionById) — nothing imported it. The casing fix itself lands intact: non-streaming OpenAI→Claude conversion now restores canonical tool names, identity echoes no longer pin lowercase, and TOOL_RENAME_MAP gained the Task* tools. Fixes the live-reproduced Claude Code 'No such tool available: bash' failures. Thank you @linhdmn — outstanding repro and root-cause writeup!
…iegosouzapw#11035) (diegosouzapw#11157)

Cherry-picked the JSDoc commit onto the current tip (authorship preserved), stripping the stale generated-count noise files. Focused: opencode-v2-config-11070 2/2. Comments now match the 128k fallback shipped in diegosouzapw#11035/diegosouzapw#11054. Thank you @rqzbeh!
…-web (diegosouzapw#11000) (diegosouzapw#11161)

Cherry-picked onto the current tip (authorship preserved), noise files stripped. Pre-merge addition: regenerated the golden snapshot with UPDATE_GOLDEN=1 because the branch's snapshot predated two legitimate tip changes — the dify bare-root from diegosouzapw#11065 and the hackclub removal from diegosouzapw#11123. The regen'd delta contains exactly those two (audited). This also drains a live base-red: provider-translate-path-golden was failing on the pure tip. 3/3 green. Thank you @rqzbeh!
…dal Enter handler (diegosouzapw#10995) (diegosouzapw#11156)

Cherry-picked onto the current tip (authorship preserved), generated-count noise stripped. Pre-merge: file-size baseline rebaselined 1080→1082 with dated annotation (the +2 lines are the Enter-handler isCheckDisabled mirror — owner-requested diegosouzapw#11056 polish; rest is Prettier reflow). Gate green; vitest add-api-key-modal-enter-key 2/2 (jsdom render test). Thank you @rqzbeh!
… oauth start (diegosouzapw#11164) (diegosouzapw#11173)

Cherry-picked onto the current tip (authorship preserved), noise stripped. Focused: oauth-device-flow-11164 green + 9474-claude-code-oauth-mismap neighbor suite green. Device-code endpoint is tried first, camelCase/snake_case fallbacks normalized, no more blank code / 'Visit: undefined'. Fixes diegosouzapw#11164. Thank you @rqzbeh!
… providers (diegosouzapw#11100) (diegosouzapw#11155)

Cherry-picked onto the current tip (authorship preserved), generated-count noise stripped. Two pre-merge adjustments: (1) dropped the unrelated localDb.ts re-export hunk (nothing in this PR uses those symbols); (2) automated security review flagged the blocked-provider list resolving once at server creation — the handler now rebuilds the schema per invocation via the resolver (advertised tools/list schema stays a creation-time snapshot, which is inherent to MCP). Vitest: new runtime-blocked-schema suite 3/3, full MCP __tests__ 118/118; contract suites (mcp-web-search-provider-enum-contract, search-blocked-providers-11100) 6/6; typecheck clean. This closes the residual gap noted when diegosouzapw#11120 was closed. Thank you @rqzbeh!
…antiation (diegosouzapw#11039) (diegosouzapw#11163)

Cherry-picked onto the current tip (authorship preserved), noise stripped. Validated on BOTH runtimes: bun test tests/unit/db-adapters/ 44/44 under the pinned Bun 1.3.14 (native bun:sqlite path), and node --test on the same suite 52 pass / 0 fail / 1 skip (Bun-only adapter skips under Node, as designed). Dockerfile.bun entrypoint now matches the standalone runner shape. Follow-up to diegosouzapw#11039. Thank you @rqzbeh!
…workflow (diegosouzapw#11039) (diegosouzapw#11168)

Cherry-picked onto the current tip (authorship preserved, Dockerfile.bun conflict with the just-merged diegosouzapw#11163 resolved additively — runner-web stage after the new entrypoint). Three pre-merge fixes on the branch: (1) generated-count noise stripped; (2) runner-web stage now returns to the non-root bun user after the apt install (mirrors the Node Dockerfile runner-web re-asserting USER node — the stage previously ended as root); (3) the 6 new build/manifest steps SHA-pinned so the zizmor ratchet stays at 191<=192 findings instead of regressing to 197 (actionlint clean). Workflow YAML parses; runner-base/runner-web targets cross-checked against the Dockerfile stages. Thank you @rqzbeh!
…apw#11122) (diegosouzapw#11154)

Validated on a worktree over the current tip: the red it fixes reproduced exactly as described (media-page-client-browser-bundle red since diegosouzapw#11122 — providerRegistry became reachable from the dashboard client bundle via node:net). Post-fix: bundle test 2/2 green, new ip-parity suite + is-local-provider 7/7, all 7 outboundUrlGuard consumer suites 76/76 (the moved normalizeHost/isPrivateHost keep their re-exports; routing behavior untouched). Thank you @yourspraveen — clean surgical extraction with a pure-JS ipVersion mirroring Node's own regexes.
… test (diegosouzapw#11160)

Validated against code before merge: 159 migration files on disk, 56 free-forever (Hack Club removal), 40 pools — counts verified, not trusted. check:docs-counts exit 0 (HARD failures drained; the 2 remaining soft executors-count notes are pre-existing on the tip) and check:test-discovery OK (orphan moved into the collected tree). These reds came from diegosouzapw#11103/diegosouzapw#11123 merging without the count regen — thanks for sweeping them @yourspraveen!
…uzapw#11071) (diegosouzapw#11165)

Validated on a worktree over the current tip: account-fallback-service 91/91 plus the five sibling lockout suites 24/24. The measurement in the body (40 of 111 passthroughModels providers uncovered on this branch) is the clincher — one lookup via getProviderById().passthroughModels beside the existing checks, closing the diegosouzapw#11071 remainder for shared-registry gateways (port of diegosouzapw#11075 which had only landed on main). Thank you @yourspraveen!
… confirmed) (diegosouzapw#11114)

Validated on the combined batch board over tip 92ef3c7: static gates clean (changelog, file-size, complexity 2624<=2774, cognitive 1182<=1223, dead-code 411<=416), typecheck:core clean, focused tests green.

Companion to diegosouzapw#11113 (consumer side): tokenrouter joins BUILTIN_PROVIDERS_SYSTEM_MUST_BE_FIRST — memory-system-first-6135 suite green. Live-confirmed 400 class documented in the body. Thank you @ggdayup!
…zapw#11162)

Validated on the combined batch board over tip 92ef3c7: static gates clean (changelog, file-size, complexity 2624<=2774, cognitive 1182<=1223, dead-code 411<=416), typecheck:core clean, focused tests green.

Combo without models is now refused at the schema boundary (API 400), the CLI flags it, and openapi.yaml matches the real contract (phantom props removed). combo-* suites + cli-combo-create-models green on the board. Closes diegosouzapw#10954. Thank you @maxmad64bis!
…e) (diegosouzapw#11158)

Validated on the combined batch board over tip 92ef3c7: static gates clean (changelog, file-size, complexity 2624<=2774, cognitive 1182<=1223, dead-code 411<=416), typecheck:core clean, focused tests green.

Empty-envelope 400 (no error field, empty content, finish_reason null) now rotates/retries instead of propagating as success; 200/streaming path never buffered; real-error 400s untouched. account-rotation + new rotation suite 34/34 on the board. Thank you @maxmad64bis!
…stem message (diegosouzapw#11113)

Validated on the combined batch board: purify-system-first suite 4/4, typecheck clean. Pre-merge: file-size baseline gained a frozen entry for contextManager.ts at 1001 (+1, this PR's merge-into-leading-system branch) with a dated annotation — the gate caps unlisted files at 1000. Producer side of the live-confirmed TokenRouter 400 class: no internal path emits a mid-array system message anymore. Thank you @ggdayup — the call-log evidence made this airtight!
…ly output (diegosouzapw#11151)

Merged after conflict resolution against the tip's diegosouzapw#11109 (per-call tool_call tracking): scanOpenAiSseText keeps the per-call finish_reason special-case AND gains reasoningText + literal finishReason; canContinue uses the in-flight predicate with the new reasoning-only-clean-stop escape. One integration fix on the branch: the PR's hallucinatedEmptyStop referenced emittedToolCall, which diegosouzapw#11109 had renamed — the branch now tracks emittedSawToolCall at the emitted level (any tool_call delta, complete or not), preserving the PR's don't-recover-after-tool-calls intent. Chain suites green: stream-continuation-wiring + stream-continuation + stream-recovery-toolcall 29/29. Thank you @maxmad64bis!
…f concatenating it raw (diegosouzapw#11152)

Merged after sibling diegosouzapw#11151 landed: streamRecovery.ts auto-merged byte-identical to the validated combined board; the test-file conflict (both PRs added suites at the same anchor) resolved keeping all 11 tests — diegosouzapw#11151's four clean-stop cases plus this PR's three threshold cases, with the PR's updated partial-tail fixture for the pre-existing overlap test. Full chain green: 32/32 (wiring + continuation + toolcall regression). The documented 8-char overlap threshold ends the silent mid-word gluing. Thank you @maxmad64bis!
diegosouzapw#11190)

* feat(api): structured ?format=json for the self-service usage endpoint

GET /api/usage/om-usage already let any key read its own usage — personal
daily/weekly USD limits and the provider quota snapshot — but only as
text/plain, which a UI cannot parse safely. OmniCopilot issue diegosouzapw#8 asks exactly
for this surface.

Adds ?format=json, returning the ApiKeyUsageLimitStatus + UsageSnapshot the
text is rendered from. Text and JSON share the same collectors
(collectUsageSnapshots, getApiKeyUsageLimitStatus), so the two can never
disagree about a number. The response is a discriminated union: a key without
allowUsageCommand (403) or an invalid key (401) returns
{ allowed:false, error:{message} }, distinct from allowed:true with empty
sections — the state a panel must render as "nothing learned yet", not a
refusal. Text form unchanged; without ?format the contract is untouched.

The endpoint was previously missing from API_REFERENCE.md; it now has a
section documenting both forms, the allowUsageCommand gate, and the
self-service auth model (caller's own key, not requireManagementAuth).

Regression guards in tests/unit/usage-command-json-format.test.ts (4 tests:
json shape, text default preserved, structured 403, sanitized 401 with no
stack trace). Existing internal-usage-command suite still 12/12.

* chore(changelog): correct the fragment to the real PR number (diegosouzapw#11190)

---------

Co-authored-by: Xiangzhe <bakryun0718@proton.me>
…-usage json (diegosouzapw#11192)

* feat(api): structured ?format=json for the self-service usage endpoint

GET /api/usage/om-usage already let any key read its own usage — personal
daily/weekly USD limits and the provider quota snapshot — but only as
text/plain, which a UI cannot parse safely. OmniCopilot issue diegosouzapw#8 asks exactly
for this surface.

Adds ?format=json, returning the ApiKeyUsageLimitStatus + UsageSnapshot the
text is rendered from. Text and JSON share the same collectors
(collectUsageSnapshots, getApiKeyUsageLimitStatus), so the two can never
disagree about a number. The response is a discriminated union: a key without
allowUsageCommand (403) or an invalid key (401) returns
{ allowed:false, error:{message} }, distinct from allowed:true with empty
sections — the state a panel must render as "nothing learned yet", not a
refusal. Text form unchanged; without ?format the contract is untouched.

The endpoint was previously missing from API_REFERENCE.md; it now has a
section documenting both forms, the allowUsageCommand gate, and the
self-service auth model (caller's own key, not requireManagementAuth).

Regression guards in tests/unit/usage-command-json-format.test.ts (4 tests:
json shape, text default preserved, structured 403, sanitized 401 with no
stack trace). Existing internal-usage-command suite still 12/12.

* chore(changelog): correct the fragment to the real PR number (diegosouzapw#11190)

* feat(api): return every connection's snapshot under providers[] in om-usage json

Closes diegosouzapw#11191. buildUsageCommandJson picked a single snapshot via selectUsageSnapshot, so a panel could only ever show one provider. The collector already had them all — the single-pick is a presentation choice for a terminal. The JSON form now also returns the full UsageSnapshot[] alongside the selected provider, so a UI can render Codex / Claude / OpenCode side by side. The text form is untouched.

---------

Co-authored-by: Xiangzhe <bakryun0718@proton.me>
… links through the shortener (diegosouzapw#11196)

- New CheaperInferenceSponsorBanner on the dashboard home, same size/shape as
  KimiSponsorBanner, no version gate (durable partnership). Uses the
  cheaperinference ProviderIcon and the brand green (#31f889) with the dark
  ink CTA (contrast, per colors.ts token).
- CTA points at https://link.omniroute.online/cheaper — the branded short
  link — so clicks land in our Kutt metrics.
- VscodeCopilotBanner CTA now points at https://link.omniroute.online/vsx
  instead of the raw Marketplace URL, for the same reason.
- i18n strings in en + pt (en is the namespace-level fallback for the other
  41 locales).
- Tests: new cheaperInferenceSponsorBanner.test.tsx (render, CTA href, dismiss
  persistence); vscodeCopilotBanner.test.tsx updated to the new CTA URL.

Co-authored-by: Xiangzhe <bakryun0718@proton.me>
… response quality validation (diegosouzapw#11036)

SSE comment lines (OpenRouter keep-alives) and leading whitespace no longer fail the combo quality gate's JSON fallback. First contribution — clean minimal fix with test. Thank you @asorourx, welcome aboard!
…#11041)

Compaction-V2 output now counts as real model output (no synthetic response.failed after response.completed), the Codex SSE filter handles CRLF framing, and terminal detection runs before scan-state bounding. 88/88 stream/readiness suites on the board. Thank you @jackjinke!
Validated on the combined batch board + this branch alone: chat-body-admission + authz/pipeline 65/65, file-size gate green with a dated frozen entry (chatBodyAdmission 1005→1009 — the +4 lease/drain wiring lines, owner-authorized rebaseline). trackRequest was never called, so SIGTERM waitForDrain saw zero in-flight and killed live SSE; leases now hold the drain counter for the stream's lifetime, and the 503 carries Retry-After. Closes diegosouzapw#11015. Thank you @RaviTharuma!
… retrieve tool (diegosouzapw#11084)

Validated on the combined batch board + this branch: ccr-non-mcp-full-prompt-loss + ccr-retrieval-ramp 20/20, file-size gate green with the ccr/index listing (1024, dated annotation — owner-authorized). The callerSupportsCcrRetrieve gate now skips the whole engine for callers whose tools[] cannot reach omniroute_ccr_retrieve — no more 15KB prompt arriving upstream as 112 tokens. Production-measured root cause, textbook TDD. Thank you @HouMinXi!
@diegosouzapw
diegosouzapw merged commit c018bb4 into diegosouzapw:release/v3.8.50 Aug 23, 2026
0 of 3 checks passed
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
Merged after conflict resolution in modelMetadataRegistry.ts: the tip's effortTiers chain (declared efforts → declared tiers → undefined-if-thinking-declared → codex extension) now carries this PR's GLM guard as the final-fallback override — GLM-family models without a provider-declared contract get the authoritative empty tier list instead of generic OpenAI tiers. GLM/ZCode suites 40/40 on the resolved branch. Closes diegosouzapw#10962. Thank you @xz-dev!
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

GLM catalog advertises unsupported reasoning effort tiers across providers