[AI-000] strip sampling params for opus 4.7/4.8/fable-5 - #20
Merged
Merged
Conversation
* feat(claude): [AI-000] enable 1m context beta and strip unsupported effort for non-opus claude Add context-1m-2025-08-07 to Anthropic-Beta header to enable 1M context window. Strip top-level effort from requests to non-Opus claude upstreams (haiku/sonnet reject it with 400 invalid_request_error); effort is Opus 4.6+ only. * fix(claude): [AI-000] remove context-1m beta flag The 1M context beta is not enabled on all subscriptions; sending context-1m-2025-08-07 made every claude request (haiku included) fail with 400 'The long context beta is not yet available for this subscription'. Drop the flag — the effort-strip fix is unaffected. * fix(claude): [AI-000] strip context-1m beta flag from all claude requests The 1M-context beta requires a subscription entitlement. Without it, requests 400 with 'The long context beta is not yet available for this subscription'. The flag leaks two ways: the static spoof header, and — critically — the Claude Code header cache (claudeHeaderCache.js), which captures a real client's anthropic-beta (incl. context-1m) and re-injects it into every later claude request, including haiku. Strip it from the final headers in DefaultExecutor so neither path can send it. Verified live (docker): with a context-1m-poisoned header cache, old image fails haiku 10/10; fixed image passes haiku 12/12 with zero 1M errors across opus/ sonnet/haiku. ---------
* feat: [AI-000] capture and aggregate per-project usage stats Add x-project request-tag capture, thread project through saveUsageStats into usageHistory.project column, and aggregate stats.byProject (project| model|provider) so usage rolls up per project with per-model token/cost. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat: [AI-000] add Projects usage dashboard page and nav Add Usage by Project view to UsageStats (project|model token/cost rows), a dedicated /dashboard/projects page locked to that view, and a Projects sidebar nav entry. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore: ignore .claude local harness directory --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
…n to details (#7) * chore: ignore local tool scratch dirs (codegraph, codex-pentest) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(usage): [AI-000] stop double-counting streaming usage across project and Untagged Streaming requests wrote a usageHistory row twice: once via logUsage (no project tag, rolled up under Untagged) and once via saveUsageStats (carrying the x-project tag). The duplicate predates per-project stats but only became visible as a project/Untagged split once requests were tagged. logUsage now only appends to the recent-requests log; saveUsageStats is the single usageHistory writer for streaming and preserves cache/reasoning token detail. Also add a Project column + filter to /dashboard/usage?tab=details: thread the project tag through buildRequestDetail, persist it in a new indexed requestDetails.project column, expose a project query param on the request-details API, and render the column, filter, and drawer field. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore: ignore local tool scratch dirs (codegraph, codex-pentest) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(usage): [AI-000] make /dashboard/projects project-first The Projects page reused the shared UsageStats component and only locked the bottom table to the project view, so the page was still model- and provider-centric: the default sort was rawModel (models led the grouped rows) and ProviderTopology, RecentRequests, and the model usage chart filled the page above the table. Add a projectFocus prop, gated so the main usage dashboard is unchanged. On the Projects page it defaults the sort to projectName so projects lead, and hides the model/provider topology, recent-requests, and chart sections. The overview totals (requests, input/output tokens, cost) and the project-grouped table with its per-project token/cost rollup remain. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * refactor(usage): [AI-000] rename projectFocus prop to isProjectFocused Boolean props must use an is/has/should prefix per meaningful-names.md. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…rts (#10) * chore: ignore local tool scratch dirs (codegraph, codex-pentest) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(apikeys): [AI-000] key rotation with manager fields, email notify, and usage charts API key management: - Add managerEmail, managerName, expiresAt, rotatedAt, lastNotifiedAt to the apiKeys schema (additive, auto-synced) and enforce expiry in validateApiKey. - Add rotateApiKey() generating a fresh key value while keeping metadata, plus a POST /api/keys/[id]/rotate route that rotates and emails the manager the new key. Create/update routes accept the new fields with shared validation. - Endpoint dashboard: create/edit modal collects manager name/email/expiry; key rows show manager, expiry (red when expired), and rotated date; a rotate action with confirmation; a one-time new-key modal reporting email status. Email infrastructure: - Add src/lib/email/mailer.js supporting Resend (HTTP) and SMTP (nodemailer), selected via settings. Secrets are write-only in the settings API (exposed as *_Configured booleans) and never logged. - Add POST /api/settings/email-test and an Email Notifications settings card (provider toggle, from address/name, Resend key or SMTP host/port/auth/TLS, and a send-test action). Visualization: - Add byApiKeyProject aggregation dimension to usageRepo (daily + raw paths). - Add ApiKeyUsageChart: stacked bars per API key, toggle By Model / By Project and Requests / Tokens / Cost. Rendered on the usage dashboard. Tests: unit coverage for key metadata validation and expiry logic. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
… reconciles (#12) Legacy usageDaily rows written before the per-project feature shipped carry day.cost / byModel / byApiKey but omit byProject / byApiKeyProject. In daily-summary periods (7d/30d/60d/all), /dashboard/projects summed day.cost for the Est. Cost card while Usage-by-Project summed byProject, so the project table and the By-Project API key chart undercounted. Today/24h was correct because it recomputes live from usageHistory. usageHistory is the complete authoritative record (no pruning), and sum(usageHistory.cost) per day == day.cost, so re-aggregating is lossless. - extract getLocalDateKey/aggregateEntryToDay into helpers/aggregate.js (leaf module) to break the migration->usageRepo->driver->migrate import cycle - add migration v2 rebuilding usageDaily from usageHistory; passes stored cost through (no pricing recompute) and parsed tokens JSON; leaves history-less daily rows untouched; idempotent - v2 is column-defensive: selects NULL for project when the column predates the feature, since versioned migrations run before the additive column sync - regression test reconciles byProject/byApiKeyProject/byModel/byApiKey with day.cost, plus the pre-project-column upgrade path Co-authored-by: dat <dat@9router.local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
… instead of orphaning (#13) Streaming requests wrote requestDetails twice: a "[Streaming in progress...]" placeholder (0 tokens) and a completion update with real content + tokens. The repo upserts on id, but each save generated its own random streamDetailId, so completion inserted a second row instead of updating the placeholder. The placeholder row stayed stuck at 0 input/output tokens forever — which is what shows up when opening a request like 1780314690566-woddaraym. Generate the id once in buildOnStreamComplete and thread it through chatCore to handleStreamingResponse so both saves share it and the completion upserts the same row. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…t/reload (#11) * chore: ignore local tool scratch dirs (codegraph, codex-pentest) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(usage): [AI-000] add per-project model chart and reload button to projects page Add a stacked bar chart (one bar per project, segmented by model) with a Requests/Tokens/Cost metric toggle, so the projects page shows which model is used most per project and model cost per project in a single chart. It reuses the byProject map already returned by /api/usage/stats — no extra fetch. Also add a manual Reload button that re-fetches stats through the shared fetchStats callback, shown on the project-first view. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(docker): [AI-000] copy nodemailer into runner image for SMTP email Next standalone output traces deps from static imports; nodemailer is loaded via dynamic import() for the SMTP path, so tracing omitted it and the runtime image lacked the module. Copy it explicitly like node-forge and next. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…t, not just clean EOF (#14) The shared-id fix made the placeholder row and the completion update share one id, but onStreamComplete only ran inside the SSE TransformStream's flush(), which WHATWG streams invoke ONLY on clean upstream EOF. When a stream ended abnormally — client disconnect, stall-timeout abort, or upstream socket error — flush() never ran, so the row stayed "[Streaming in progress...]" with 0 tokens forever (e.g. request 1780316467100-r0lewav58). Add a cancel(reason) hook to the SSE TransformStream and route both flush() and cancel() through a once-guarded finalize. cancel() fires on every abnormal termination path (verified: downstream reader.cancel and upstream controller .error both trigger it through the pipeWithDisconnect chain). Interrupted rows are marked status="interrupted" with the reason in the content, and usage is only recorded when the provider actually returned token counts. Co-authored-by: dat <dat@9router.local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…e thinking gap (#15) Claude with extended thinking can stream zero bytes for 60s+ while reasoning before the first token. A proxy in front of 9router (Cloudflare Tunnel, nginx) treats that silence as an idle connection and kills it at ~55-60s, so the request aborts before any token arrives — ttft null, 0 tokens, row finalized as interrupted. Emit an SSE comment line (": 9router-keepalive") on a 10s idle timer from stream start until the first real chunk. Comment lines are ignored by every spec-compliant SSE client (Anthropic/OpenAI SDKs included), so callers see nothing. Timer resets on each upstream chunk and is cleared on finalize. Co-authored-by: dat <dat@9router.local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore: ignore local tool scratch dirs (codegraph, codex-pentest) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(auth): [AI-001] derive request locality from bind address, not Host header Two proven findings from the security assessment shared one root cause: the guard trusted spoofable client headers as proof of network origin. AUTH-VULN-01 (High): isLocalRequest() treated `Host: localhost` as proof of loopback, so a network client reaching a 0.0.0.0-bound instance could call the public LLM API without an API key. Locality now derives from the server bind address (HOSTNAME) — header checks remain only as defense-in-depth CSRF guards. Default CLI bind flipped 0.0.0.0 -> 127.0.0.1 so network exposure is explicit opt-in (--host 0.0.0.0), which then correctly requires auth. AUTH-VULN-02 (Medium): getClientIp() trusted X-Forwarded-For/X-Real-IP from every request, letting an attacker rotate the header to escape the login lockout. These headers are now honored only when NINE_ROUTER_TRUSTED_PROXY is set; otherwise all attempts share one bucket so rotation cannot reset lockout. Regression tests cover spoofed-Host-on-non-loopback and rotated-XFF cases. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
# Conflicts: # open-sse/handlers/chatCore.js # open-sse/utils/stream.js # src/app/api/usage/[connectionId]/route.js
* Fix model test routing for image providers * Fix STT model test routing * Use a valid WAV sample for STT model tests * Harden STT ping input and expand model-test coverage * fix(minimax): Bổ sung MiniMax-M3 + cập nhật Quota Tracker coding/CN Squash-merge PR decolua#1631 (decolua/9router) — chỉ lấy file code + test, bỏ docs. - feat(minimax): add MiniMax-M3 to intl + cn provider models (targetFormat claude) - feat(minimax): add MiniMax-M3 pricing entry - fix(minimax): translate Claude body khi content=null (M3 thinking-only) - fix(minimax): hiển thị quota M-series bucket "general"/"MiniMax-M*" + percent-only - test: minimax usage / model registration / pricing Co-Authored-By: Claude <noreply@anthropic.com> Co-authored-by: Cursor <cursoragent@cursor.com> * fix(providers): restore one-connection guard for compatible/embedding nodes OpenAI/Anthropic Compatible and Custom Embedding nodes allow exactly one connection each. The guards were dropped during the bun:sqlite refactor (v0.4.28), so duplicate POSTs were accepted (201) instead of rejected (400). Restore the per-node existing-connection check in the POST handler. Test: tests/unit/compatible-provider-connections.test.js now passes. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(qoder): fetch latest model + nút import model trên dashboard Merge PR decolua#1642 (decolua/9router) — chỉ lấy code + i18n, bỏ test/docs/package.json. - qoder.js: bỏ guard QODER_MODEL_MAP cứng, resolve model_config qua dynamic API (hỗ trợ qmodel_latest không cần sửa code) - dashboard page: thêm nút "Fetch Qoder Models" tự import model list vào aliases - i18n zh-CN: thêm key cho nút fetch Co-Authored-By: Claude <noreply@anthropic.com> Co-authored-by: Cursor <cursoragent@cursor.com> * fix(kiro): add mappable "auto" model slot for Kiro agent mode Kiro sends modelId "auto" for the main agent turn; without a defaultModels slot getMappedModel returned null and the call leaked to AWS instead of the configured provider. Adds the slot + guard test. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(xiaomi-tokenplan): add Claude-native MiMo V2.5 Pro alias via dedicated executor Add mimo-v2.5-pro-claude alias routing to the Xiaomi TokenPlan Anthropic-compatible /anthropic/v1/messages endpoint. Logic lives in a dedicated XiaomiTokenplanExecutor (config-driven via targetFormat) instead of the shared DefaultExecutor. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(antigravity): add gemini-3.5-flash-extra-low (Low) model - Add gemini-3.5-flash-extra-low across CLI menu, provider models, usage, pricing - Add MITM synonyms (high/medium/extra-low) and split pattern so Low no longer falls through to Medium - Strip models/ prefix in getMappedModel for AG public name normalization Co-authored-by: Cursor <cursoragent@cursor.com> * fix(qoder): allow qmodel_latest model key - Add qmodel_latest to QODER_MODEL_MAP - Expose qmodel_latest in static Qoder provider catalog (qd) - Generalize executor comment so model set does not go stale - Add unit coverage for the new model key + catalog Closes decolua#1638 Co-authored-by: Cursor <cursoragent@cursor.com> * fix(antigravity): passthrough tab-autocomplete + mark default agent slot mandatory MODEL_NO_MAP guard never re-routes Antigravity tab-autocomplete (tab_* models) so latency-critical inline completion stays native. Flags gemini-3.5-flash-low (agent/Default) as mandatory in the dashboard. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(codex): durable OAuth refresh lifecycle Add shared OAuth credential lifecycle manager with provider-aware refresh decisions. Implement CodexExecutor.refreshCredentials so 401/403 retry refresh works for Codex, track lastRefreshAt and refresh before the upstream stale-token window, preserve omitted idToken, and add per-connection single-flight refresh to avoid refresh-token rotation races. Merged from PR decolua#1664. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(mitm): Kiro binary EventStream crash + add models & TTS tool filtering - server.js: isBinaryData() skips binary AWS EventStream bodies (fix JSON parse crash) - kiro.js: isBinaryEventStream detection + migrate to pipeTransformedEventStream pipeline - base.js: add pipeTransformedSSE / pipeTransformedEventStream helpers - chatCore.js: filter tool messages + tools for TTS models via getModelType() - providerModels.js: add getModelType() - cliTools.js: add gpt-5-mini (Copilot), glm-5 & minimax-m2.5 (Kiro) Co-authored-by: Cursor <cursoragent@cursor.com> * docs: add Russian README, remove unused testFromFile script Co-authored-by: Cursor <cursoragent@cursor.com> * test(translator): add data-driven coverage, bug-exposing cases, and real provider smoke - matrix.js generates N providers x M models from PROVIDER_MODELS - coverage-all-models + format-roundtrip for structural/semantic checks - bugs-* files expose known translation issues via it.fails - real/smoke-providers runs full handleChatCore path against live providers (RUN_REAL=1), concurrent - registerAll.js eagerly imports translators (Vitest ESM require fix) - vitest.config: array aliases for subpaths + maxConcurrency 60 Co-authored-by: Cursor <cursoragent@cursor.com> * fix(kiro): handle 400 on tool-bearing history without client tools Kiro requires a non-empty currentMessage tools array whenever history references any tool use, else returns "Improperly formed request" (400). Clients trip this by omitting tools on follow-ups after client-side compaction. - flattenToolInteractions(): no client tools -> collapse tool_use/result to text so the "tools required" rule never fires - reconcileOrphanedToolResults(): client tools -> salvage orphaned results as text, keep matched ones, guard co-located tools array - safeJSONParse(): guard tool-call argument parsing against bad JSON - merge consecutive user userInputMessageContext; null-guard currentMessage for assistant-only input Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: Cursor <cursoragent@cursor.com> * fix(minimax): echo reasoning_content on follow-up turns to avoid 400 MiniMax requires reasoning_content echoed back on assistant messages in multi-turn/tool-call conversations. Add minimax and minimax-cn to PROVIDER_RULES (scope all), same fix as DeepSeek (decolua#1543). Co-authored-by: Cursor <cursoragent@cursor.com> * fix(proxy): raise Next client body limit to 128MB (configurable) Avoid 10MB truncation of large /v1 payloads via NINEROUTER_PROXY_CLIENT_MAX_BODY_SIZE (decolua#1529, decolua#1572) Co-authored-by: Cursor <cursoragent@cursor.com> * feat(caveman): add wenyan classical Chinese levels and sync upstream prompts Add wenyan-lite/wenyan/wenyan-ultra levels for max token compression, sync SHARED_EXAMPLES/AUTO_CLARITY/PERSISTENCE across all levels, and expose 3 wenyan buttons in endpoint settings UI. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(i18n): add endpoint exposure notice across multiple languages Added a new translation for the message "Endpoint is exposed without an API key." to various language files, enhancing user awareness regarding API security. This update ensures that users are informed about potential risks associated with unprotected endpoints in their respective languages. * fix(claude): forced tool_choice 400 on cc/ OAuth route convertOpenAIToolChoice mapped {type:"function"} verbatim and cloakClaudeTools left tool_choice.name unsuffixed, both rejected by Claude on the cc/ path. Map forced-function to {type:"tool",name}, allowlist Claude-valid types, and suffix tool_choice.name when it targets a renamed client tool. Fixes decolua#1592 Co-authored-by: Cursor <cursoragent@cursor.com> * fix(tunnel): skip virtual interfaces to prevent false netchange watchdog - Add VIRTUAL_IFACE_REGEX to filter utun/awdl/bridge from network fingerprint - Trust cloudflared/tailscale while process is alive, never kill on force restart Co-authored-by: Cursor <cursoragent@cursor.com> * refactor(dashboard): reorganize menu actions across sidebar/header/profile Move shutdown into header popup + profile, move remote into sidebar above settings, add flag-only language switcher in header, and add language card plus shutdown/logout actions to the profile page. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(codex): harden streaming timeouts + Responses terminal events Raise stall/connect timeouts to 60s (configurable per-provider), accept codex response.done, and always emit a terminal response.failed + [DONE] for Responses passthrough when a stream closes, stalls, or aborts before a terminal event — preventing codex clients from hanging. Co-authored-by: jonathanli12 <jonathanli12@users.noreply.github.com> Co-authored-by: rifuki <rifuki@users.noreply.github.com> Co-authored-by: nguyenha935 <nguyenha935@users.noreply.github.com> Co-authored-by: trananhtung <trananhtung@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com> * feat(endpoint): implement locale-based visibility for wenyan caveman levels * # v0.4.71 (2026-06-06) ## Features - Caveman: add wenyan classical Chinese levels and sync upstream prompts; locale-based visibility on endpoint page - i18n: endpoint exposure notice across multiple languages + Russian README - Antigravity: add gemini-3.5-flash-extra-low (Low) model - xiaomi-tokenplan: add Claude-native MiMo V2.5 Pro alias via dedicated executor - Qoder: fetch latest model + dashboard import-model button (decolua#1642) - MiniMax: add MiniMax-M3 + update Quota Tracker coding/CN (decolua#1631) ## Fixes - Codex: harden streaming timeouts (stall/connect raised to 60s, configurable per-provider), accept `response.done` event, and always emit a terminal `response.failed` + `[DONE]` for Responses passthrough when a stream closes, stalls, or aborts before a terminal event — prevents codex clients from hanging (decolua#1648, decolua#1680, decolua#1688, decolua#1618) - Codex: durable OAuth refresh lifecycle (decolua#1664) - Tunnel: skip virtual interfaces to prevent false netchange watchdog - Claude: fix forced tool_choice 400 on cc/ OAuth route (decolua#1592) - Proxy: raise Next client body limit to 128MB via `NINEROUTER_PROXY_CLIENT_MAX_BODY_SIZE` (decolua#1529, decolua#1572) - MiniMax: echo `reasoning_content` on follow-up turns to avoid 400 (decolua#1543) - Kiro: handle 400 on tool-bearing history without client tools; add mappable "auto" model slot; fix binary EventStream crash + add models & TTS tool filtering - Antigravity: passthrough tab-autocomplete + mark default agent slot mandatory - Qoder: allow `qmodel_latest` model key (decolua#1638) - Providers: restore one-connection guard for compatible/embedding nodes - Model-test: route image/STT probes to their real endpoints, harden STT ping; add opencode-go + xiaomi-tokenplan to connection test (decolua#1576, decolua#1628) ## Improvements - Dashboard: reorganize menu actions across sidebar/header/profile - Translator: add data-driven coverage, bug-exposing cases, and real provider smoke tests * feat(streaming): keep proxied JSON + SSE connections alive through thinking gaps Add downstream keepalive so long upstream silences (reasoning providers can stay quiet 60s+) no longer trip Cloudflare/nginx idle read timeouts: - jsonKeepalive: stream leading JSON whitespace on a timer for forced SSE→JSON and non-streaming proxied responses, then write the final buffered document — whitespace is valid before the JSON value so clients parse normally. - stream.js: reset the SSE comment heartbeat after every upstream chunk instead of stopping it at first token, covering later thinking gaps. - runtimeConfig: make stall timeout + heartbeat/keepalive cadences env configurable; raise STREAM_STALL_TIMEOUT_MS default 60s → 5m. - Dockerfile/.dockerignore: npm ci with lockfile, build cache mount, and --chown at copy time to drop the slow recursive chown -R /app. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: yicone <yicone@gmail.com> Co-authored-by: hodtien <tienhd@scm.devops.vnpt.vn> Co-authored-by: Claude <noreply@anthropic.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: decolua <decoluadt@example.com> Co-authored-by: Delcado19 <85063283+Delcado19@users.noreply.github.com> Co-authored-by: zhangweihong <757058941@qq.com> Co-authored-by: Mr_NoboDy <dnxk@dnxk-pop.tailc90f5d.ts.net> Co-authored-by: AbdoKnbGit <abdoknabo102030@gmail.com> Co-authored-by: therunnas <bcviniciuslc@gmail.com> Co-authored-by: Kevin Le <anhle12892@gmail.com> Co-authored-by: Giao Ho <joutvhu@gmail.com> Co-authored-by: Simon Shi <simonsmh@gmail.com> Co-authored-by: Farhan Usman <farhanusman421@gmail.com> Co-authored-by: arden1601 <corneliusardensatwikahermawan@mail.ugm.ac.id> Co-authored-by: Claude Code <nadimtuhin@gmail.com> Co-authored-by: jonathanli12 <jonathanli12@users.noreply.github.com> Co-authored-by: rifuki <rifuki@users.noreply.github.com> Co-authored-by: nguyenha935 <nguyenha935@users.noreply.github.com> Co-authored-by: trananhtung <trananhtung@users.noreply.github.com> Co-authored-by: dat <dat@9router.local>
# Conflicts: # open-sse/config/runtimeConfig.js
* Fix model test routing for image providers * Fix STT model test routing * Use a valid WAV sample for STT model tests * Harden STT ping input and expand model-test coverage * fix(minimax): Bổ sung MiniMax-M3 + cập nhật Quota Tracker coding/CN Squash-merge PR decolua#1631 (decolua/9router) — chỉ lấy file code + test, bỏ docs. - feat(minimax): add MiniMax-M3 to intl + cn provider models (targetFormat claude) - feat(minimax): add MiniMax-M3 pricing entry - fix(minimax): translate Claude body khi content=null (M3 thinking-only) - fix(minimax): hiển thị quota M-series bucket "general"/"MiniMax-M*" + percent-only - test: minimax usage / model registration / pricing Co-Authored-By: Claude <noreply@anthropic.com> Co-authored-by: Cursor <cursoragent@cursor.com> * fix(providers): restore one-connection guard for compatible/embedding nodes OpenAI/Anthropic Compatible and Custom Embedding nodes allow exactly one connection each. The guards were dropped during the bun:sqlite refactor (v0.4.28), so duplicate POSTs were accepted (201) instead of rejected (400). Restore the per-node existing-connection check in the POST handler. Test: tests/unit/compatible-provider-connections.test.js now passes. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(qoder): fetch latest model + nút import model trên dashboard Merge PR decolua#1642 (decolua/9router) — chỉ lấy code + i18n, bỏ test/docs/package.json. - qoder.js: bỏ guard QODER_MODEL_MAP cứng, resolve model_config qua dynamic API (hỗ trợ qmodel_latest không cần sửa code) - dashboard page: thêm nút "Fetch Qoder Models" tự import model list vào aliases - i18n zh-CN: thêm key cho nút fetch Co-Authored-By: Claude <noreply@anthropic.com> Co-authored-by: Cursor <cursoragent@cursor.com> * fix(kiro): add mappable "auto" model slot for Kiro agent mode Kiro sends modelId "auto" for the main agent turn; without a defaultModels slot getMappedModel returned null and the call leaked to AWS instead of the configured provider. Adds the slot + guard test. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(xiaomi-tokenplan): add Claude-native MiMo V2.5 Pro alias via dedicated executor Add mimo-v2.5-pro-claude alias routing to the Xiaomi TokenPlan Anthropic-compatible /anthropic/v1/messages endpoint. Logic lives in a dedicated XiaomiTokenplanExecutor (config-driven via targetFormat) instead of the shared DefaultExecutor. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(antigravity): add gemini-3.5-flash-extra-low (Low) model - Add gemini-3.5-flash-extra-low across CLI menu, provider models, usage, pricing - Add MITM synonyms (high/medium/extra-low) and split pattern so Low no longer falls through to Medium - Strip models/ prefix in getMappedModel for AG public name normalization Co-authored-by: Cursor <cursoragent@cursor.com> * fix(qoder): allow qmodel_latest model key - Add qmodel_latest to QODER_MODEL_MAP - Expose qmodel_latest in static Qoder provider catalog (qd) - Generalize executor comment so model set does not go stale - Add unit coverage for the new model key + catalog Closes decolua#1638 Co-authored-by: Cursor <cursoragent@cursor.com> * fix(antigravity): passthrough tab-autocomplete + mark default agent slot mandatory MODEL_NO_MAP guard never re-routes Antigravity tab-autocomplete (tab_* models) so latency-critical inline completion stays native. Flags gemini-3.5-flash-low (agent/Default) as mandatory in the dashboard. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(codex): durable OAuth refresh lifecycle Add shared OAuth credential lifecycle manager with provider-aware refresh decisions. Implement CodexExecutor.refreshCredentials so 401/403 retry refresh works for Codex, track lastRefreshAt and refresh before the upstream stale-token window, preserve omitted idToken, and add per-connection single-flight refresh to avoid refresh-token rotation races. Merged from PR decolua#1664. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(mitm): Kiro binary EventStream crash + add models & TTS tool filtering - server.js: isBinaryData() skips binary AWS EventStream bodies (fix JSON parse crash) - kiro.js: isBinaryEventStream detection + migrate to pipeTransformedEventStream pipeline - base.js: add pipeTransformedSSE / pipeTransformedEventStream helpers - chatCore.js: filter tool messages + tools for TTS models via getModelType() - providerModels.js: add getModelType() - cliTools.js: add gpt-5-mini (Copilot), glm-5 & minimax-m2.5 (Kiro) Co-authored-by: Cursor <cursoragent@cursor.com> * docs: add Russian README, remove unused testFromFile script Co-authored-by: Cursor <cursoragent@cursor.com> * test(translator): add data-driven coverage, bug-exposing cases, and real provider smoke - matrix.js generates N providers x M models from PROVIDER_MODELS - coverage-all-models + format-roundtrip for structural/semantic checks - bugs-* files expose known translation issues via it.fails - real/smoke-providers runs full handleChatCore path against live providers (RUN_REAL=1), concurrent - registerAll.js eagerly imports translators (Vitest ESM require fix) - vitest.config: array aliases for subpaths + maxConcurrency 60 Co-authored-by: Cursor <cursoragent@cursor.com> * fix(kiro): handle 400 on tool-bearing history without client tools Kiro requires a non-empty currentMessage tools array whenever history references any tool use, else returns "Improperly formed request" (400). Clients trip this by omitting tools on follow-ups after client-side compaction. - flattenToolInteractions(): no client tools -> collapse tool_use/result to text so the "tools required" rule never fires - reconcileOrphanedToolResults(): client tools -> salvage orphaned results as text, keep matched ones, guard co-located tools array - safeJSONParse(): guard tool-call argument parsing against bad JSON - merge consecutive user userInputMessageContext; null-guard currentMessage for assistant-only input Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: Cursor <cursoragent@cursor.com> * fix(minimax): echo reasoning_content on follow-up turns to avoid 400 MiniMax requires reasoning_content echoed back on assistant messages in multi-turn/tool-call conversations. Add minimax and minimax-cn to PROVIDER_RULES (scope all), same fix as DeepSeek (decolua#1543). Co-authored-by: Cursor <cursoragent@cursor.com> * fix(proxy): raise Next client body limit to 128MB (configurable) Avoid 10MB truncation of large /v1 payloads via NINEROUTER_PROXY_CLIENT_MAX_BODY_SIZE (decolua#1529, decolua#1572) Co-authored-by: Cursor <cursoragent@cursor.com> * feat(caveman): add wenyan classical Chinese levels and sync upstream prompts Add wenyan-lite/wenyan/wenyan-ultra levels for max token compression, sync SHARED_EXAMPLES/AUTO_CLARITY/PERSISTENCE across all levels, and expose 3 wenyan buttons in endpoint settings UI. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(i18n): add endpoint exposure notice across multiple languages Added a new translation for the message "Endpoint is exposed without an API key." to various language files, enhancing user awareness regarding API security. This update ensures that users are informed about potential risks associated with unprotected endpoints in their respective languages. * fix(claude): forced tool_choice 400 on cc/ OAuth route convertOpenAIToolChoice mapped {type:"function"} verbatim and cloakClaudeTools left tool_choice.name unsuffixed, both rejected by Claude on the cc/ path. Map forced-function to {type:"tool",name}, allowlist Claude-valid types, and suffix tool_choice.name when it targets a renamed client tool. Fixes decolua#1592 Co-authored-by: Cursor <cursoragent@cursor.com> * fix(tunnel): skip virtual interfaces to prevent false netchange watchdog - Add VIRTUAL_IFACE_REGEX to filter utun/awdl/bridge from network fingerprint - Trust cloudflared/tailscale while process is alive, never kill on force restart Co-authored-by: Cursor <cursoragent@cursor.com> * refactor(dashboard): reorganize menu actions across sidebar/header/profile Move shutdown into header popup + profile, move remote into sidebar above settings, add flag-only language switcher in header, and add language card plus shutdown/logout actions to the profile page. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(codex): harden streaming timeouts + Responses terminal events Raise stall/connect timeouts to 60s (configurable per-provider), accept codex response.done, and always emit a terminal response.failed + [DONE] for Responses passthrough when a stream closes, stalls, or aborts before a terminal event — preventing codex clients from hanging. Co-authored-by: jonathanli12 <jonathanli12@users.noreply.github.com> Co-authored-by: rifuki <rifuki@users.noreply.github.com> Co-authored-by: nguyenha935 <nguyenha935@users.noreply.github.com> Co-authored-by: trananhtung <trananhtung@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com> * feat(endpoint): implement locale-based visibility for wenyan caveman levels * # v0.4.71 (2026-06-06) ## Features - Caveman: add wenyan classical Chinese levels and sync upstream prompts; locale-based visibility on endpoint page - i18n: endpoint exposure notice across multiple languages + Russian README - Antigravity: add gemini-3.5-flash-extra-low (Low) model - xiaomi-tokenplan: add Claude-native MiMo V2.5 Pro alias via dedicated executor - Qoder: fetch latest model + dashboard import-model button (decolua#1642) - MiniMax: add MiniMax-M3 + update Quota Tracker coding/CN (decolua#1631) ## Fixes - Codex: harden streaming timeouts (stall/connect raised to 60s, configurable per-provider), accept `response.done` event, and always emit a terminal `response.failed` + `[DONE]` for Responses passthrough when a stream closes, stalls, or aborts before a terminal event — prevents codex clients from hanging (decolua#1648, decolua#1680, decolua#1688, decolua#1618) - Codex: durable OAuth refresh lifecycle (decolua#1664) - Tunnel: skip virtual interfaces to prevent false netchange watchdog - Claude: fix forced tool_choice 400 on cc/ OAuth route (decolua#1592) - Proxy: raise Next client body limit to 128MB via `NINEROUTER_PROXY_CLIENT_MAX_BODY_SIZE` (decolua#1529, decolua#1572) - MiniMax: echo `reasoning_content` on follow-up turns to avoid 400 (decolua#1543) - Kiro: handle 400 on tool-bearing history without client tools; add mappable "auto" model slot; fix binary EventStream crash + add models & TTS tool filtering - Antigravity: passthrough tab-autocomplete + mark default agent slot mandatory - Qoder: allow `qmodel_latest` model key (decolua#1638) - Providers: restore one-connection guard for compatible/embedding nodes - Model-test: route image/STT probes to their real endpoints, harden STT ping; add opencode-go + xiaomi-tokenplan to connection test (decolua#1576, decolua#1628) ## Improvements - Dashboard: reorganize menu actions across sidebar/header/profile - Translator: add data-driven coverage, bug-exposing cases, and real provider smoke tests * chore: add zero-downtime deploy + rollback scripts and maintenance image - deploy.sh: rebuild app image, snapshot current image as 9router:previous BEFORE the build (re-tag tag->tag; tagging the running container's digest fails on the containerd image store once the build GCs old layers), swap containers behind a maintenance page, health-check, auto-restore the maintenance page + print rollback command on failure. - rollback.sh: recreate the app container from 9router:previous (or an explicit tag) with the same volume/port/DATA_DIR, behind the maintenance page, then health-check. - scripts/maintenance/: tiny standalone container that holds host port 20128 during the swap — 503 + Retry-After (JSON) for API clients, HTML page for browsers. - .gitignore: un-ignore deploy.sh (the deploy*.sh glob otherwise hides it). Both scripts pin DATA_DIR=/app/data to the 9router-data volume so usage data survives container recreation. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: yicone <yicone@gmail.com> Co-authored-by: hodtien <tienhd@scm.devops.vnpt.vn> Co-authored-by: Claude <noreply@anthropic.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: decolua <decoluadt@example.com> Co-authored-by: Delcado19 <85063283+Delcado19@users.noreply.github.com> Co-authored-by: zhangweihong <757058941@qq.com> Co-authored-by: Mr_NoboDy <dnxk@dnxk-pop.tailc90f5d.ts.net> Co-authored-by: AbdoKnbGit <abdoknabo102030@gmail.com> Co-authored-by: therunnas <bcviniciuslc@gmail.com> Co-authored-by: Kevin Le <anhle12892@gmail.com> Co-authored-by: Giao Ho <joutvhu@gmail.com> Co-authored-by: Simon Shi <simonsmh@gmail.com> Co-authored-by: Farhan Usman <farhanusman421@gmail.com> Co-authored-by: arden1601 <corneliusardensatwikahermawan@mail.ugm.ac.id> Co-authored-by: Claude Code <nadimtuhin@gmail.com> Co-authored-by: jonathanli12 <jonathanli12@users.noreply.github.com> Co-authored-by: rifuki <rifuki@users.noreply.github.com> Co-authored-by: nguyenha935 <nguyenha935@users.noreply.github.com> Co-authored-by: trananhtung <trananhtung@users.noreply.github.com> Co-authored-by: dat <dat@9router.local>
The model-usage-by-project chart on /dashboard/projects rendered a bar for every project, including ones with non-zero tokens but no configured pricing (or any other metric that summed to zero). Those rows show as flat /usr/bin/bash.00 bars and add visual noise without any actionable signal. Track the underlying byProject entries alongside the metric-aggregated project rows, then drop any project whose total requests, total tokens (prompt + completion), and total cost are all > 0. The order-by-busiest metric sort is unchanged. When every project is filtered out, the existing empty-state copy remains the right message.
# Conflicts: # Dockerfile # src/app/(dashboard)/dashboard/usage/components/ProviderLimits/utils.js # src/dashboardGuard.js # src/lib/auth/loginLimiter.js # src/shared/components/UsageStats.js
Opus 4.7+, Opus 4.8, and Fable 5 reject temperature/top_p/top_k with a
400 invalid_request_error ("`temperature` is deprecated"). Claude Code
sends temperature on every request, so the native claude→claude
passthrough path forwarded it straight through and these models 400'd.
Strip the sampling params at the final-body dispatch point (alongside the
existing effort strip), so it covers both the passthrough body and
translated bodies. Version-scoped: Opus 4.6 and earlier, sonnet, and
haiku still accept these params and are left untouched. Matches against
both the requested model and the resolved upstream id so aliases that map
to an opus-4-8 upstream are also caught.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
datj9-reader
approved these changes
Jun 15, 2026
datj9
force-pushed
the
fix/AI-000-strip-sampling-params-opus48
branch
from
June 15, 2026 21:47
63f0f26 to
97cd383
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Requests to
claude-opus-4-8(and any opus-4-7 / fable-5 upstream) fail with:Opus 4.7+, Opus 4.8, and Fable 5 removed the sampling parameters (
temperature,top_p,top_k) — sending any of them returns a 400. Claude Code attachestemperatureto every request, and 9router's nativeclaude → claudepassthrough path forwards the body lossless, so the param reached Anthropic unmodified and the request 400'd.Fix
Strip
temperature/top_p/top_kat the final-body dispatch point inhandleChatCore, right alongside the existingeffortstrip — so it covers both the passthrough body and translated bodies in one place.The strip is version-scoped: only
opus-4-(7|8)andfable-5reject these params. Opus 4.6 and earlier, sonnet, and haiku still accept them and are left untouched. The match runs against both the requested model and the resolved upstream id, so a custom alias that maps to an opus-4-8 upstream is also caught.Testing
max_tokensetc. intact).npm run build→ ✓ compiled successfully.🤖 Generated with Claude Code