Skip to content

[AI-000] strip sampling params for opus 4.7/4.8/fable-5 - #20

Merged
datj9 merged 24 commits into
masterfrom
fix/AI-000-strip-sampling-params-opus48
Jun 15, 2026
Merged

datj9 merged 24 commits into
masterfrom
fix/AI-000-strip-sampling-params-opus48

Conversation

@datj9

@datj9 datj9 commented Jun 14, 2026

Copy link
Copy Markdown
Owner

Problem

Requests to claude-opus-4-8 (and any opus-4-7 / fable-5 upstream) fail with:

API Error: 400 [claude/claude-opus-4-8] [400]: {"type":"error","error":{"type":"invalid_request_error","message":"`temperature` is deprecated...

Opus 4.7+, Opus 4.8, and Fable 5 removed the sampling parameters (temperature, top_p, top_k) — sending any of them returns a 400. Claude Code attaches temperature to every request, and 9router's native claude → claude passthrough path forwards the body lossless, so the param reached Anthropic unmodified and the request 400'd.

Fix

Strip temperature / top_p / top_k at the final-body dispatch point in handleChatCore, right alongside the existing effort strip — so it covers both the passthrough body and translated bodies in one place.

The strip is version-scoped: only opus-4-(7|8) and fable-5 reject these params. Opus 4.6 and earlier, sonnet, and haiku still accept them and are left untouched. The match runs against both the requested model and the resolved upstream id, so a custom alias that maps to an opus-4-8 upstream is also caught.

const rejectsSamplingParams = (...models) =>
  models.some((model = "") => /opus-4-(7|8)|fable-5/.test(model));

Testing

  • Unit-level: verified the model matcher (opus-4-8/4-7/fable-5 → strip; opus-4-6/4-5/sonnet/haiku/empty → keep) and the strip logic (removes only the present sampling params, leaves max_tokens etc. intact).
  • npm run build → ✓ compiled successfully.
  • Deployed to the self-hosted Docker host and confirmed the container is healthy on :20128 with the fix present in the running image.

🤖 Generated with Claude Code

datj9 and others added 24 commits May 31, 2026 12:04
* feat(claude): [AI-000] enable 1m context beta and strip unsupported effort for non-opus claude

Add context-1m-2025-08-07 to Anthropic-Beta header to enable 1M context window.
Strip top-level effort from requests to non-Opus claude upstreams (haiku/sonnet
reject it with 400 invalid_request_error); effort is Opus 4.6+ only.


* fix(claude): [AI-000] remove context-1m beta flag

The 1M context beta is not enabled on all subscriptions; sending
context-1m-2025-08-07 made every claude request (haiku included) fail with
400 'The long context beta is not yet available for this subscription'.
Drop the flag — the effort-strip fix is unaffected.


* fix(claude): [AI-000] strip context-1m beta flag from all claude requests

The 1M-context beta requires a subscription entitlement. Without it, requests
400 with 'The long context beta is not yet available for this subscription'.
The flag leaks two ways: the static spoof header, and — critically — the
Claude Code header cache (claudeHeaderCache.js), which captures a real client's
anthropic-beta (incl. context-1m) and re-injects it into every later claude
request, including haiku. Strip it from the final headers in DefaultExecutor so
neither path can send it.

Verified live (docker): with a context-1m-poisoned header cache, old image fails
haiku 10/10; fixed image passes haiku 12/12 with zero 1M errors across opus/
sonnet/haiku.

---------
* feat: [AI-000] capture and aggregate per-project usage stats

Add x-project request-tag capture, thread project through saveUsageStats
into usageHistory.project column, and aggregate stats.byProject (project|
model|provider) so usage rolls up per project with per-model token/cost.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat: [AI-000] add Projects usage dashboard page and nav

Add Usage by Project view to UsageStats (project|model token/cost rows),
a dedicated /dashboard/projects page locked to that view, and a Projects
sidebar nav entry.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: ignore .claude local harness directory

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
…n to details (#7)

* chore: ignore local tool scratch dirs (codegraph, codex-pentest)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(usage): [AI-000] stop double-counting streaming usage across project and Untagged

Streaming requests wrote a usageHistory row twice: once via logUsage
(no project tag, rolled up under Untagged) and once via saveUsageStats
(carrying the x-project tag). The duplicate predates per-project stats
but only became visible as a project/Untagged split once requests were
tagged. logUsage now only appends to the recent-requests log;
saveUsageStats is the single usageHistory writer for streaming and
preserves cache/reasoning token detail.

Also add a Project column + filter to /dashboard/usage?tab=details:
thread the project tag through buildRequestDetail, persist it in a new
indexed requestDetails.project column, expose a project query param on
the request-details API, and render the column, filter, and drawer field.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore: ignore local tool scratch dirs (codegraph, codex-pentest)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(usage): [AI-000] make /dashboard/projects project-first

The Projects page reused the shared UsageStats component and only locked
the bottom table to the project view, so the page was still model- and
provider-centric: the default sort was rawModel (models led the grouped
rows) and ProviderTopology, RecentRequests, and the model usage chart
filled the page above the table.

Add a projectFocus prop, gated so the main usage dashboard is unchanged.
On the Projects page it defaults the sort to projectName so projects lead,
and hides the model/provider topology, recent-requests, and chart sections.
The overview totals (requests, input/output tokens, cost) and the
project-grouped table with its per-project token/cost rollup remain.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(usage): [AI-000] rename projectFocus prop to isProjectFocused

Boolean props must use an is/has/should prefix per meaningful-names.md.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…rts (#10)

* chore: ignore local tool scratch dirs (codegraph, codex-pentest)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(apikeys): [AI-000] key rotation with manager fields, email notify, and usage charts

API key management:
- Add managerEmail, managerName, expiresAt, rotatedAt, lastNotifiedAt to the
  apiKeys schema (additive, auto-synced) and enforce expiry in validateApiKey.
- Add rotateApiKey() generating a fresh key value while keeping metadata, plus
  a POST /api/keys/[id]/rotate route that rotates and emails the manager the
  new key. Create/update routes accept the new fields with shared validation.
- Endpoint dashboard: create/edit modal collects manager name/email/expiry;
  key rows show manager, expiry (red when expired), and rotated date; a rotate
  action with confirmation; a one-time new-key modal reporting email status.

Email infrastructure:
- Add src/lib/email/mailer.js supporting Resend (HTTP) and SMTP (nodemailer),
  selected via settings. Secrets are write-only in the settings API (exposed as
  *_Configured booleans) and never logged.
- Add POST /api/settings/email-test and an Email Notifications settings card
  (provider toggle, from address/name, Resend key or SMTP host/port/auth/TLS,
  and a send-test action).

Visualization:
- Add byApiKeyProject aggregation dimension to usageRepo (daily + raw paths).
- Add ApiKeyUsageChart: stacked bars per API key, toggle By Model / By Project
  and Requests / Tokens / Cost. Rendered on the usage dashboard.

Tests: unit coverage for key metadata validation and expiry logic.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
… reconciles (#12)

Legacy usageDaily rows written before the per-project feature shipped carry
day.cost / byModel / byApiKey but omit byProject / byApiKeyProject. In
daily-summary periods (7d/30d/60d/all), /dashboard/projects summed day.cost for
the Est. Cost card while Usage-by-Project summed byProject, so the project table
and the By-Project API key chart undercounted. Today/24h was correct because it
recomputes live from usageHistory.

usageHistory is the complete authoritative record (no pruning), and
sum(usageHistory.cost) per day == day.cost, so re-aggregating is lossless.

- extract getLocalDateKey/aggregateEntryToDay into helpers/aggregate.js (leaf
  module) to break the migration->usageRepo->driver->migrate import cycle
- add migration v2 rebuilding usageDaily from usageHistory; passes stored cost
  through (no pricing recompute) and parsed tokens JSON; leaves history-less
  daily rows untouched; idempotent
- v2 is column-defensive: selects NULL for project when the column predates the
  feature, since versioned migrations run before the additive column sync
- regression test reconciles byProject/byApiKeyProject/byModel/byApiKey with
  day.cost, plus the pre-project-column upgrade path

Co-authored-by: dat <dat@9router.local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
… instead of orphaning (#13)

Streaming requests wrote requestDetails twice: a "[Streaming in progress...]"
placeholder (0 tokens) and a completion update with real content + tokens.
The repo upserts on id, but each save generated its own random streamDetailId,
so completion inserted a second row instead of updating the placeholder. The
placeholder row stayed stuck at 0 input/output tokens forever — which is what
shows up when opening a request like 1780314690566-woddaraym.

Generate the id once in buildOnStreamComplete and thread it through chatCore to
handleStreamingResponse so both saves share it and the completion upserts the
same row.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…t/reload (#11)

* chore: ignore local tool scratch dirs (codegraph, codex-pentest)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(usage): [AI-000] add per-project model chart and reload button to projects page

Add a stacked bar chart (one bar per project, segmented by model) with a
Requests/Tokens/Cost metric toggle, so the projects page shows which model
is used most per project and model cost per project in a single chart. It
reuses the byProject map already returned by /api/usage/stats — no extra fetch.

Also add a manual Reload button that re-fetches stats through the shared
fetchStats callback, shown on the project-first view.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(docker): [AI-000] copy nodemailer into runner image for SMTP email

Next standalone output traces deps from static imports; nodemailer is loaded
via dynamic import() for the SMTP path, so tracing omitted it and the runtime
image lacked the module. Copy it explicitly like node-forge and next.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…t, not just clean EOF (#14)

The shared-id fix made the placeholder row and the completion update share one
id, but onStreamComplete only ran inside the SSE TransformStream's flush(), which
WHATWG streams invoke ONLY on clean upstream EOF. When a stream ended abnormally
— client disconnect, stall-timeout abort, or upstream socket error — flush()
never ran, so the row stayed "[Streaming in progress...]" with 0 tokens forever
(e.g. request 1780316467100-r0lewav58).

Add a cancel(reason) hook to the SSE TransformStream and route both flush() and
cancel() through a once-guarded finalize. cancel() fires on every abnormal
termination path (verified: downstream reader.cancel and upstream controller
.error both trigger it through the pipeWithDisconnect chain). Interrupted rows
are marked status="interrupted" with the reason in the content, and usage is
only recorded when the provider actually returned token counts.

Co-authored-by: dat <dat@9router.local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…e thinking gap (#15)

Claude with extended thinking can stream zero bytes for 60s+ while reasoning
before the first token. A proxy in front of 9router (Cloudflare Tunnel, nginx)
treats that silence as an idle connection and kills it at ~55-60s, so the
request aborts before any token arrives — ttft null, 0 tokens, row finalized as
interrupted.

Emit an SSE comment line (": 9router-keepalive") on a 10s idle timer from
stream start until the first real chunk. Comment lines are ignored by every
spec-compliant SSE client (Anthropic/OpenAI SDKs included), so callers see
nothing. Timer resets on each upstream chunk and is cleared on finalize.

Co-authored-by: dat <dat@9router.local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore: ignore local tool scratch dirs (codegraph, codex-pentest)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(auth): [AI-001] derive request locality from bind address, not Host header

Two proven findings from the security assessment shared one root cause: the
guard trusted spoofable client headers as proof of network origin.

AUTH-VULN-01 (High): isLocalRequest() treated `Host: localhost` as proof of
loopback, so a network client reaching a 0.0.0.0-bound instance could call the
public LLM API without an API key. Locality now derives from the server bind
address (HOSTNAME) — header checks remain only as defense-in-depth CSRF
guards. Default CLI bind flipped 0.0.0.0 -> 127.0.0.1 so network exposure is
explicit opt-in (--host 0.0.0.0), which then correctly requires auth.

AUTH-VULN-02 (Medium): getClientIp() trusted X-Forwarded-For/X-Real-IP from
every request, letting an attacker rotate the header to escape the login
lockout. These headers are now honored only when NINE_ROUTER_TRUSTED_PROXY is
set; otherwise all attempts share one bucket so rotation cannot reset lockout.

Regression tests cover spoofed-Host-on-non-loopback and rotated-XFF cases.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
# Conflicts:
#	open-sse/handlers/chatCore.js
#	open-sse/utils/stream.js
#	src/app/api/usage/[connectionId]/route.js
* Fix model test routing for image providers

* Fix STT model test routing

* Use a valid WAV sample for STT model tests

* Harden STT ping input and expand model-test coverage

* fix(minimax): Bổ sung MiniMax-M3 + cập nhật Quota Tracker coding/CN

Squash-merge PR decolua#1631 (decolua/9router) — chỉ lấy file code + test, bỏ docs.

- feat(minimax): add MiniMax-M3 to intl + cn provider models (targetFormat claude)
- feat(minimax): add MiniMax-M3 pricing entry
- fix(minimax): translate Claude body khi content=null (M3 thinking-only)
- fix(minimax): hiển thị quota M-series bucket "general"/"MiniMax-M*" + percent-only
- test: minimax usage / model registration / pricing

Co-Authored-By: Claude <noreply@anthropic.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(providers): restore one-connection guard for compatible/embedding nodes

OpenAI/Anthropic Compatible and Custom Embedding nodes allow exactly one
connection each. The guards were dropped during the bun:sqlite refactor
(v0.4.28), so duplicate POSTs were accepted (201) instead of rejected (400).

Restore the per-node existing-connection check in the POST handler.
Test: tests/unit/compatible-provider-connections.test.js now passes.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(qoder): fetch latest model + nút import model trên dashboard

Merge PR decolua#1642 (decolua/9router) — chỉ lấy code + i18n, bỏ test/docs/package.json.

- qoder.js: bỏ guard QODER_MODEL_MAP cứng, resolve model_config qua dynamic API (hỗ trợ qmodel_latest không cần sửa code)
- dashboard page: thêm nút "Fetch Qoder Models" tự import model list vào aliases
- i18n zh-CN: thêm key cho nút fetch

Co-Authored-By: Claude <noreply@anthropic.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(kiro): add mappable "auto" model slot for Kiro agent mode

Kiro sends modelId "auto" for the main agent turn; without a defaultModels
slot getMappedModel returned null and the call leaked to AWS instead of the
configured provider. Adds the slot + guard test.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(xiaomi-tokenplan): add Claude-native MiMo V2.5 Pro alias via dedicated executor

Add mimo-v2.5-pro-claude alias routing to the Xiaomi TokenPlan Anthropic-compatible
/anthropic/v1/messages endpoint. Logic lives in a dedicated XiaomiTokenplanExecutor
(config-driven via targetFormat) instead of the shared DefaultExecutor.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(antigravity): add gemini-3.5-flash-extra-low (Low) model

- Add gemini-3.5-flash-extra-low across CLI menu, provider models, usage, pricing
- Add MITM synonyms (high/medium/extra-low) and split pattern so Low no longer falls through to Medium
- Strip models/ prefix in getMappedModel for AG public name normalization

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(qoder): allow qmodel_latest model key

- Add qmodel_latest to QODER_MODEL_MAP
- Expose qmodel_latest in static Qoder provider catalog (qd)
- Generalize executor comment so model set does not go stale
- Add unit coverage for the new model key + catalog

Closes decolua#1638

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(antigravity): passthrough tab-autocomplete + mark default agent slot mandatory

MODEL_NO_MAP guard never re-routes Antigravity tab-autocomplete (tab_* models)
so latency-critical inline completion stays native. Flags gemini-3.5-flash-low
(agent/Default) as mandatory in the dashboard.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(codex): durable OAuth refresh lifecycle

Add shared OAuth credential lifecycle manager with provider-aware refresh
decisions. Implement CodexExecutor.refreshCredentials so 401/403 retry
refresh works for Codex, track lastRefreshAt and refresh before the
upstream stale-token window, preserve omitted idToken, and add
per-connection single-flight refresh to avoid refresh-token rotation races.

Merged from PR decolua#1664.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(mitm): Kiro binary EventStream crash + add models & TTS tool filtering

- server.js: isBinaryData() skips binary AWS EventStream bodies (fix JSON parse crash)
- kiro.js: isBinaryEventStream detection + migrate to pipeTransformedEventStream pipeline
- base.js: add pipeTransformedSSE / pipeTransformedEventStream helpers
- chatCore.js: filter tool messages + tools for TTS models via getModelType()
- providerModels.js: add getModelType()
- cliTools.js: add gpt-5-mini (Copilot), glm-5 & minimax-m2.5 (Kiro)

Co-authored-by: Cursor <cursoragent@cursor.com>

* docs: add Russian README, remove unused testFromFile script

Co-authored-by: Cursor <cursoragent@cursor.com>

* test(translator): add data-driven coverage, bug-exposing cases, and real provider smoke

- matrix.js generates N providers x M models from PROVIDER_MODELS
- coverage-all-models + format-roundtrip for structural/semantic checks
- bugs-* files expose known translation issues via it.fails
- real/smoke-providers runs full handleChatCore path against live providers (RUN_REAL=1), concurrent
- registerAll.js eagerly imports translators (Vitest ESM require fix)
- vitest.config: array aliases for subpaths + maxConcurrency 60

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(kiro): handle 400 on tool-bearing history without client tools

Kiro requires a non-empty currentMessage tools array whenever history
references any tool use, else returns "Improperly formed request" (400).
Clients trip this by omitting tools on follow-ups after client-side
compaction.

- flattenToolInteractions(): no client tools -> collapse tool_use/result
  to text so the "tools required" rule never fires
- reconcileOrphanedToolResults(): client tools -> salvage orphaned
  results as text, keep matched ones, guard co-located tools array
- safeJSONParse(): guard tool-call argument parsing against bad JSON
- merge consecutive user userInputMessageContext; null-guard
  currentMessage for assistant-only input

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(minimax): echo reasoning_content on follow-up turns to avoid 400

MiniMax requires reasoning_content echoed back on assistant messages in
multi-turn/tool-call conversations. Add minimax and minimax-cn to
PROVIDER_RULES (scope all), same fix as DeepSeek (decolua#1543).

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(proxy): raise Next client body limit to 128MB (configurable)

Avoid 10MB truncation of large /v1 payloads via NINEROUTER_PROXY_CLIENT_MAX_BODY_SIZE (decolua#1529, decolua#1572)

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(caveman): add wenyan classical Chinese levels and sync upstream prompts

Add wenyan-lite/wenyan/wenyan-ultra levels for max token compression,
sync SHARED_EXAMPLES/AUTO_CLARITY/PERSISTENCE across all levels, and
expose 3 wenyan buttons in endpoint settings UI.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(i18n): add endpoint exposure notice across multiple languages

Added a new translation for the message "Endpoint is exposed without an API key." to various language files, enhancing user awareness regarding API security. This update ensures that users are informed about potential risks associated with unprotected endpoints in their respective languages.

* fix(claude): forced tool_choice 400 on cc/ OAuth route

convertOpenAIToolChoice mapped {type:"function"} verbatim and cloakClaudeTools
left tool_choice.name unsuffixed, both rejected by Claude on the cc/ path.
Map forced-function to {type:"tool",name}, allowlist Claude-valid types, and
suffix tool_choice.name when it targets a renamed client tool.

Fixes decolua#1592

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(tunnel): skip virtual interfaces to prevent false netchange watchdog

- Add VIRTUAL_IFACE_REGEX to filter utun/awdl/bridge from network fingerprint
- Trust cloudflared/tailscale while process is alive, never kill on force restart

Co-authored-by: Cursor <cursoragent@cursor.com>

* refactor(dashboard): reorganize menu actions across sidebar/header/profile

Move shutdown into header popup + profile, move remote into sidebar above
settings, add flag-only language switcher in header, and add language card
plus shutdown/logout actions to the profile page.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(codex): harden streaming timeouts + Responses terminal events

Raise stall/connect timeouts to 60s (configurable per-provider), accept
codex response.done, and always emit a terminal response.failed + [DONE]
for Responses passthrough when a stream closes, stalls, or aborts before
a terminal event — preventing codex clients from hanging.

Co-authored-by: jonathanli12 <jonathanli12@users.noreply.github.com>
Co-authored-by: rifuki <rifuki@users.noreply.github.com>
Co-authored-by: nguyenha935 <nguyenha935@users.noreply.github.com>
Co-authored-by: trananhtung <trananhtung@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(endpoint): implement locale-based visibility for wenyan caveman levels

* # v0.4.71 (2026-06-06)

## Features
- Caveman: add wenyan classical Chinese levels and sync upstream prompts; locale-based visibility on endpoint page
- i18n: endpoint exposure notice across multiple languages + Russian README
- Antigravity: add gemini-3.5-flash-extra-low (Low) model
- xiaomi-tokenplan: add Claude-native MiMo V2.5 Pro alias via dedicated executor
- Qoder: fetch latest model + dashboard import-model button (decolua#1642)
- MiniMax: add MiniMax-M3 + update Quota Tracker coding/CN (decolua#1631)

## Fixes
- Codex: harden streaming timeouts (stall/connect raised to 60s, configurable per-provider), accept `response.done` event, and always emit a terminal `response.failed` + `[DONE]` for Responses passthrough when a stream closes, stalls, or aborts before a terminal event — prevents codex clients from hanging (decolua#1648, decolua#1680, decolua#1688, decolua#1618)
- Codex: durable OAuth refresh lifecycle (decolua#1664)
- Tunnel: skip virtual interfaces to prevent false netchange watchdog
- Claude: fix forced tool_choice 400 on cc/ OAuth route (decolua#1592)
- Proxy: raise Next client body limit to 128MB via `NINEROUTER_PROXY_CLIENT_MAX_BODY_SIZE` (decolua#1529, decolua#1572)
- MiniMax: echo `reasoning_content` on follow-up turns to avoid 400 (decolua#1543)
- Kiro: handle 400 on tool-bearing history without client tools; add mappable "auto" model slot; fix binary EventStream crash + add models & TTS tool filtering
- Antigravity: passthrough tab-autocomplete + mark default agent slot mandatory
- Qoder: allow `qmodel_latest` model key (decolua#1638)
- Providers: restore one-connection guard for compatible/embedding nodes
- Model-test: route image/STT probes to their real endpoints, harden STT ping; add opencode-go + xiaomi-tokenplan to connection test (decolua#1576, decolua#1628)

## Improvements
- Dashboard: reorganize menu actions across sidebar/header/profile
- Translator: add data-driven coverage, bug-exposing cases, and real provider smoke tests

* feat(streaming): keep proxied JSON + SSE connections alive through thinking gaps

Add downstream keepalive so long upstream silences (reasoning providers can
stay quiet 60s+) no longer trip Cloudflare/nginx idle read timeouts:

- jsonKeepalive: stream leading JSON whitespace on a timer for forced
  SSE→JSON and non-streaming proxied responses, then write the final
  buffered document — whitespace is valid before the JSON value so clients
  parse normally.
- stream.js: reset the SSE comment heartbeat after every upstream chunk
  instead of stopping it at first token, covering later thinking gaps.
- runtimeConfig: make stall timeout + heartbeat/keepalive cadences env
  configurable; raise STREAM_STALL_TIMEOUT_MS default 60s → 5m.
- Dockerfile/.dockerignore: npm ci with lockfile, build cache mount, and
  --chown at copy time to drop the slow recursive chown -R /app.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: yicone <yicone@gmail.com>
Co-authored-by: hodtien <tienhd@scm.devops.vnpt.vn>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: decolua <decoluadt@example.com>
Co-authored-by: Delcado19 <85063283+Delcado19@users.noreply.github.com>
Co-authored-by: zhangweihong <757058941@qq.com>
Co-authored-by: Mr_NoboDy <dnxk@dnxk-pop.tailc90f5d.ts.net>
Co-authored-by: AbdoKnbGit <abdoknabo102030@gmail.com>
Co-authored-by: therunnas <bcviniciuslc@gmail.com>
Co-authored-by: Kevin Le <anhle12892@gmail.com>
Co-authored-by: Giao Ho <joutvhu@gmail.com>
Co-authored-by: Simon Shi <simonsmh@gmail.com>
Co-authored-by: Farhan Usman <farhanusman421@gmail.com>
Co-authored-by: arden1601 <corneliusardensatwikahermawan@mail.ugm.ac.id>
Co-authored-by: Claude Code <nadimtuhin@gmail.com>
Co-authored-by: jonathanli12 <jonathanli12@users.noreply.github.com>
Co-authored-by: rifuki <rifuki@users.noreply.github.com>
Co-authored-by: nguyenha935 <nguyenha935@users.noreply.github.com>
Co-authored-by: trananhtung <trananhtung@users.noreply.github.com>
Co-authored-by: dat <dat@9router.local>
# Conflicts:
#	open-sse/config/runtimeConfig.js
* Fix model test routing for image providers

* Fix STT model test routing

* Use a valid WAV sample for STT model tests

* Harden STT ping input and expand model-test coverage

* fix(minimax): Bổ sung MiniMax-M3 + cập nhật Quota Tracker coding/CN

Squash-merge PR decolua#1631 (decolua/9router) — chỉ lấy file code + test, bỏ docs.

- feat(minimax): add MiniMax-M3 to intl + cn provider models (targetFormat claude)
- feat(minimax): add MiniMax-M3 pricing entry
- fix(minimax): translate Claude body khi content=null (M3 thinking-only)
- fix(minimax): hiển thị quota M-series bucket "general"/"MiniMax-M*" + percent-only
- test: minimax usage / model registration / pricing

Co-Authored-By: Claude <noreply@anthropic.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(providers): restore one-connection guard for compatible/embedding nodes

OpenAI/Anthropic Compatible and Custom Embedding nodes allow exactly one
connection each. The guards were dropped during the bun:sqlite refactor
(v0.4.28), so duplicate POSTs were accepted (201) instead of rejected (400).

Restore the per-node existing-connection check in the POST handler.
Test: tests/unit/compatible-provider-connections.test.js now passes.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(qoder): fetch latest model + nút import model trên dashboard

Merge PR decolua#1642 (decolua/9router) — chỉ lấy code + i18n, bỏ test/docs/package.json.

- qoder.js: bỏ guard QODER_MODEL_MAP cứng, resolve model_config qua dynamic API (hỗ trợ qmodel_latest không cần sửa code)
- dashboard page: thêm nút "Fetch Qoder Models" tự import model list vào aliases
- i18n zh-CN: thêm key cho nút fetch

Co-Authored-By: Claude <noreply@anthropic.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(kiro): add mappable "auto" model slot for Kiro agent mode

Kiro sends modelId "auto" for the main agent turn; without a defaultModels
slot getMappedModel returned null and the call leaked to AWS instead of the
configured provider. Adds the slot + guard test.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(xiaomi-tokenplan): add Claude-native MiMo V2.5 Pro alias via dedicated executor

Add mimo-v2.5-pro-claude alias routing to the Xiaomi TokenPlan Anthropic-compatible
/anthropic/v1/messages endpoint. Logic lives in a dedicated XiaomiTokenplanExecutor
(config-driven via targetFormat) instead of the shared DefaultExecutor.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(antigravity): add gemini-3.5-flash-extra-low (Low) model

- Add gemini-3.5-flash-extra-low across CLI menu, provider models, usage, pricing
- Add MITM synonyms (high/medium/extra-low) and split pattern so Low no longer falls through to Medium
- Strip models/ prefix in getMappedModel for AG public name normalization

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(qoder): allow qmodel_latest model key

- Add qmodel_latest to QODER_MODEL_MAP
- Expose qmodel_latest in static Qoder provider catalog (qd)
- Generalize executor comment so model set does not go stale
- Add unit coverage for the new model key + catalog

Closes decolua#1638

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(antigravity): passthrough tab-autocomplete + mark default agent slot mandatory

MODEL_NO_MAP guard never re-routes Antigravity tab-autocomplete (tab_* models)
so latency-critical inline completion stays native. Flags gemini-3.5-flash-low
(agent/Default) as mandatory in the dashboard.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(codex): durable OAuth refresh lifecycle

Add shared OAuth credential lifecycle manager with provider-aware refresh
decisions. Implement CodexExecutor.refreshCredentials so 401/403 retry
refresh works for Codex, track lastRefreshAt and refresh before the
upstream stale-token window, preserve omitted idToken, and add
per-connection single-flight refresh to avoid refresh-token rotation races.

Merged from PR decolua#1664.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(mitm): Kiro binary EventStream crash + add models & TTS tool filtering

- server.js: isBinaryData() skips binary AWS EventStream bodies (fix JSON parse crash)
- kiro.js: isBinaryEventStream detection + migrate to pipeTransformedEventStream pipeline
- base.js: add pipeTransformedSSE / pipeTransformedEventStream helpers
- chatCore.js: filter tool messages + tools for TTS models via getModelType()
- providerModels.js: add getModelType()
- cliTools.js: add gpt-5-mini (Copilot), glm-5 & minimax-m2.5 (Kiro)

Co-authored-by: Cursor <cursoragent@cursor.com>

* docs: add Russian README, remove unused testFromFile script

Co-authored-by: Cursor <cursoragent@cursor.com>

* test(translator): add data-driven coverage, bug-exposing cases, and real provider smoke

- matrix.js generates N providers x M models from PROVIDER_MODELS
- coverage-all-models + format-roundtrip for structural/semantic checks
- bugs-* files expose known translation issues via it.fails
- real/smoke-providers runs full handleChatCore path against live providers (RUN_REAL=1), concurrent
- registerAll.js eagerly imports translators (Vitest ESM require fix)
- vitest.config: array aliases for subpaths + maxConcurrency 60

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(kiro): handle 400 on tool-bearing history without client tools

Kiro requires a non-empty currentMessage tools array whenever history
references any tool use, else returns "Improperly formed request" (400).
Clients trip this by omitting tools on follow-ups after client-side
compaction.

- flattenToolInteractions(): no client tools -> collapse tool_use/result
  to text so the "tools required" rule never fires
- reconcileOrphanedToolResults(): client tools -> salvage orphaned
  results as text, keep matched ones, guard co-located tools array
- safeJSONParse(): guard tool-call argument parsing against bad JSON
- merge consecutive user userInputMessageContext; null-guard
  currentMessage for assistant-only input

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(minimax): echo reasoning_content on follow-up turns to avoid 400

MiniMax requires reasoning_content echoed back on assistant messages in
multi-turn/tool-call conversations. Add minimax and minimax-cn to
PROVIDER_RULES (scope all), same fix as DeepSeek (decolua#1543).

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(proxy): raise Next client body limit to 128MB (configurable)

Avoid 10MB truncation of large /v1 payloads via NINEROUTER_PROXY_CLIENT_MAX_BODY_SIZE (decolua#1529, decolua#1572)

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(caveman): add wenyan classical Chinese levels and sync upstream prompts

Add wenyan-lite/wenyan/wenyan-ultra levels for max token compression,
sync SHARED_EXAMPLES/AUTO_CLARITY/PERSISTENCE across all levels, and
expose 3 wenyan buttons in endpoint settings UI.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(i18n): add endpoint exposure notice across multiple languages

Added a new translation for the message "Endpoint is exposed without an API key." to various language files, enhancing user awareness regarding API security. This update ensures that users are informed about potential risks associated with unprotected endpoints in their respective languages.

* fix(claude): forced tool_choice 400 on cc/ OAuth route

convertOpenAIToolChoice mapped {type:"function"} verbatim and cloakClaudeTools
left tool_choice.name unsuffixed, both rejected by Claude on the cc/ path.
Map forced-function to {type:"tool",name}, allowlist Claude-valid types, and
suffix tool_choice.name when it targets a renamed client tool.

Fixes decolua#1592

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(tunnel): skip virtual interfaces to prevent false netchange watchdog

- Add VIRTUAL_IFACE_REGEX to filter utun/awdl/bridge from network fingerprint
- Trust cloudflared/tailscale while process is alive, never kill on force restart

Co-authored-by: Cursor <cursoragent@cursor.com>

* refactor(dashboard): reorganize menu actions across sidebar/header/profile

Move shutdown into header popup + profile, move remote into sidebar above
settings, add flag-only language switcher in header, and add language card
plus shutdown/logout actions to the profile page.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(codex): harden streaming timeouts + Responses terminal events

Raise stall/connect timeouts to 60s (configurable per-provider), accept
codex response.done, and always emit a terminal response.failed + [DONE]
for Responses passthrough when a stream closes, stalls, or aborts before
a terminal event — preventing codex clients from hanging.

Co-authored-by: jonathanli12 <jonathanli12@users.noreply.github.com>
Co-authored-by: rifuki <rifuki@users.noreply.github.com>
Co-authored-by: nguyenha935 <nguyenha935@users.noreply.github.com>
Co-authored-by: trananhtung <trananhtung@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(endpoint): implement locale-based visibility for wenyan caveman levels

* # v0.4.71 (2026-06-06)

## Features
- Caveman: add wenyan classical Chinese levels and sync upstream prompts; locale-based visibility on endpoint page
- i18n: endpoint exposure notice across multiple languages + Russian README
- Antigravity: add gemini-3.5-flash-extra-low (Low) model
- xiaomi-tokenplan: add Claude-native MiMo V2.5 Pro alias via dedicated executor
- Qoder: fetch latest model + dashboard import-model button (decolua#1642)
- MiniMax: add MiniMax-M3 + update Quota Tracker coding/CN (decolua#1631)

## Fixes
- Codex: harden streaming timeouts (stall/connect raised to 60s, configurable per-provider), accept `response.done` event, and always emit a terminal `response.failed` + `[DONE]` for Responses passthrough when a stream closes, stalls, or aborts before a terminal event — prevents codex clients from hanging (decolua#1648, decolua#1680, decolua#1688, decolua#1618)
- Codex: durable OAuth refresh lifecycle (decolua#1664)
- Tunnel: skip virtual interfaces to prevent false netchange watchdog
- Claude: fix forced tool_choice 400 on cc/ OAuth route (decolua#1592)
- Proxy: raise Next client body limit to 128MB via `NINEROUTER_PROXY_CLIENT_MAX_BODY_SIZE` (decolua#1529, decolua#1572)
- MiniMax: echo `reasoning_content` on follow-up turns to avoid 400 (decolua#1543)
- Kiro: handle 400 on tool-bearing history without client tools; add mappable "auto" model slot; fix binary EventStream crash + add models & TTS tool filtering
- Antigravity: passthrough tab-autocomplete + mark default agent slot mandatory
- Qoder: allow `qmodel_latest` model key (decolua#1638)
- Providers: restore one-connection guard for compatible/embedding nodes
- Model-test: route image/STT probes to their real endpoints, harden STT ping; add opencode-go + xiaomi-tokenplan to connection test (decolua#1576, decolua#1628)

## Improvements
- Dashboard: reorganize menu actions across sidebar/header/profile
- Translator: add data-driven coverage, bug-exposing cases, and real provider smoke tests

* chore: add zero-downtime deploy + rollback scripts and maintenance image

- deploy.sh: rebuild app image, snapshot current image as 9router:previous
  BEFORE the build (re-tag tag->tag; tagging the running container's digest
  fails on the containerd image store once the build GCs old layers), swap
  containers behind a maintenance page, health-check, auto-restore the
  maintenance page + print rollback command on failure.
- rollback.sh: recreate the app container from 9router:previous (or an
  explicit tag) with the same volume/port/DATA_DIR, behind the maintenance
  page, then health-check.
- scripts/maintenance/: tiny standalone container that holds host port 20128
  during the swap — 503 + Retry-After (JSON) for API clients, HTML page for
  browsers.
- .gitignore: un-ignore deploy.sh (the deploy*.sh glob otherwise hides it).

Both scripts pin DATA_DIR=/app/data to the 9router-data volume so usage data
survives container recreation.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: yicone <yicone@gmail.com>
Co-authored-by: hodtien <tienhd@scm.devops.vnpt.vn>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: decolua <decoluadt@example.com>
Co-authored-by: Delcado19 <85063283+Delcado19@users.noreply.github.com>
Co-authored-by: zhangweihong <757058941@qq.com>
Co-authored-by: Mr_NoboDy <dnxk@dnxk-pop.tailc90f5d.ts.net>
Co-authored-by: AbdoKnbGit <abdoknabo102030@gmail.com>
Co-authored-by: therunnas <bcviniciuslc@gmail.com>
Co-authored-by: Kevin Le <anhle12892@gmail.com>
Co-authored-by: Giao Ho <joutvhu@gmail.com>
Co-authored-by: Simon Shi <simonsmh@gmail.com>
Co-authored-by: Farhan Usman <farhanusman421@gmail.com>
Co-authored-by: arden1601 <corneliusardensatwikahermawan@mail.ugm.ac.id>
Co-authored-by: Claude Code <nadimtuhin@gmail.com>
Co-authored-by: jonathanli12 <jonathanli12@users.noreply.github.com>
Co-authored-by: rifuki <rifuki@users.noreply.github.com>
Co-authored-by: nguyenha935 <nguyenha935@users.noreply.github.com>
Co-authored-by: trananhtung <trananhtung@users.noreply.github.com>
Co-authored-by: dat <dat@9router.local>
The model-usage-by-project chart on /dashboard/projects rendered a bar
for every project, including ones with non-zero tokens but no configured
pricing (or any other metric that summed to zero). Those rows show as
flat /usr/bin/bash.00 bars and add visual noise without any actionable signal.

Track the underlying byProject entries alongside the metric-aggregated
project rows, then drop any project whose total requests, total tokens
(prompt + completion), and total cost are all > 0. The order-by-busiest
metric sort is unchanged. When every project is filtered out, the
existing empty-state copy remains the right message.
# Conflicts:
#	Dockerfile
#	src/app/(dashboard)/dashboard/usage/components/ProviderLimits/utils.js
#	src/dashboardGuard.js
#	src/lib/auth/loginLimiter.js
#	src/shared/components/UsageStats.js
Opus 4.7+, Opus 4.8, and Fable 5 reject temperature/top_p/top_k with a
400 invalid_request_error ("`temperature` is deprecated"). Claude Code
sends temperature on every request, so the native claude→claude
passthrough path forwarded it straight through and these models 400'd.

Strip the sampling params at the final-body dispatch point (alongside the
existing effort strip), so it covers both the passthrough body and
translated bodies. Version-scoped: Opus 4.6 and earlier, sonnet, and
haiku still accept these params and are left untouched. Matches against
both the requested model and the resolved upstream id so aliases that map
to an opus-4-8 upstream are also caught.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@datj9 datj9 self-assigned this Jun 15, 2026
@datj9
datj9 force-pushed the fix/AI-000-strip-sampling-params-opus48 branch from 63f0f26 to 97cd383 Compare June 15, 2026 21:47
@datj9
datj9 merged commit 8d32414 into master Jun 15, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants