Skip to content

fix(nvidia): drop EOL models, repoint DeepSeek V4 Flash at its live id - #3397

Open
ggfto wants to merge 8 commits into
decolua:masterfrom
ggfto:fix/nvidia-eol-models
Open

ggfto wants to merge 8 commits into
decolua:masterfrom
ggfto:fix/nvidia-eol-models

Conversation

@ggfto

@ggfto ggfto commented Aug 17, 2026 •

Copy link
Copy Markdown

NVIDIA retired three models that the registry still advertises. Each one answers 410 Gone with an explicit end-of-life date, so the catalog offers them and every route to them fails at call time:

minimaxai/minimax-m2.7 EOL 2026-07-27
deepseek-ai/deepseek-v4-pro EOL 2026-08-07
deepseek-ai/deepseek-v4-flash EOL 2026-08-07

deepseek-v4-flash lives on under a dated id — deepseek-v4-flash-0731 — and is verified answering, so it is repointed rather than removed. The other two have no successor in NVIDIA's live catalog and are dropped; minimax-m3 already covers the MiniMax slot.

Scope is deliberately narrow. moonshotai/kimi-k2.6 and nvidia/nemotron-3-ultra-550b-a55b also fail here, but they are still listed in NVIDIA's /v1/models and return "Not found for account", which is per-account access rather than a stale registry entry — left untouched. The bare minimax-m2.7 / deepseek-v4-* ids under codebuddy-cn and poolside in capabilities.js belong to other providers and are also untouched.

Verified against NVIDIA's live /v1/models (102 entries) and by real inference through a local instance.

Co-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com

NVIDIA retired three models that the registry still advertises. Each one
answers 410 Gone with an explicit end-of-life date, so the catalog offers
them and every route to them fails at call time:

  minimaxai/minimax-m2.7        EOL 2026-07-27
  deepseek-ai/deepseek-v4-pro   EOL 2026-08-07
  deepseek-ai/deepseek-v4-flash EOL 2026-08-07

deepseek-v4-flash lives on under a dated id — deepseek-v4-flash-0731 — and
is verified answering, so it is repointed rather than removed. The other
two have no successor in NVIDIA's live catalog and are dropped; minimax-m3
already covers the MiniMax slot.

Scope is deliberately narrow. moonshotai/kimi-k2.6 and
nvidia/nemotron-3-ultra-550b-a55b also fail here, but they are still listed
in NVIDIA's /v1/models and return "Not found for account", which is
per-account access rather than a stale registry entry — left untouched.
The bare minimax-m2.7 / deepseek-v4-* ids under codebuddy-cn and poolside
in capabilities.js belong to other providers and are also untouched.

Verified against NVIDIA's live /v1/models (102 entries) and by real
inference through a local instance.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The workflow pushed to a hardcoded `decolua/9router` Docker Hub namespace,
which a fork cannot write to. Publish to `ghcr.io/<owner>/<repo>` only, so it
works on any fork with the built-in GITHUB_TOKEN and no configured secrets.

Also fixes three ways the tag list could come out wrong:
- `latest` was gated on `is_default_branch`, which is never true on a tag
  push, so a release never actually moved `latest`.
- a manual run from a branch produced no tags at all (semver needs a tag ref),
  which fails the push step; `type=ref` plus an always-on `sha-` tag fixes it.
- the arm64 leg relied on binfmt already being registered on the runner.
…ce disabled models when routing

Two switches in the dashboard did nothing to the traffic they described.

Proxy pool:
- `strictProxy` never reached the request path — `auth.js` dropped it when
  building `providerSpecificData` and `chatCore` never read it — so a strict
  pool fell back to a direct connection on any proxy error, which is exactly
  what strict mode exists to prevent.
- OAuth token refresh went out unproxied everywhere: `chatCore` called
  `refreshCredentials(credentials, log)` without the third argument, and the
  same omission ran through `embeddingsCore`, `imageGenerationCore`, the
  proactive `checkAndRefreshToken`, the connection Test button and the
  translator playground. `proxyOptions` is now threaded through the whole
  refresh chain down to each provider's token endpoint.
- Only `chatCore` implemented pooling at all; embeddings, images, TTS, STT and
  video had no proxy handling whatsoever. The first three now pass
  `proxyOptions` explicitly. TTS adapters and the executors that call bare
  `fetch` (grok-web, perplexity-web, devin-cli) never accepted the argument, so
  they inherit the connection's egress from an ambient AsyncLocalStorage
  context instead; an explicit argument still wins over it. The wrapped region
  is kept tight so local sidecars (headroom, ollama-local) stay direct.
- A pool that is bound but unusable — deleted, or deactivated because a failed
  connectivity test flips `isActive` off — was completely silent while every
  request went out direct. It now says so, and the PROXY log line covers the
  direct case too.

Disabled models:
- `disabledModels` was read only by `/api/models`, `/v1/models` and the model
  picker, so switching a model off just hid it from the UI. Combos saved before
  the change kept the id and kept routing to it, and a direct `/v1` call with
  that model still worked.
- Disabled entries are dropped from a combo before rotation, so they cost no
  round trip, and a combo left with nothing enabled reports that instead of
  failing model by model. A direct request for a disabled model answers 403
  across chat, embeddings, images, TTS, STT and video.

The four test files that stub `global.fetch` now mock `proxyFetch.js` the way
`base-executor-retry.test.js` already did, since those handlers no longer call
the global directly. Full suite: 90 failures before and after, no regressions.
Both platforms were cross-built on a single amd64 runner, so the arm64 leg ran
the entire builder stage under QEMU — compiling better-sqlite3 from source and
running the full Next build emulated. A dispatch run sat in `Build and push`
for over 75 minutes without publishing anything.

Split into a matrix that builds amd64 on ubuntu-latest and arm64 on
ubuntu-24.04-arm, each pushing an untagged image by digest, then a merge job
assembles the tagged multi-arch manifest with `imagetools create`. Native ARM
runners are free for public repositories.

Build cache is now keyed per platform; a single shared ref had the two legs
overwriting each other's layers.
… latest tambem em build manual do branch default

Um pool com strictProxy=true que ficou inutilizavel (deletado ou
desativado por falha de teste) apenas avisava e caia para conexao
direta/legado — exatamente o que o modo estrito existe para impedir.
Agora resolveConnectionProxyConfig lanca erro marcado
(strictProxyRefusal) que o catch propaga em vez de engolir.

docker-publish.yml: alem de tags v*, latest agora tambem e tagueada em
workflow_dispatch do branch default, para que 'docker compose pull
latest' sempre rastreie a build mais recente mesmo sem tag v* — era isso
que faltava quando o compose precisou ser apontado para sha-4c32cc4.
@ggfto ggfto changed the title fix(nvidia): drop EOL models, repoint DeepSeek V4 Flash at its live id fix: proxy pool estrito, undici no runtime e disabled models Sep 22, 2026
@ggfto ggfto changed the title fix: proxy pool estrito, undici no runtime e disabled models fix(nvidia): drop EOL models, repoint DeepSeek V4 Flash at its live id Sep 22, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant