fix(inference): drop the retired DeepSeek V4 Pro from the featured menu - #9626
Conversation
NVIDIA retired deepseek-ai/deepseek-v4-pro on 2026-08-07 and its chat route now returns HTTP 410, but the public featured feed still lists it, so the onboard menu offered a model that fails on first inference. Add it to the featured-catalog retirement deny-list. Signed-off-by: Tinson Lai <tinsonl@nvidia.com>
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Enterprise Run ID: 📒 Files selected for processing (2)
Included review availability: Your plan provides up to 12 included reviews per hour; 11 remain after this review. 📝 WalkthroughWalkthroughThe NVIDIA featured-model parser now excludes ChangesNVIDIA featured-model filtering
Estimated code review effort: 2 (Simple) | ~10 minutes Merge Risk: 🟡 Moderate · up to This change removes an unavailable NVIDIA model from the featured picker and adds a regression test, preventing users from selecting that retired entry. The implementation is localized, but merge readiness remains pending the required sensitive-path review and completion or waiver of the outstanding quality-gate items. Suggested reviewers: 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Comment |
Code Coverage OverviewLanguages: TypeScript TypeScript / code-coverage/pluginThe overall line coverage in commit f5c86b9 in the TypeScript / code-coverage/cliThe overall line coverage in commit f5c86b9 in the Show a line coverage summary of the most impacted files.
Updated |
PR Review Advisor — No blocking findings reportedAdvisor assessment: No blocking advisor findings reported Model lanes
1 additional E2E selection from the second opinionAdvisory only. The primary lane did not select these E2E jobs or targets.
Second-opinion terminology and E2E selections are advisory. Live E2E does not run automatically for pull requests. Since last review: 0 prior items resolved · 0 still apply · 0 new items found 1 semantic terminology decisionTerminology decisions are advisory. They affect the assessment only when a separate finding identifies concrete semantic impact.
E2E guidanceAdvisory only. A maintainer can dispatch the default E2E suite for the commit under review. Recommended E2E: Manual-only E2E: This automated review informs maintainers. Warnings and suggestions do not require a response. A maintainer decides whether to merge. |
prekshivyas
left a comment
There was a problem hiding this comment.
Reviewed the focused deny-list change and regression coverage; no blocking findings.
<!-- markdownlint-disable MD041 --> ## Summary Add the canonical dated changelog entry required before planning the v0.0.112 release. The entry summarizes the 75 merged PRs in `v0.0.111..af56158`, links user-facing themes to published documentation routes, and links every included source PR. ## Changes - Add `docs/changelog/2026-08-20.mdx` with the exact `## v0.0.112` release heading and parser-safe MDX SPDX comment. - Cover managed local inference, onboarding and sandbox lifecycle recovery, messaging continuity, review and release automation, E2E qualification, dependency updates, and cumulative documentation catch-up. - Preserve the documentation skip list and supported-agent matrix; the release entry contains none of the blocked terms or excluded experimental surfaces. ### Source-to-doc mapping - #8620 -> `docs/changelog/2026-08-20.mdx`: Record the LangChain Deep Agents Code 0.1.55 update. - #9192 -> `docs/changelog/2026-08-20.mdx`: Record the OpenShell 0.0.106 update. - #9240 -> `docs/changelog/2026-08-20.mdx`: Record the cold base-image pull heartbeat. - #9412 -> `docs/changelog/2026-08-20.mdx`: Record voice context preservation across sequential turns. - #9483 -> `docs/changelog/2026-08-20.mdx`: Record Ollama model verification through the sandbox endpoint. - #9493 -> `docs/changelog/2026-08-20.mdx`: Record E2E cloud-check wiring coverage. - #9495 -> `docs/changelog/2026-08-20.mdx`: Record Model Router endpoint health validation. - #9534 -> `docs/changelog/2026-08-20.mdx`: Record default-sandbox resolution for tunnel status. - #9537 -> `docs/changelog/2026-08-20.mdx`: Record Linux AMD64 Muse and Lightning profiles. - #9543 -> `docs/changelog/2026-08-20.mdx`: Record corrected network-policy preset examples. - #9545 -> `docs/changelog/2026-08-20.mdx`: Record shared runtime-adapter port validation. - #9578 -> `docs/changelog/2026-08-20.mdx`: Record Portable network creation before host aliases. - #9589 -> `docs/changelog/2026-08-20.mdx`: Record running vLLM profile validation. - #9590 -> `docs/changelog/2026-08-20.mdx`: Record the two-turn atomic advisor review. - #9597 -> `docs/changelog/2026-08-20.mdx`: Record Portable uninstall without host-owned lifecycle resources. - #9605 -> `docs/changelog/2026-08-20.mdx`: Record release automation for an initially empty tag history. - #9607 -> `docs/changelog/2026-08-20.mdx`: Record credential retry navigation. - #9626 -> `docs/changelog/2026-08-20.mdx`: Record retirement of DeepSeek V4 Pro from the featured menu. - #9631 -> `docs/changelog/2026-08-20.mdx`: Record reduction-directed advisor design blockers. - #9632 -> `docs/changelog/2026-08-20.mdx`: Record Portable Ollama under Podman. - #9633 -> `docs/changelog/2026-08-20.mdx`: Record llama.cpp attachment without `/props` model aliases. - #9636 -> `docs/changelog/2026-08-20.mdx`: Record Docker authority independent of terminal state. - #9641 -> `docs/changelog/2026-08-20.mdx`: Record the separate Portable host-gateway subnet. - #9642 -> `docs/changelog/2026-08-20.mdx`: Record cumulative command documentation catch-up. - #9645 -> `docs/changelog/2026-08-20.mdx`: Record removal of completed advisor rollout compatibility. - #9647 -> `docs/changelog/2026-08-20.mdx`: Record diagnostics for OpenShell deletion handoffs. - #9650 -> `docs/changelog/2026-08-20.mdx`: Record OpenClaw pairing settlement after route changes. - #9652 -> `docs/changelog/2026-08-20.mdx`: Record repaired same-turn advisor submissions. - #9653 -> `docs/changelog/2026-08-20.mdx`: Record llama.cpp authority preservation on resume. - #9654 -> `docs/changelog/2026-08-20.mdx`: Record the schema-owned Microsoft Teams webhook field. - #9655 -> `docs/changelog/2026-08-20.mdx`: Record configured managed vLLM ports. - #9656 -> `docs/changelog/2026-08-20.mdx`: Record interrupted managed vLLM installation recovery. - #9660 -> `docs/changelog/2026-08-20.mdx`: Record catalog-owned vLLM profiles and refreshed llama.cpp pins. - #9663 -> `docs/changelog/2026-08-20.mdx`: Record attested LKG production-image requests. - #9664 -> `docs/changelog/2026-08-20.mdx`: Record corrected documented environment-variable handling. - #9665 -> `docs/changelog/2026-08-20.mdx`: Record retired gateway evidence validation. - #9666 -> `docs/changelog/2026-08-20.mdx`: Record Docker authority across terminal sessions. - #9667 -> `docs/changelog/2026-08-20.mdx`: Record contribution intake and product-decision guidance. - #9669 -> `docs/changelog/2026-08-20.mdx`: Record bounded DGX Spark llama.cpp request bodies. - #9670 -> `docs/changelog/2026-08-20.mdx`: Record managed llama.cpp bridge authentication. - #9671 -> `docs/changelog/2026-08-20.mdx`: Record gateway recreation after Docker network loss. - #9672 -> `docs/changelog/2026-08-20.mdx`: Record bounded WSL Ollama host probes. - #9674 -> `docs/changelog/2026-08-20.mdx`: Record cumulative inference and command documentation catch-up. - #9675 -> `docs/changelog/2026-08-20.mdx`: Record Muse Glimmer vLLM image revision handling. - #9676 -> `docs/changelog/2026-08-20.mdx`: Record the grouped CodeQL Actions update. - #9677 -> `docs/changelog/2026-08-20.mdx`: Record the actions/setup-go 7.0.0 update. - #9678 -> `docs/changelog/2026-08-20.mdx`: Record resumable failed llama.cpp cleanup. - #9681 -> `docs/changelog/2026-08-20.mdx`: Record Docker executable injection in the state-mutation harness. - #9683 -> `docs/changelog/2026-08-20.mdx`: Record Windows Docker path fixtures. - #9684 -> `docs/changelog/2026-08-20.mdx`: Record isolated macOS status subprocess cleanup. - #9686 -> `docs/changelog/2026-08-20.mdx`: Record managed-inference catalog compilation for Portable E2E. - #9687 -> `docs/changelog/2026-08-20.mdx`: Record cumulative uninstall documentation catch-up. - #9688 -> `docs/changelog/2026-08-20.mdx`: Record DCode model-selector loading through tsx. - #9689 -> `docs/changelog/2026-08-20.mdx`: Record bounded docs-parity process starts. - #9690 -> `docs/changelog/2026-08-20.mdx`: Record reduced advisor review protocol failures. - #9691 -> `docs/changelog/2026-08-20.mdx`: Record managed llama.cpp bridge cleanup coverage. - #9692 -> `docs/changelog/2026-08-20.mdx`: Record upstream credential rejection diagnostics. - #9693 -> `docs/changelog/2026-08-20.mdx`: Record cumulative managed vLLM documentation catch-up. - #9694 -> `docs/changelog/2026-08-20.mdx`: Record the pinned Portable rootless Podman runtime. - #9695 -> `docs/changelog/2026-08-20.mdx`: Record owned llama.cpp image publication. - #9697 -> `docs/changelog/2026-08-20.mdx`: Record Windows-host Ollama resume behavior. - #9699 -> `docs/changelog/2026-08-20.mdx`: Record the separate trusted Windows path oracle. - #9702 -> `docs/changelog/2026-08-20.mdx`: Record sandbox bridge cleanup coverage. - #9703 -> `docs/changelog/2026-08-20.mdx`: Record hardened Ollama installer downloads. - #9704 -> `docs/changelog/2026-08-20.mdx`: Record supervised dashboard recovery evidence. - #9706 -> `docs/changelog/2026-08-20.mdx`: Record reused model and reasoning health validation. - #9708 -> `docs/changelog/2026-08-20.mdx`: Record fixed local vLLM profile preservation. - #9711 -> `docs/changelog/2026-08-20.mdx`: Record local registry authority in E2E runs. - #9712 -> `docs/changelog/2026-08-20.mdx`: Record Hermes dashboard migration before gateway health. - #9720 -> `docs/changelog/2026-08-20.mdx`: Record default OpenClaw session admission during uninstall. - #9721 -> `docs/changelog/2026-08-20.mdx`: Record MCP credential republishing after policy binding. - #9722 -> `docs/changelog/2026-08-20.mdx`: Record provider republishing after Docker recreation. - #9724 -> `docs/changelog/2026-08-20.mdx`: Record reclamation of dead Shields lifecycle owners. - #9725 -> `docs/changelog/2026-08-20.mdx`: Record fail-closed unscripted onboarding prompts. - #9729 -> `docs/changelog/2026-08-20.mdx`: Record aligned sandbox launch forward ports. ## Type of Change - [ ] Code change (feature, bug fix, or refactor) - [ ] Code change with doc updates - [x] Doc only (prose changes, no code sample modifications) - [ ] Doc only (includes code sample changes) ## Quality Gates - [ ] Tests added or updated for changed behavior - [x] Existing tests cover changed behavior — justification: `test/changelog-docs.test.ts` validates the dated release-entry contract. - [ ] Tests not applicable — justification: - [ ] Sensitive paths changed (security, policy, credentials, preflight, onboarding, inference, runner, sandbox, or messaging) - [ ] Sensitive-path review completed or maintainer-approved waiver recorded — reviewer/approval link/justification: - [ ] Non-success, skipped, or missing CI check accepted by maintainer — check name, approval link, and follow-up issue: ## DGX Station Hardware Evidence - [ ] Tested on DGX Station - Tested commit: Not applicable; documentation-only change. - Station profile/scenario: Not applicable. - Result: Not applicable. - Supporting evidence: Not applicable. ## Verification - [x] PR description includes a `Signed-off-by:` line and every commit appears as `Verified` in GitHub - [x] Normal `pre-commit`, `commit-msg`, and `pre-push` hooks passed, or `npm run validate:pr` passed after refreshing `origin/main` when hooks were skipped or unavailable - [x] Targeted behavior tests pass for the current change set, or tests are marked not applicable above — `npx vitest run test/changelog-docs.test.ts` (7 passed). - [ ] Applicable broad gate passed — `npm test` for broad runtime/test-harness changes; `npm run check` for repo-wide validation/coverage changes — command/result: Not applicable to one prose-only changelog page. - [x] Quality Gates section completed with required justifications or waivers - [x] No secrets, API keys, or credentials committed - [ ] `npm run docs` builds without warnings (doc changes only) — passed with 0 errors and the 2 existing Fern warnings. - [x] Doc pages follow the [style guide](https://github.com/NVIDIA/NemoClaw/blob/main/docs/CONTRIBUTING.md) (doc changes only) - [ ] New doc pages include SPDX header and frontmatter (new pages only) — the parser-safe MDX SPDX comment is present; native changelog pages intentionally do not use frontmatter. --- Signed-off-by: Charan Jagwani <cjagwani@nvidia.com> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Added release notes for v0.0.112. * Documented improvements to managed model runtimes, sandbox recovery, MCP and provider handling, messaging, Shields, and PR Review Advisor. * Added details on release provenance, end-to-end qualification, dependency updates, and documentation alignment. <!-- end of auto-generated comment: release notes by coderabbit.ai -->
<!-- markdownlint-disable MD041 --> ## Summary The NVIDIA Endpoints onboard menu still listed `z-ai/glm-5.2` from NVIDIA's public featured feed, even though the authenticated `/v1/models` catalog for the same account does not serve any GLM ID. Numbered featured choices skip live catalog validation, so a user who picked GLM 5.2 could finish onboard and fail later. This change adds that model ID to the existing featured-feed retirement deny-list. ## Related Issue Fixes #10222 ## Changes - Add `z-ai/glm-5.2` to `RETIRED_NVIDIA_FEATURED_MODEL_IDS` in `src/lib/inference/nvidia-featured-models.ts`. This is a data entry in the existing deny-list, not a new abstraction or fallback path. The deny-list and its policy comment already exist for models whose catalogs outlive their NVIDIA Endpoints routes. - Update `src/lib/inference/nvidia-featured-models.test.ts` so the live-feed fixture used for #9611 also drops GLM 5.2. The rendered menu keeps Nemotron Ultra, Nemotron Super, and Minimax M3. - Assert `CLOUD_MODEL_OPTIONS` still omits `z-ai/glm-5.2` in `src/lib/inference/config.test.ts`. That bundled fallback never contained this ID. - Record the retired featured route in `docs/inference/model-capability-audit.mdx`, matching the DeepSeek V4 Pro row. Deliberately unchanged, with reasons: - `validateNvidiaEndpointModel` still runs only on the `Other...` path. Featured numbered choices stay deny-list filtered, which is the #9626 contract. - OpenRouter continues to share the NVIDIA featured session in `setup-nim-flow.ts`, so this ID also leaves that picker. That coupling already exists for Kimi K2.6 and DeepSeek V4 Pro. - This PR does not intersect the featured feed with authenticated `/v1/models` at menu load time. ## Type of Change - [ ] Code change (feature, bug fix, or refactor) - [x] Code change with doc updates - [ ] Doc only (prose changes, no code sample modifications) - [ ] Doc only (includes code sample changes) ## Quality Gates <!-- Check one tests line. Check other lines when applicable. Add every requested justification or approval reference. --> - [x] Tests added or updated for changed behavior - [ ] Existing tests cover changed behavior — justification: - [ ] Tests not applicable — justification: - [x] Sensitive paths changed (security, policy, credentials, preflight, onboarding, inference, runner, sandbox, or messaging) - [x] Sensitive-path review completed or maintainer-approved waiver recorded — reviewer/approval link/justification: [security review PASS](#10242 (review)) - [ ] Non-success, skipped, or missing CI check accepted by maintainer — check name, approval link, and follow-up issue: ## Documentation Writer Review - [x] Documentation writer reviewed the completed changes - Result: `docs-updated` - Evidence: At current revision `d263fb3abd`, reviewed `docs/inference/model-capability-audit.mdx` against issue #10222 and the retired-model filter and tests. The row links the authenticated WSL and DGX Spark catalog evidence, keeps the retirement scope limited to NVIDIA Endpoints featured selections, and preserves the existing custom-endpoint caveat. Reviewed source blob: `576b8d6bbf`. - Validation: `npm run docs` passed with zero errors and two pre-existing Fern warnings. Generated guide variants completed, and `git diff --check` passed. - Agent: Codex Desktop <!-- docs-review-revision: d263fb3 --> <!-- docs-review-source-blob: 576b8d6 --> ## DGX Station Hardware Evidence <!-- Required only when scripts/prepare-dgx-station-host.sh changes. Maintainers must review the linked evidence before approving or merging. This is human-reviewed evidence, not authenticated hardware provenance. Exceptional bypasses use existing repository governance and must be documented on the PR. --> - [ ] Tested on DGX Station - Tested commit: not applicable, `scripts/prepare-dgx-station-host.sh` is unchanged - Station profile/scenario: not applicable - Result: not applicable - Supporting evidence: not applicable ## Verification <!-- Check each applicable item only when supported by the requested evidence. Run targeted tests once per relevant change set and rerun after later edits or hook autofixes that can affect the tested behavior. Do not rerun hook-covered checks. --> - [x] PR description includes a `Signed-off-by:` line and every commit appears as `Verified` in GitHub - [x] Normal `pre-commit`, `commit-msg`, and `pre-push` hooks passed, or `npm run validate:pr` passed after refreshing `origin/main` when hooks were skipped or unavailable - [x] Targeted behavior tests pass for the current change set, or tests are marked not applicable above — command/result or justification: After `npm run build:cli`, `npx vitest run --project cli src/lib/inference/config.test.ts src/lib/inference/nvidia-featured-models.test.ts` passed 109 tests across 2 files. `npm run validate:pr` also passed after the main-branch refresh. - [ ] Applicable broad gate passed — `npm test` for broad runtime/test-harness changes; `npm run check` for repo-wide validation/coverage changes — command/result: Not applicable to one deny-list entry, its focused regressions, and one evidence row. Branch-wide `npm run validate:pr` passed. - [x] Quality Gates section completed with required justifications or waivers - [x] No secrets, API keys, or credentials committed - [ ] `npm run docs` builds without warnings (doc changes only) — passed with zero errors; Fern reported two pre-existing warnings. - [x] Doc pages follow the [style guide](https://github.com/NVIDIA/NemoClaw/blob/main/docs/CONTRIBUTING.md) (doc changes only) - [ ] New doc pages include SPDX header and frontmatter (new pages only) --- <!-- DCO sign-off is required in this PR description, and every commit must appear as Verified in GitHub. Run: git config user.name && git config user.email --> Signed-off-by: Rui Luo <ruluo@nvidia.com> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Removed the retired NVIDIA GLM 5.2 model from available cloud and featured model selections. * Prevented the retired model from appearing in authenticated model catalogs or live model options. * Ensured model listings consistently exclude GLM 5.2 across supported NVIDIA integrations. * **Documentation** * Updated the model capability audit to record GLM 5.2’s retirement and exclusion from featured selections. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Rui Luo <ruluo@nvidia.com> Signed-off-by: Apurv Kumaria <akumaria@nvidia.com> Co-authored-by: Apurv Kumaria <akumaria@nvidia.com>
Summary
The onboard wizard listed
deepseek-ai/deepseek-v4-proin the NVIDIA Endpoints featured model menu after NVIDIA retired the model, so a user who picked it completed onboarding and then failed on the first inference call. The wizard builds that menu from NVIDIA's public featured feed, which still advertises the retired entry, so this change adds the model ID to the featured-catalog retirement deny-list. The menu now shows only the live entries.Related Issue
Fixes #9611
Changes
deepseek-ai/deepseek-v4-protoRETIRED_NVIDIA_FEATURED_MODEL_IDSinsrc/lib/inference/nvidia-featured-models.ts. This is a data entry in the existing deny-list, not a new abstraction or fallback path; the deny-list and its policy comment already exist for models whose catalogs outlive their routes.src/lib/inference/nvidia-featured-models.test.tsthat drivesgetNvidiaFeaturedModelPromptOptionswith the current live feed payload and asserts the rendered menu keeps the four live entries and drops the retired one.Deliberately unchanged, with reasons:
isDeepSeekV4ProModelinsrc/lib/inference/openai-probe-models.tsand the matching blueprint payload rule stay. Commit 778811d set this precedent when it retired Kimi K2.6: the retirement is scoped to NVIDIA Endpoints, and those branches still serve custom and compatible routes that host the same model ID. The manualOther...entry path also stays open, and it is already covered becausevalidateNvidiaEndpointModelchecks the authenticated/v1/modelscatalog, which no longer lists the model.CLOUD_MODEL_OPTIONSneeds no edit. It never contained this model, andsrc/lib/inference/config.test.tsalready asserts its absence.it.eachcase inconfig.test.tstitledretires %s only from the NVIDIA Endpoints pickerwas not extended, because it also asserts the model remains inHERMES_PROVIDER_MODEL_OPTIONS, and this model was never in that list.Known side effect, consistent with existing behavior:
src/lib/onboard/setup-nim-flow.tsassignsopenRouterFeaturedModels = nvidiaFeaturedModels, so the OpenRouter picker shares this deny-list and also stops offering the ID.moonshotai/kimi-k2.6is already denied under the same shared session, so this PR introduces no new coupling.Runtime evidence
Collected against
https://integrate.api.nvidia.com/v1on 2026-08-19 with a valid key:GET /v1/models: HTTP 200, 102 models,deepseek-ai/deepseek-v4-proabsent. This matches the count in the issue report.POST /v1/chat/completionswith that model: HTTP 410, body{"type":"about:blank","title":"Gone","status":410,"detail":"The model 'deepseek-ai/deepseek-v4-pro' has reached its end of life on 2026-08-07T09:00:00Z and is no longer available."}. NVIDIA reports the exact retirement timestamp, which confirms the 2026-08-07 date in the issue.GET https://assets.ngc.nvidia.com/products/api-catalog/featured-models.json: HTTP 200, 5 entries, still includingdeepseek-ai/deepseek-v4-pro. The regression test uses this exact payload.Two corrections to the issue report, both minor: the reported menu is produced by the live feed rather than a bundled catalog JSON, and the same shared session means the OpenRouter picker was affected as well, which the report does not mention.
Type of Change
Quality Gates
DGX Station Hardware Evidence
scripts/prepare-dgx-station-host.shis unchangedVerification
Signed-off-by:line and every commit appears asVerifiedin GitHubpre-commit,commit-msg, andpre-pushhooks passed, ornpm run validate:prpassed after refreshingorigin/mainwhen hooks were skipped or unavailablenpx vitest run --project cli src/lib/inference src/lib/onboard/nvidia-featured-model-selection.test.ts src/lib/onboard/setup-nim-flow.test.tspassed 2129 tests across 103 files;npx vitest run --project integration test/onboard-selection.test.tspassed 67 tests;npm run typecheck:cliandnpm run checks:repositorypassed. Reverting only the source change fails the new test and leaves the other 17 in that file passing.npm testfor broad runtime/test-harness changes;npm run checkfor repo-wide validation/coverage changes — command/result: not run; the change is one deny-list entry with no runtime or harness surfacenpm run docsbuilds without warnings (doc changes only)Deferred documentation
docs/inference/model-capability-audit.mdxrecords retired NVIDIA Endpoints routes and carries a row for Kimi K2.6 from the precedent commit, but has no row for this model. That row is proposed forDocs / Post-Merge Catch-Up.test/model-capability-audit-doc.test.tsasserts the state vocabulary, evidence field names, header row, and navigation entries only, so it does not require a row per retired model and deferral cannot redden CI. No other page needs an edit:docs/inference/choose-model.mdxalready states that NVIDIA Endpoints excludes retired choices and corrects catalog lag, anddocs/inference/use-nvidia-endpoints.mdxdocuments the bundled fallback list, which this change does not touch.Signed-off-by: Tinson Lai tinsonl@nvidia.com
Summary by CodeRabbit