Skip to content

feat: add NVIDIA NIM and MiniMax provider support - #552

Merged
kevincodex1 merged 13 commits into
Twigpine:mainfrom
LifeJiggy:feature/nvidia-minimax-providers
Apr 15, 2026
Merged

kevincodex1 merged 13 commits into
Twigpine:mainfrom
LifeJiggy:feature/nvidia-minimax-providers

Conversation

@LifeJiggy

Copy link
Copy Markdown
Contributor
  • Add nvidia-nim and minimax to --provider CLI flag
  • Add model discovery for NVIDIA NIM (160+ models) and MiniMax
  • Update /model picker to show provider-specific models
  • Fix provider detection in startup banner
  • Update .env.example with new provider options

Supported providers:

Summary

  • what changed
  • why it changed

Impact

  • user-facing impact:
  • developer/maintainer impact:

Testing

  • bun run build
  • bun run smoke
  • focused tests:

Notes

  • provider/model path tested:
  • screenshots attached (if UI changed):
  • follow-up work or known limitations:

- Add nvidia-nim and minimax to --provider CLI flag
- Add model discovery for NVIDIA NIM (160+ models) and MiniMax
- Update /model picker to show provider-specific models
- Fix provider detection in startup banner
- Update .env.example with new provider options

Supported providers:
- NVIDIA NIM: https://integrate.api.nvidia.com/v1
- MiniMax: https://api.minimax.io/v1
@hamzafelix

Copy link
Copy Markdown

we need this merged ASAP @LifeJiggy

@LifeJiggy

Copy link
Copy Markdown
Contributor Author

Hi @maintainers — this PR is ready for review.

It adds NVIDIA NIM and MiniMax providers with model discovery and CLI integration.
No breaking changes — all providers are optional via flags and existing behavior is unaffected.

Local checks pass (bun run build, bun run smoke).
Looks like workflows need approval to run CI — could someone help approve so checks can complete?

Appreciate your time.

@kevincodex1
kevincodex1 requested review from Vasanthdev2004 and gnanam1990 and removed request for Vasanthdev2004 April 13, 2026 13:48
@kevincodex1

Copy link
Copy Markdown
Member

hello @LifeJiggy kindly fix conflicts

@LifeJiggy

Copy link
Copy Markdown
Contributor Author

All conflicts have been resolved.
The PR adds NVIDIA NIM and MiniMax providers while maintaining compatibility with the latest main branch changes.

@Vasanthdev2004 Vasanthdev2004 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: PR #552 — Add NVIDIA NIM and MiniMax provider support

Reviewed on head 6aea85f. CI green ✅. 14 files, +559/-65.

Thanks for the contribution! The overall structure is solid — provider detection, CLI flag, profile env, model picker, and validation all wired up correctly. There are several issues that need addressing before merge though.


🔧 Blockers

1. No context window entries for NVIDIA NIM models

openaiContextWindows.ts has zero entries for any NVIDIA NIM model. The default model nvidia/llama-3.1-nemotron-70b-instruct falls back to the 128k generic default, which is incorrect (Nemotron 70B supports 128k context, but many NIM models have different limits). Without context window entries, users selecting NVIDIA NIM models will hit the same infinite auto-compact loop bug as Issue #635.

At minimum, add entries for the default and commonly-used models:

  • nvidia/llama-3.1-nemotron-70b-instruct → 128k context
  • nvidia/llama-3.1-nemotron-ultra-253b-v1 → 128k context
  • meta/llama-3.1-405b-instruct → 128k context
  • Other popular models users are likely to pick

2. Missing context window entries for new MiniMax models

The existing MiniMax M2.x entries are present (204,800), but the models newly added to minimaxModels.ts are missing:

  • MiniMax-Text-01 → 512k (per PR description)
  • MiniMax-Text-01-Preview → 256k
  • MiniMax-Vision-01 → 32k
  • MiniMax-Vision-01-Fast → 16k

3. Inconsistent default models across files

The default model for NVIDIA NIM differs between files:

  • configs.ts → nvidia/llama-3.1-nemotron-70b-instruct
  • providerFlag.ts → nvidia/nemotron-mini-4b-instruct
  • providerProfiles.ts preset → nvidia/nemotron-mini-4b-instruct
  • providerProfile.ts buildNvidiaNimProfileEnv → nvidia/llama-3.1-nemotron-70b-instruct

Same for MiniMax:

  • configs.ts → MiniMax-M2.5
  • providerFlag.ts → MiniMax-M2
  • providerProfiles.ts preset → MiniMax-M2
  • providerProfile.ts buildMiniMaxProfileEnv → MiniMax-M2.5

Pick one consistent default per provider. The flagship model (nemotron-70b-instruct / MiniMax-M2.5) seems like the right choice for configs and profile env (used at runtime), while the lighter model (nemotron-mini-4b / MiniMax-M2) might make sense for the initial preset.

4. Inconsistent MiniMax API base URL

.env.example uses https://api.minimax.chat/v1 but providerFlag.ts, providerProfile.ts, and providerProfiles.ts all use https://api.minimax.io/v1. These are different domains. Verify which is correct and use it consistently.


🟡 Non-blocking concerns

5. ~30 non-chat models in NVIDIA NIM picker

The model list includes embedding models (nv-embedcode-7b, nemoretriever-*, nv-embed-*), reward models (nemotron-70b-reward, nemotron-4-340b-reward), safety guard models (nemotron-safety-guard-8b, nemoguard-*, content-safety-reasoning-*), PII detection (gliner-pii), parse models, translation, and speech models. These are not suitable for a coding assistant — selecting an embedding or reward model would produce garbage or errors. Consider filtering to only instruct/chat/code models, or at least grouping them with clear warnings.

6. sanitizeProviderConfigValue signature changes

The PR removes the third processEnv argument from several sanitizeProviderConfigValue calls and adds SecretValueSource type casts in buildLaunchEnv. The function signature on main is (value, ...sources: Array<SecretValueSource | null | undefined>), so removing processEnv (a NodeJS.ProcessEnv, not a SecretValueSource) is correct — passing it was a type mismatch. But the processEnv as SecretValueSource casts in buildLaunchEnv (lines 540, 543, etc.) are also type-unsafe. This should be cleaned up properly rather than casting.

7. isMiniMaxProvider() detection relies on MINIMAX_API_KEY

isMiniMaxProvider() returns true if MINIMAX_API_KEY is set. But if someone sets MINIMAX_API_KEY in their env for other reasons while using a different provider, they'd get MiniMax behavior unexpectedly. Consider using a dedicated env var like CLAUDE_CODE_USE_MINIMAX (similar to CLAUDE_CODE_USE_GEMINI / CLAUDE_CODE_USE_OPENAI) for explicit provider selection, with MINIMAX_API_KEY only for credentials.

8. isNvidiaNimProvider() uses NVIDIA_NIM env var inconsistently

The detection checks isEnvTruthy(process.env.NVIDIA_NIM) which is set by providerFlag.ts to '1'. But .env.example documents NVIDIA_API_KEY, not NVIDIA_NIM. The NVIDIA_NIM flag is an implementation detail of the --provider nvidia-nim flow — it shouldn't be a user-facing env var. If users are expected to set NVIDIA_NIM=1 directly, document it; otherwise, consider using CLAUDE_CODE_USE_NVIDIA_NIM for consistency with other providers.

9. configs.ts adds nvidia-nim and minimax keys to every model config

Every model config (Claude 3.7 Sonnet, Haiku, Opus 4, etc.) now has nvidia-nim and minimax entries. This means selecting "Claude Sonnet 4" with the NVIDIA NIM provider maps to nvidia/llama-3.1-nemotron-70b-instruct, which is a very different model. This is conceptually similar to how other providers work (Gemini maps Claude configs to Gemini models), but the naming could confuse users who expect a Claude model and get Nemotron instead. Just flagging this for awareness.

10. .ts → .js import extensions in providerProfile.ts

The diff changes .ts extensions to .js in imports (e.g., from './providerRecommendation.ts' → from './providerRecommendation.js'). This is correct for ESM module resolution, but it's mixed with the rest of the codebase that may still use .ts extensions. Worth verifying this matches the project convention consistently.


Verdict: Needs changes 🔧

Four blockers: (1) missing NVIDIA NIM context window entries, (2) missing new MiniMax context window entries, (3) inconsistent default models across files, (4) inconsistent MiniMax API URL. These will cause real user-facing bugs (auto-compact loops, wrong models, wrong endpoints). The non-blocking items (model list filtering, detection logic, import style) can be addressed in a follow-up.

@LifeJiggy

LifeJiggy commented Apr 13, 2026 •

Copy link
Copy Markdown
Contributor Author

@Vasanthdev2004

Blockers Fixed:

  1. NVIDIA NIM context window entries - Added entries for nemotron-70b, nemotron-mini-4b, llama-3.1 models (128k), deepseek-r1/v3, qwen3

  2. New MiniMax context window entries - Added MiniMax-Text-01 (512k), Text-01-Preview (256k), Vision-01 (32k), Vision-01-Fast (16k), M2 (204k)

  3. Inconsistent default models - Now consistent across all files:

    • NVIDIA NIM: z-ai/glm5 (GLM-5)
    • MiniMax: MiniMax-M2.5
  4. MiniMax API URL - Kept api.minimax.io (correct domain)

Non-blocking (addressed in follow-up):

  • Model list filtering for non-chat models
  • Provider detection logic improvements
  • Import extension consistency

@Vasanthdev2004 Vasanthdev2004 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-review: PR #552 — Add NVIDIA NIM and MiniMax provider support (head 2b4f982)

CI green ✅. Thanks for addressing the blockers, @LifeJiggy — good progress. Several issues remain though.


🔧 Still blocking

1. Default model z-ai/glm5 has no context window entry

The default NVIDIA NIM model was changed to z-ai/glm5 across configs/providerFlag/providerProfiles, but there's no entry for z-ai/glm5 (or z-ai/glm4.7) in openaiContextWindows.ts. GLM-5 has a 200K context window — without an entry, it falls back to 128k. This means users of the default NVIDIA NIM model will get a suboptimal context window (128k instead of 200k), and won't benefit from the full model capability.

Must add:

  • z-ai/glm5 → 200,000 context (GLM-5 specs)
  • z-ai/glm4.7 → context window per Zhipu AI docs

Similarly, many models in the NVIDIA NIM picker (moonshotai/kimi-*, writer/palmyra-*, 01-ai/yi-large, databricks/dbrx-instruct, ai21labs/jamba-*) have no context window entries. Every model in the picker that a user might select needs a context entry or will hit the 128k fallback.

2. No OPENAI_MAX_OUTPUT_TOKENS entries for NVIDIA NIM or new MiniMax models

OPENAI_MAX_OUTPUT_TOKENS has entries for existing MiniMax M2.x variants (131,072) but nothing for:

  • Any NVIDIA NIM model (nemotron-70b, nemotron-mini, llama variants, etc.)
  • New MiniMax models: MiniMax-Text-01, MiniMax-Text-01-Preview, MiniMax-Vision-01, MiniMax-Vision-01-Fast, MiniMax-M2

Without these, getModelMaxOutputTokens() falls back to conservative defaults, which could cause truncated outputs or unnecessary max_tokens capping.

3. .env.example MiniMax URL still inconsistent

.env.example still shows https://api.minimax.chat/v1 while all code files use https://api.minimax.io/v1. The contributor said "kept api.minimax.io (correct domain)" but didn't update .env.example to match. These are different domains — users following .env.example will hit the wrong endpoint.


🟡 Non-blocking but important

4. Default model changed from Nemotron 70B to GLM-5 — surprising choice

The original default was nvidia/llama-3.1-nemotron-70b-instruct (NVIDIA's flagship instruct model). It was changed to z-ai/glm5 (a Zhipu AI model). While GLM-5 is a strong model, it's a third-party model hosted on NIM, not an NVIDIA native model. The NVIDIA NIM provider might be more intuitively expected to default to an NVIDIA-first model. Just flagging — maintainer call on whether this makes sense.

5. configs.ts maps all Claude configs to z-ai/glm5 and MiniMax-M2.5

Every Claude model config (Sonnet, Haiku, Opus, etc.) now maps to z-ai/glm5 for nvidia-nim and MiniMax-M2.5 for minimax. This is the same pattern used for other providers (Gemini maps to gemini-2.0-flash, etc.), so it's consistent — just making sure this is intentional.

6. ~30 non-chat models still in NVIDIA NIM picker

Embedding models (nv-embedcode-7b, nemoretriever-*), reward models (nemotron-70b-reward), safety guards (nemotron-safety-guard-8b, nemoguard-*), PII detection (gliner-pii), parse models, translation, and speech models are still in the model picker. Selecting any of these in a coding assistant will produce garbage output or errors. The contributor acknowledged this as a follow-up, but it significantly impacts first-run user experience.

7. .env.example model also inconsistent

.env.example shows OPENAI_MODEL=MiniMax-Text-01 but code defaults to MiniMax-M2.5. Minor but could confuse users setting up via .env.example.


Verdict: Still needs changes 🔧

Three remaining blockers: (1) z-ai/glm5 default model missing context window entry, (2) no MAX_OUTPUT_TOKENS entries for any NVIDIA NIM or new MiniMax models, (3) .env.example MiniMax URL still inconsistent with code. The context window gaps will cause suboptimal behavior for the very default model users will start with.

…ax output tokens, fix .env.example, revert to Nemotron default
@LifeJiggy

Copy link
Copy Markdown
Contributor Author

Good day

Additional Fixes Applied

All items from second review addressed:

  1. z-ai/glm5 context window - Added z-ai/glm5 (200k), z-ai/glm4.7, moonshotai/kimi-k2.5, databricks/dbrx-instruct, ai21labs/jamba entries to context windows

  2. MAX_OUTPUT_TOKENS entries - Added for NVIDIA NIM models (nemotron-70b, nemotron-mini, llama variants, GLM models, deepseek, kimi) and new MiniMax models (M2, Text-01, Vision-01)

  3. .env.example URL - Fixed to https://api.minimax.io/v1 and model to MiniMax-M2.5

  4. Default model - Reverted to nvidia/llama-3.1-nemotron-70b-instruct (NVIDIA's flagship native model)

  5. configs.ts mapping - Intentional, matches other providers pattern (Gemini → gemini-2.0-flash)

  6. Non-chat models in picker - Will address in follow-up PR (filter embeddings, reward models, safety guards)

  7. .env.example model - Fixed to MiniMax-M2.5

All blocking items resolved. Ready for re-review.

@Vasanthdev2004 Vasanthdev2004 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-review: PR #552 — NVIDIA NIM + MiniMax provider support (head f5245d7)

CI green ✅. 15 files, +870/-31. Thanks for the updates, @LifeJiggy — all three previous blockers are addressed:

  • ✅ z-ai/glm5 context window entry present (200,000) and MAX_OUTPUT_TOKENS (32,768)
  • ✅ MAX_OUTPUT_TOKENS entries added for NVIDIA NIM and new MiniMax models
  • ✅ .env.example uses api.minimax.io/v1

The provider onboarding is much more complete now. Still a few issues:


🔧 Blockers

1. 36 non-chat models in the NVIDIA NIM model picker

The nvidiaNimModels.ts file lists 144 models. At least 36 of them are not coding/chat models and will produce garbage or errors if selected by a user:

  • 8 embedding models: nv-embedcode-7b, nemoretriever-*, nv-embedqa-*, nemotron-embed-*, arctic-embed-l
  • 2 reward models: nemotron-70b-reward, nemotron-4-340b-reward
  • 6 safety/guard models: nemotron-safety-guard, nemoguard-*, llama-guard-4, shieldgemma, granite-guardian, gliner-pii
  • 2 parse models: nemoretriever-parse, nemotron-parse
  • 8 code completion/fill models (labeled "AGENTIC CODE SECURITY"): codestral-22b, mamba-codestral-7b, deepseek-coder-6.7b, codegemma-*, codellama-70b, starcoder2-*, granite-*-code
  • 10+ others: vision-only (vila, neva-22b, paligemma, nvclip), speech (streampetr), translation (riva-translate), charts (deplot), kosmos-2

Additionally, all the embedding/reward/parse models are mislabeled with description: 'AGENTIC SECURITY' — they're not security models, and this label makes the categorization meaningless.

Fix: Remove all non-chat models from the picker, or add a category field and filter to show only chat/instruct models in the /model command. At minimum, remove embedding, reward, safety-guard, PII-detection, parse, vision-only, and speech models — these will not work as coding assistants.

2. applyProviderProfileToProcessEnv doesn't set MINIMAX_API_KEY or NVIDIA_API_KEY

When a user saves a MiniMax or NVIDIA NIM profile with an API key, applyProviderProfileToProcessEnv() sets OPENAI_API_KEY from profile.apiKey but never sets MINIMAX_API_KEY or NVIDIA_API_KEY.

This means openaiShim.ts's isMiniMax = !!process.env.MINIMAX_API_KEY is false when using a saved profile, even though the API key IS set (just under OPENAI_API_KEY). Currently this only affects the isMiniMax flag used for API key fallback — the key still works via OPENAI_API_KEY. But any future MiniMax-specific logic gated on isMiniMax will silently fail for saved profiles.

This is the same class of bug as PR #623's blocker for Mistral/Gemini API keys.

Fix in applyProviderProfileToProcessEnv:

// After setting OPENAI_API_KEY:
if (profile.name === 'MiniMax' || process.env.MINIMAX_API_KEY) {
  process.env.MINIMAX_API_KEY ??= profile.apiKey
}
if (profile.name === 'NVIDIA NIM' || process.env.NVIDIA_API_KEY) {
  process.env.NVIDIA_API_KEY ??= profile.apiKey
}

Or more robustly: detect by base URL pattern.


🟡 Non-blocking but important

3. sanitizeProfile still forces provider to 'anthropic' | 'openai'

nvidia-nim and minimax provider profiles get their provider field overwritten to 'openai' by sanitizeProfile. This means saved NVIDIA NIM profiles lose their provider identity. The runtime detection (URL pattern + env var) still works, but the saved profile is misleading. Not a blocker since the runtime path is correct, but worth noting for consistency with the ProviderProfile type now including 'nvidia-nim' | 'minimax'.

4. Default model is Nemotron 70B — reasonable but not the most capable coding model

The default for nvidia-nim is nvidia/llama-3.1-nemotron-70b-instruct. Among the 144 models, several are stronger for coding (e.g., deepseek-ai/deepseek-v3.2, qwen/qwen3-coder-480b-a35b-instruct). Nemotron 70B is a solid general model but not NVIDIA's best for code. This is subjective — leaving as non-blocking.

5. 111 of 144 NVIDIA NIM models lack context window entries

Only 33 NVIDIA NIM models have entries in openaiContextWindows.ts. The remaining 111 will fall back to 128K (after PR #636 merges) or 8K (current main). This isn't blocking since the fallback exists, but many NVIDIA models have much smaller context windows (4K-8K), and a 128K fallback for a 4K model could cause API errors.


✅ What looks good

  • .env.example correctly uses api.minimax.io/v1 ✅
  • MiniMax model list is clean and appropriate (11 models, all chat)
  • MiniMax model descriptions include context window sizes
  • providerFlag.ts correctly sets env vars for both NVIDIA NIM and MiniMax
  • providerDiscovery.ts detects both providers by URL pattern and env var
  • validateModel.ts correctly validates against cached model lists
  • Default models in configs.ts are consistent (Nemotron 70B for NIM, M2.5 for MiniMax)
  • Context window and MAX_OUTPUT_TOKENS entries for key models are present
  • isNvidiaNimProvider() and isMiniMaxProvider() detection in both runtime and model options
  • CI green ✅

Verdict: Needs changes 🔧

Two blockers: (1) 36 non-chat models in the NVIDIA NIM picker will produce garbage or errors — remove embedding, reward, safety-guard, PII, parse, vision-only, and speech models at minimum; (2) applyProviderProfileToProcessEnv doesn't set MINIMAX_API_KEY/NVIDIA_API_KEY from saved profiles, causing isMiniMax detection to fail. The rest is solid.

@LifeJiggy

Copy link
Copy Markdown
Contributor Author

All Blockers Resolved

Blockers Fixed:

  1. 36 non-chat models in NVIDIA NIM picker - Filtered to ~100 chat/instruct models only. Removed:

    • Embedding models (nv-embedcode, nemoretriever, nv-embed, arctic-embed)
    • Reward models (nemotron-70b-reward, nemotron-4-340b-reward)
    • Safety/guard models (nemotron-safety-guard, nemoguard, llama-guard, shieldgemma, granite-guardian, gliner-pii)
    • Parse models (nemoretriever-parse, nemotron-parse)
    • Vision-only models (neva-22b, vila, nvclip, paligemma, deplot)
    • Speech/translation (streampetr, riva-translate)
    • Also fixed mislabeled descriptions (was "AGENTIC SECURITY" for embeddings)
  2. applyProviderProfileToProcessEnv doesn't set MINIMAX_API_KEY/NVIDIA_API_KEY - Now detects provider by base URL pattern and sets the appropriate provider-specific API key when applying saved profiles.

Non-blocking Items:

  • Provider profile sanitize forces 'openai' - noted, runtime detection still works
  • Default model is Nemotron 70B - reasonable NVIDIA flagship
  • 111 of 144 models lack context entries - fallback exists, can expand in follow-up

All blocking issues resolved. Ready for merge ✅

Vasanthdev2004
Vasanthdev2004 previously approved these changes Apr 14, 2026

@Vasanthdev2004 Vasanthdev2004 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-review: PR #552 — NVIDIA NIM + MiniMax provider support (head a0cec05)

CI green ✅. 15 files, +870/-31. Both previous blockers are now fixed:

  • ✅ Non-chat models removed: 144 → 111 models. All embedding, reward, safety-guard, PII-detection, parse, vision-only, and speech models are gone. The remaining models are all chat/instruct/code/reasoning models.
  • ✅ MINIMAX_API_KEY / NVIDIA_API_KEY set from saved profiles: applyProviderProfileToProcessEnv now detects MiniMax and NVIDIA by base URL pattern and sets the provider-specific API key env var. The isMiniMax detection in openaiShim.ts will now work correctly for saved profiles.

Thanks for iterating on this, @LifeJiggy — the provider onboarding is in much better shape now.


🟡 Non-blocking

1. sanitizeProfile still forces provider to 'anthropic' | 'openai'

Saved NVIDIA NIM and MiniMax profiles get their provider field overwritten to 'openai'. Runtime detection (URL pattern + env var) still works correctly, so this doesn't break anything — the saved profile is just misleading. Consider updating sanitizeProfile to preserve 'nvidia-nim' | 'minimax' provider values for consistency with the ProviderProfile type.

2. 94 of 111 NVIDIA NIM models lack context window entries

Only 17 NVIDIA NIM models have entries in openaiContextWindows.ts. The remaining 94 will fall back to 128K (after PR #636 merges). Many of these models have 4K–8K context windows, so the 128K fallback is too generous and could cause the tool to send more context than the model supports. The provider enforces server-side limits, so this won't crash — it just means wasted API calls that get rejected, or premature auto-compact that could be avoided with correct entries. Worth adding entries for the popular models (Gemma, Llama, Phi, Granite, DeepSeek) in a follow-up.

3. mixtral-8x22b-v0.1 is a base model, not instruct

The model list includes mistralai/mixtral-8x22b-v0.1 (labeled "Mixtral 8x22B Base"). This is a base/completion model, not an instruct-tuned chat model. Selecting it will work but produce less helpful, unformatted output. The -instruct-v0.1 variant is also available and would be a better choice. Non-blocking since the label says "Base."

4. clearProviderProfileEnvFromProcessEnv doesn't clear MINIMAX_API_KEY / NVIDIA_API_KEY

When switching away from MiniMax or NVIDIA NIM, these env vars linger in process.env. Since the USE_* flags are cleared and runtime detection is URL-based, this doesn't cause routing bugs — the vars are just dead. Same class of issue as PR #623. Worth fixing in a follow-up along with the other provider-specific env var cleanup.


✅ All blockers resolved

  • ✅ Non-chat models removed from NVIDIA NIM picker
  • ✅ MINIMAX_API_KEY / NVIDIA_API_KEY set from saved profiles
  • ✅ .env.example uses api.minimax.io/v1
  • ✅ z-ai/glm5 context window + MAX_OUTPUT_TOKENS entries present
  • ✅ MiniMax model list is clean (11 models, all chat)
  • ✅ providerFlag.ts correctly sets env vars for both providers
  • ✅ providerDiscovery.ts detects both providers by URL pattern
  • ✅ Default models consistent (Nemotron 70B for NIM, M2.5 for MiniMax)
  • ✅ validateModel.ts validates against cached model lists
  • ✅ CI green

Verdict: Approve-ready ✅

Both blockers are properly addressed. The non-blocking items (sanitizeProfile, missing context entries, base model in list, env var cleanup) are follow-up quality improvements. Ship-ready from my perspective.

@LifeJiggy

Copy link
Copy Markdown
Contributor Author

Final Updates - All Review Items Addressed

Non-blocking Items Fixed:

  1. mixtral-8x22b base model - Changed to mixtral-8x22b-instruct-v0.1 (instruct variant)

  2. Context window entries - Added ~40 more entries for popular NVIDIA NIM models:

    • Meta Llama 3.3, Google Gemma 2/3, Mistral, Microsoft Phi, IBM Granite, DeepSeek v3.2, Qwen 2.5, Yi Large
    • Plus corresponding MAX_OUTPUT_TOKENS entries
  3. clearProviderProfileEnvFromProcessEnv - Now clears MINIMAX_API_KEY, NVIDIA_API_KEY, NVIDIA_NIM when switching away from these providers

  4. sanitizeProfile provider field - Noted for follow-up (requires config.ts type change to preserve 'nvidia-nim' | 'minimax')

Summary:

  • ✅ Blockers: All 3 resolved
  • ✅ Non-blocking: All 4 addressed (3 fixed, 1 noted for follow-up)
  • ✅ Model picker filtered to ~110 chat/instruct models
  • ✅ Context window entries for 50+ NVIDIA NIM models
  • ✅ Saved profiles correctly set provider API keys
  • ✅ .env.example consistent with code

Ready now

@kevincodex1

Copy link
Copy Markdown
Member

please have a look again bro @Vasanthdev2004

@Vasanthdev2004 Vasanthdev2004 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-review: PR #552 — head 52215afa (changes since approved a0cec05)

CI green ✅. 3 files changed since previous approval, +52/-12.

Changes since a0cec05

1. nvidiaNimModels.ts — Fixed Mixtral base→instruct ✅
mixtral-8x22b-v0.1 (base) → mixtral-8x22b-instruct-v0.1 (instruct). This was my non-blocking suggestion — base models work but produce less helpful output.

2. openaiContextWindows.ts — Added context window entries for popular NVIDIA NIM models ✅

  • Google Gemma 2/3 (8k for Gemma 2, 131k for Gemma 3)
  • DeepSeek v3.2 (128k)
  • Qwen 2.5-7b (32k)
  • Mistral Large 3 (256k), Small 3.1 (32k), Mixtral 8x7b (32k)
  • Microsoft Phi-4-mini (16k), Phi-3.5-mini (16k), Phi-3-mini-128k (128k)
  • IBM Granite 3.3 (8k), Granite 8b-code (8k)
  • Yi Large (32k)
  • Also fixed nvidia/llama-3.2-* → meta/llama-3.2-* prefix, added Llama 3.3-70b

This addresses my concern about 94 of 111 models lacking entries. Popular models now have correct values.

3. openaiContextWindows.ts — Added max_output_tokens entries ✅
New entries for Gemma (4k/16k), DeepSeek v3.2 (32k), Qwen, Mistral, Phi, Granite, Yi, DBRX, Jamba. Removed Ollama-specific entries (gemma2:27b, codellama:13b, etc.) which belong in Ollama-specific tables.

4. providerProfiles.ts — Added MINIMAX_API_KEY, NVIDIA_API_KEY, NVIDIA_NIM clearing ✅
This addresses my concern about stale env vars lingering on provider switch. Same class of bug as #623.

Assessment

All changes address feedback from my previous review. No new concerns. Approval stands.

Verdict: Approve-ready ✅ (re-confirmed)

@kevincodex1
kevincodex1 merged commit 51191d6 into Twigpine:main Apr 15, 2026
1 check passed
C1ph3r404 pushed a commit to C1ph3r404/openclaude that referenced this pull request Apr 29, 2026
* feat: add NVIDIA NIM and MiniMax provider support

- Add nvidia-nim and minimax to --provider CLI flag
- Add model discovery for NVIDIA NIM (160+ models) and MiniMax
- Update /model picker to show provider-specific models
- Fix provider detection in startup banner
- Update .env.example with new provider options

Supported providers:
- NVIDIA NIM: https://integrate.api.nvidia.com/v1
- MiniMax: https://api.minimax.io/v1

* fix: resolve conflict in StartupScreen (keep NVIDIA/MiniMax + add Codex detection)

* fix: resolve providerProfile conflict (add imports from main, keep NVIDIA/MiniMax)

* fix: revert providerSecrets to match main (NVIDIA/MiniMax handled elsewhere)

* fix: add context window entries for NVIDIA NIM and new MiniMax models

* fix: use GLM-5 as NVIDIA NIM default and MiniMax-M2.5 for consistency

* fix: address remaining review items - add GLM/Kimi context entries, max output tokens, fix .env.example, revert to Nemotron default

* fix: filter NVIDIA NIM picker to chat/instruct models only, set provider-specific API keys from saved profiles

* chore: add more NVIDIA NIM context window entries for popular models

* fix: address remaining non-blocking items - fix base model, clear provider API keys on profile switch
The-FOOL-00 pushed a commit to The-FOOL-00/openclaude that referenced this pull request May 24, 2026
* feat: add NVIDIA NIM and MiniMax provider support

- Add nvidia-nim and minimax to --provider CLI flag
- Add model discovery for NVIDIA NIM (160+ models) and MiniMax
- Update /model picker to show provider-specific models
- Fix provider detection in startup banner
- Update .env.example with new provider options

Supported providers:
- NVIDIA NIM: https://integrate.api.nvidia.com/v1
- MiniMax: https://api.minimax.io/v1

* fix: resolve conflict in StartupScreen (keep NVIDIA/MiniMax + add Codex detection)

* fix: resolve providerProfile conflict (add imports from main, keep NVIDIA/MiniMax)

* fix: revert providerSecrets to match main (NVIDIA/MiniMax handled elsewhere)

* fix: add context window entries for NVIDIA NIM and new MiniMax models

* fix: use GLM-5 as NVIDIA NIM default and MiniMax-M2.5 for consistency

* fix: address remaining review items - add GLM/Kimi context entries, max output tokens, fix .env.example, revert to Nemotron default

* fix: filter NVIDIA NIM picker to chat/instruct models only, set provider-specific API keys from saved profiles

* chore: add more NVIDIA NIM context window entries for popular models

* fix: address remaining non-blocking items - fix base model, clear provider API keys on profile switch
discopops pushed a commit to discopops/openclaude that referenced this pull request May 28, 2026
* feat: add NVIDIA NIM and MiniMax provider support

- Add nvidia-nim and minimax to --provider CLI flag
- Add model discovery for NVIDIA NIM (160+ models) and MiniMax
- Update /model picker to show provider-specific models
- Fix provider detection in startup banner
- Update .env.example with new provider options

Supported providers:
- NVIDIA NIM: https://integrate.api.nvidia.com/v1
- MiniMax: https://api.minimax.io/v1

* fix: resolve conflict in StartupScreen (keep NVIDIA/MiniMax + add Codex detection)

* fix: resolve providerProfile conflict (add imports from main, keep NVIDIA/MiniMax)

* fix: revert providerSecrets to match main (NVIDIA/MiniMax handled elsewhere)

* fix: add context window entries for NVIDIA NIM and new MiniMax models

* fix: use GLM-5 as NVIDIA NIM default and MiniMax-M2.5 for consistency

* fix: address remaining review items - add GLM/Kimi context entries, max output tokens, fix .env.example, revert to Nemotron default

* fix: filter NVIDIA NIM picker to chat/instruct models only, set provider-specific API keys from saved profiles

* chore: add more NVIDIA NIM context window entries for popular models

* fix: address remaining non-blocking items - fix base model, clear provider API keys on profile switch
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants