fix(providers): replace retired/stale model runtime defaults - #1823
Conversation
Several silent runtime defaults pointed at models that are either hard-retired (Google gemini-1.5-pro/flash: "SHUT DOWN. Returns 404." per src/lib/constants/enums.ts) or stale (OpenAI gpt-4o/gpt-4o-mini, Anthropic claude-3-haiku), so a caller who never overrides the relevant env var — or who follows the SDK's own error-recovery suggestion — was handed a model that 404s or is materially out of date. Fixed, all verified against marketing/data/factory/model-currency-map-2026-09.normalized.json and src/lib/constants/enums.ts: - RAGAS/guardrails/multi-judge evaluation judge default: gemini-1.5-flash -> gemini-2.5-flash (EvaluatorFactory.ts, ragasEvaluator.ts, EvaluatorRegistry.ts, guardrailsUtils.ts, multiJudgeWorkflow.ts) - LiteLLM provider default and fallback list: openai/gpt-4o-mini -> openai/gpt-5.4-mini, anthropic/claude-3-haiku -> anthropic/claude-haiku-4-5-20251001 (litellm/client.ts, and the mirrored getConfiguredLiteLLMModel()/getCommonModelsForProvider() guesses in providerHealth.ts) - OpenAI provider default: gpt-4o -> gpt-5.4 (openAI/client.ts) - Retired gemini-1.5-pro/flash dropped from CLI model-choice lists, health-check recommendations, and error-recovery suggestion lists (modelChoices.ts, providerHealth.ts, openRouter/client.ts, googleVertex/client.ts) Deliberately left alone (separate, reported findings, not in scope for this minimal fix): the wider Anthropic Claude-3.x family, which spans 30+ files well beyond this bug; GoogleAIModels.GEMINI_2_0_FLASH (announced, not confirmed-retired); GEMINI_3_PRO_PREVIEW; and the pricing/context-window/capability tables that legitimately key on retired ids. Adds test/continuous-test-suite-retired-model-defaults.ts (13 tests, offline, wired into test:unit via test:retired-model-defaults) as a static source-text regression guard: fails on origin/release (every retired/stale literal still present) and passes after this fix.
|
Warning Review limit reachedNext included review available in 15 minutes. View limit detailsLimit details: You’ve used all 2 included reviews currently available. You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. Review configuration: ⚙️ Run configurationConfiguration used: Repository: juspay/neurolink/.coderabbit.yaml Review profile: CHILL Plan: Advanced Run ID: 📒 Files selected for processing (13)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
✅ Single Commit Policy - COMPLIANTStatus: Policy requirements met • 1 commit • Valid format • Ready for merge 📊 View validation details📝 Commit Details
✅ Validation Results
🤖 Automated validation by NeuroLink Single Commit Enforcement |
Tara-ag
left a comment
There was a problem hiding this comment.
NEEDS_WORK — the retirement is mostly correct but leaves a stale openai/gpt-4o in the LiteLLM fallback path and lets the OpenAI runtime default drift from the public default surface.
| "openai/gpt-4o", | ||
| "anthropic/claude-3-haiku", | ||
| "anthropic/claude-haiku-4-5-20251001", |
There was a problem hiding this comment.
MAJOR — getFallbackModels() still advertises stale openai/gpt-4o as its first fallback.
This PR's whole purpose is retiring the stale OpenAI defaults: FALLBACK_LITELLM_MODEL moved to openai/gpt-5.4-mini, and the health-check "try this instead" probe moved off gpt-4o-mini. But the first entry of this fallback list — the model handed to any caller whose /v1/models fetch fails and falls back to the static suggestion list — is still "openai/gpt-4o" (line 335, left unchanged two lines above the edited line). A caller who exhausts the model-not-found path is handed the exact stale/retired default this PR claims to remove, contradicting the other edits in the same function.
The new test also misses this (see the test-file finding). Suggest aligning the head of the list with FALLBACK_LITELLM_MODEL:
| "openai/gpt-4o", | |
| "anthropic/claude-3-haiku", | |
| "anthropic/claude-haiku-4-5-20251001", | |
| "openai/gpt-5.4-mini", | |
| "anthropic/claude-haiku-4-5-20251001", |
| const getOpenAIApiKey = (): string => validateApiKey(createOpenAIConfig()); | ||
|
|
||
| const getOpenAIModel = (): string => getProviderModel("OPENAI_MODEL", "gpt-4o"); | ||
| const getOpenAIModel = (): string => |
There was a problem hiding this comment.
MAJOR — OpenAI runtime default diverges from the public default surface.
You moved the provider runtime default here to "gpt-5.4", but the sibling OpenAI default surfaces were left pointing at gpt-4o:
| Surface | Current value |
|---|---|
getOpenAIModel() (this line, fixed) |
gpt-5.4 |
DEFAULT_MODELS[OPENAI] → getDefaultModel("openai") (src/lib/utils/modelChoices.ts) |
GPT_4O |
TOP_MODELS_CONFIG[OPENAI] headline, labelled "Recommended - Latest multimodal model" |
GPT_4O |
providerHealth getCommonModelsForProvider OPENAI case ("try this instead") |
GPT_4O, GPT_4O_MINI first |
So getDefaultModel(OPENAI) and the health-check recommendation now contradict the actual model a bare new OpenAIProvider() runs (rule 5 surface: unmodified callers reading the default get a different answer than the runtime produces). This is exactly the divergence the PR set out to close, but applied only on one surface. Align modelChoices.ts and providerHealth.ts to GPT_5_4/GPT_5_4_MINI (or whichever current id the currency map uses) for consistency.
| "litellm/client.ts FALLBACK_LITELLM_MODEL must be openai/gpt-5.4-mini (was stale openai/gpt-4o-mini)", | ||
| ); | ||
| assert( | ||
| !content.includes('"openai/gpt-4o-mini"'), |
There was a problem hiding this comment.
MINOR — LiteLLM assertions cover gpt-4o-mini but not openai/gpt-4o, so the primary fallback regression above slips through.
This suite's LiteLLM test asserts !includes('"openai/gpt-4o-mini"') (this line) and the claude-haiku change, but never checks that getFallbackModels() no longer leads with "openai/gpt-4o". That is precisely why the litellm-fallback-gpt4o finding passes CI: the file can still contain the stale openai/gpt-4o fallback and this test stays green. Strengthen the guard:
| !content.includes('"openai/gpt-4o-mini"'), | |
| !content.includes('"openai/gpt-4o-mini"') && | |
| !content.includes('"openai/gpt-4o"'), | |
| "litellm/client.ts must not contain the stale openai/gpt-4o-mini or openai/gpt-4o literal anywhere", |
NEEDS_WORK — retirement done right, but one stale fallback and one surface drift remainThe core refactor is sound:
Checked and clean
Note for the authorReply to the two MAJORs confirming whether the additional model-id surfaces ( |
Tara-ag
left a comment
There was a problem hiding this comment.
NEEDS_WORK — two MAJOR findings remain (see inline comments): the LiteLLM fallback still leads with the retired openai/gpt-4o, and the OpenAI runtime default diverges from the modelChoices.ts/providerHealth.ts surfaces. Please address before merge.
|
🎉 This PR is included in version 12.34.0 🎉 The release is available on: Your semantic-release bot 📦🚀 |
…PAA claims Re-audit of the 173-post corpus (prompted by a hunch the last audit missed some) found 10 posts still citing stale or retired model IDs and 4 flagged for SOC2/HIPAA wording. On close read, only 5 files needed real changes: - model-evaluation-scoring.md cited gemini-1.5-flash as NeuroLink's judge fallback -- that model is hard-retired (404s). Now gemini-2.5-flash, matching the SDK fix in juspay/neurolink#1823. - openai-integration-guide.md and litellm-unified-routing.md cited the old gpt-4o/gpt-4o-mini/claude-3-haiku fallbacks -- also matching the same SDK fix; now gpt-5.4/gpt-5.4-mini/claude-haiku-4-5. - provider-abstraction.md had one stale "legacy gpt-4o" clause. - ai-ethics-responsible-use.md's two SOC2/HIPAA mentions implied NeuroLink could grant compliance; reworded to state plainly that NeuroLink holds no such certification and describe the real building blocks it provides. The other 9 flagged files were read in full and left untouched: their model mentions are legitimately historical/legacy-labeled, and 3 of the 4 SOC2/HIPAA flags were already honestly worded (attributing certification to third-party providers or explicitly disclaiming NeuroLink's own).
Each change below answers a review thread on an already-merged PR where the reviewer's claim held on the current tree. Where behaviour changed, the suite that guards it was made to fail without the change (reversed or mutated, then restored) and pass with it. Source - litellm/client.ts: getFallbackModels() still offered the stale openai/gpt-4o; it now offers openai/gpt-5.4. retired-model-defaults asserts both the absence of the old literal and the presence of the new one. (F-test-misses-litellm-gpt4o, #1823) Suites - providers-mocked: the OpenAI-compat banner and header named seven providers and no Mistral although the loop runs every catalog entry plus Cohere (T3806479869, T3807355730, #1353: one defect raised twice). The detached-pump subprocess case now also requires at least one fetch call, exiting 5 when the right RateLimitError arrives without a request (PF-T3847750289, #1531). The pre-aborted Perplexity case now expects no request, see mockFetch below. - realtime-unit: clearHandlers test reads the registry from inside disconnect(), so it fails if the registry is cleared first (T3813998725, #1354). - stt-unit: captures the registry before the first test and restores it at the end, so the e2e-stt-suite-* handlers no longer outlive the file (T3813998716-b, #1354). - vertex-loop-characterization: the stand-in now records headers and answers 401 without the Express Mode key, the Express case asserts the key header, the tool round trip asserts the payload equals { result: { found: true } }, and a new case covers a blank or whitespace baseURL falling through to GOOGLE_VERTEX_BASE_URL (T3827707056-a, T3827707069-a, T3827842925-a, #1408). - video-no-ffprobe: the PATH is now only the ffmpeg link directory and node's directory; /usr/bin and /bin, where a distro ffprobe lives, no longer make both cases skip before asserting. The ffprobe preflight checks those two directories directly and the Skip stays (T4135201681, #1861). - acceptance-gate: credentialFreeEnv now also strips names that carry a secret without saying key or token (service-account and private keys, speech keys, OTLP headers, REDIS_URL, auth config, SSH_AUTH_SOCK) and ambient *_BASE_URL, *_ENDPOINT, OTEL_EXPORTER_OTLP_* and AWS_PROFILE. Because the gate's own check reused the strip pattern, ambientSecretCanaries plants a literal list of dummy values before the strip and the gate fails if any survives (T4126861003-env-isolation, #1849). - harness: withCaseTimeout now scales by NEUROLINK_TEST_TIMEOUT_SCALE like defineSuite does, so the JSDoc claim that the two cannot drift is true (T3837277828-timeout-scale-drift, #1487). Both go through scaleTimeoutMs, which floors a budget at 1ms so a tiny valid scale cannot round it to 0 (T4053369060-1, #1732). harness-offline-timeout covers both. - mockFetch: an already-aborted signal is rejected before the call is recorded, as real fetch does (F-mockfetch-aborted-records-call, #1357). Docs and tooling - model-not-found-retryable makes a real generate() on Anthropic and OpenAI and skips without both keys; it moves from test:unit to test:live in package.json, test/README.md and the CI comments (T3790920852, #1334). - sse-bisection-findings.md no longer claims created, in_progress and output_item.done are required; only output_item.added before the deltas was isolated (F4-T4087478004-minimum-event-overclaim, #1783). - acceptance-gate.md describes the wider strip and the canary check; docs-site search index regenerated. Not changed - T3813998753 (#1354): video-generation-unit sets OPENAI_API_KEY and never restores it. Deliberate (the file's own comment says why), runSuite() calls process.exit, CI runs each suite in its own process and nothing imports the file, so nothing can observe the leaked value.
Summary
gemini-1.5-flash→gemini-2.5-flash— the old model is hard-retired ("SHUT DOWN. Returns 404." persrc/lib/constants/enums.ts), so any caller relying on the default judge was silently getting a 404.openai/gpt-4o-mini→openai/gpt-5.4-mini, andanthropic/claude-3-haiku→anthropic/claude-haiku-4-5-20251001(the old Anthropic id is retired).gpt-4o→gpt-5.4.gemini-1.5-pro/gemini-1.5-flashfrom CLI model-choice lists, health-check recommendations, and error-recovery suggestion lists (they'd just 404 again).GoogleAIModels.GEMINI_2_0_FLASH(announced but not confirmed-retired),GEMINI_3_PRO_PREVIEW, and pricing/context-window/capability lookup tables that legitimately key on retired ids.All replacements verified against
src/lib/constants/enums.tsand an internal model-currency map cross-checked against provider docs.Test plan
test/continuous-test-suite-retired-model-defaults.ts(13 tests, offline static source-text regression guard), wired intotest:retired-model-defaultsand appended totest:unit. Fails on the old code (retired/stale literals present), passes after this fix.pnpm run typecheck: pre-existing ~296-errortest/baseline unaffected; zero new errors from this change.pnpm run lint: 0 errors, pre-existing warning count unchanged.pnpm run build: succeeds.pnpm run test:unit: 302/302 pass.