Skip to content

fix(providers): replace retired/stale model runtime defaults - #1823

Merged
murdore merged 1 commit into
releasefrom
fix/retired-model-defaults
Sep 27, 2026
Merged

murdore merged 1 commit into
releasefrom
fix/retired-model-defaults

Conversation

@murdore

@murdore murdore commented Sep 27, 2026

Copy link
Copy Markdown
Contributor

Summary

  • RAGAS/guardrails/multi-judge judge default gemini-1.5-flash → gemini-2.5-flash — the old model is hard-retired ("SHUT DOWN. Returns 404." per src/lib/constants/enums.ts), so any caller relying on the default judge was silently getting a 404.
  • LiteLLM default/fallback openai/gpt-4o-mini → openai/gpt-5.4-mini, and anthropic/claude-3-haiku → anthropic/claude-haiku-4-5-20251001 (the old Anthropic id is retired).
  • OpenAI provider default gpt-4o → gpt-5.4.
  • Dropped retired gemini-1.5-pro/gemini-1.5-flash from CLI model-choice lists, health-check recommendations, and error-recovery suggestion lists (they'd just 404 again).
  • Deliberately left alone (separate, reported, not in scope for this minimal fix): the wider Anthropic Claude-3.x family (30+ files), GoogleAIModels.GEMINI_2_0_FLASH (announced but not confirmed-retired), GEMINI_3_PRO_PREVIEW, and pricing/context-window/capability lookup tables that legitimately key on retired ids.

All replacements verified against src/lib/constants/enums.ts and an internal model-currency map cross-checked against provider docs.

Test plan

  • Added test/continuous-test-suite-retired-model-defaults.ts (13 tests, offline static source-text regression guard), wired into test:retired-model-defaults and appended to test:unit. Fails on the old code (retired/stale literals present), passes after this fix.
  • pnpm run typecheck: pre-existing ~296-error test/ baseline unaffected; zero new errors from this change.
  • pnpm run lint: 0 errors, pre-existing warning count unchanged.
  • pnpm run build: succeeds.
  • pnpm run test:unit: 302/302 pass.
  • Pre-commit hook (format, codegen:catalog --check, security scan, commit-msg policy) passed clean.

Several silent runtime defaults pointed at models that are either
hard-retired (Google gemini-1.5-pro/flash: "SHUT DOWN. Returns 404."
per src/lib/constants/enums.ts) or stale (OpenAI gpt-4o/gpt-4o-mini,
Anthropic claude-3-haiku), so a caller who never overrides the
relevant env var — or who follows the SDK's own error-recovery
suggestion — was handed a model that 404s or is materially out of
date.

Fixed, all verified against
marketing/data/factory/model-currency-map-2026-09.normalized.json
and src/lib/constants/enums.ts:

- RAGAS/guardrails/multi-judge evaluation judge default:
  gemini-1.5-flash -> gemini-2.5-flash (EvaluatorFactory.ts,
  ragasEvaluator.ts, EvaluatorRegistry.ts, guardrailsUtils.ts,
  multiJudgeWorkflow.ts)
- LiteLLM provider default and fallback list: openai/gpt-4o-mini ->
  openai/gpt-5.4-mini, anthropic/claude-3-haiku ->
  anthropic/claude-haiku-4-5-20251001 (litellm/client.ts, and the
  mirrored getConfiguredLiteLLMModel()/getCommonModelsForProvider()
  guesses in providerHealth.ts)
- OpenAI provider default: gpt-4o -> gpt-5.4 (openAI/client.ts)
- Retired gemini-1.5-pro/flash dropped from CLI model-choice lists,
  health-check recommendations, and error-recovery suggestion lists
  (modelChoices.ts, providerHealth.ts, openRouter/client.ts,
  googleVertex/client.ts)

Deliberately left alone (separate, reported findings, not in scope
for this minimal fix): the wider Anthropic Claude-3.x family, which
spans 30+ files well beyond this bug; GoogleAIModels.GEMINI_2_0_FLASH
(announced, not confirmed-retired); GEMINI_3_PRO_PREVIEW; and the
pricing/context-window/capability tables that legitimately key on
retired ids.

Adds test/continuous-test-suite-retired-model-defaults.ts (13 tests,
offline, wired into test:unit via test:retired-model-defaults) as a
static source-text regression guard: fails on origin/release (every
retired/stale literal still present) and passes after this fix.
@coderabbitai

coderabbitai Bot commented Sep 27, 2026

Copy link
Copy Markdown

Warning

Review limit reached

Next included review available in 15 minutes.

Check out review usage here.

View limit details

Limit details: You’ve used all 2 included reviews currently available.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Repository: juspay/neurolink/.coderabbit.yaml

Review profile: CHILL

Plan: Advanced

Run ID: 3d307bd4-62b1-47bf-98b7-484722566e77

📥 Commits

Reviewing files that changed from the base of the PR and between 0ced8f5 and 676ae67.

📒 Files selected for processing (13)
  • package.json
  • src/lib/evaluation/EvaluatorFactory.ts
  • src/lib/evaluation/EvaluatorRegistry.ts
  • src/lib/evaluation/ragasEvaluator.ts
  • src/lib/middleware/utils/guardrailsUtils.ts
  • src/lib/providers/googleVertex/client.ts
  • src/lib/providers/litellm/client.ts
  • src/lib/providers/openAI/client.ts
  • src/lib/providers/openRouter/client.ts
  • src/lib/utils/modelChoices.ts
  • src/lib/utils/providerHealth.ts
  • src/lib/workflow/workflows/multiJudgeWorkflow.ts
  • test/continuous-test-suite-retired-model-defaults.ts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

Copy link
Copy Markdown
Contributor

✅ Single Commit Policy - COMPLIANT

Status: Policy requirements met • 1 commit • Valid format • Ready for merge

📊 View validation details

📝 Commit Details

  • Hash: 676ae67f6e1cd53acda4b623392ffefba7359e15
  • Message: fix(providers): replace retired/stale model runtime defaults
  • Author: Sachin Sharma

✅ Validation Results

  • Single commit requirement met
  • No merge commits in branch
  • Semantic commit message format verified
  • Ready for squash merge to release branch

🤖 Automated validation by NeuroLink Single Commit Enforcement

@Tara-ag Tara-ag left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

NEEDS_WORK — the retirement is mostly correct but leaves a stale openai/gpt-4o in the LiteLLM fallback path and lets the OpenAI runtime default drift from the public default surface.

Comment on lines 335 to +336
"openai/gpt-4o",
"anthropic/claude-3-haiku",
"anthropic/claude-haiku-4-5-20251001",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

MAJOR — getFallbackModels() still advertises stale openai/gpt-4o as its first fallback.

This PR's whole purpose is retiring the stale OpenAI defaults: FALLBACK_LITELLM_MODEL moved to openai/gpt-5.4-mini, and the health-check "try this instead" probe moved off gpt-4o-mini. But the first entry of this fallback list — the model handed to any caller whose /v1/models fetch fails and falls back to the static suggestion list — is still "openai/gpt-4o" (line 335, left unchanged two lines above the edited line). A caller who exhausts the model-not-found path is handed the exact stale/retired default this PR claims to remove, contradicting the other edits in the same function.

The new test also misses this (see the test-file finding). Suggest aligning the head of the list with FALLBACK_LITELLM_MODEL:

Suggested change
"openai/gpt-4o",
"anthropic/claude-3-haiku",
"anthropic/claude-haiku-4-5-20251001",
"openai/gpt-5.4-mini",
"anthropic/claude-haiku-4-5-20251001",

const getOpenAIApiKey = (): string => validateApiKey(createOpenAIConfig());

const getOpenAIModel = (): string => getProviderModel("OPENAI_MODEL", "gpt-4o");
const getOpenAIModel = (): string =>

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

MAJOR — OpenAI runtime default diverges from the public default surface.

You moved the provider runtime default here to "gpt-5.4", but the sibling OpenAI default surfaces were left pointing at gpt-4o:

Surface Current value
getOpenAIModel() (this line, fixed) gpt-5.4
DEFAULT_MODELS[OPENAI] → getDefaultModel("openai") (src/lib/utils/modelChoices.ts) GPT_4O
TOP_MODELS_CONFIG[OPENAI] headline, labelled "Recommended - Latest multimodal model" GPT_4O
providerHealth getCommonModelsForProvider OPENAI case ("try this instead") GPT_4O, GPT_4O_MINI first

So getDefaultModel(OPENAI) and the health-check recommendation now contradict the actual model a bare new OpenAIProvider() runs (rule 5 surface: unmodified callers reading the default get a different answer than the runtime produces). This is exactly the divergence the PR set out to close, but applied only on one surface. Align modelChoices.ts and providerHealth.ts to GPT_5_4/GPT_5_4_MINI (or whichever current id the currency map uses) for consistency.

"litellm/client.ts FALLBACK_LITELLM_MODEL must be openai/gpt-5.4-mini (was stale openai/gpt-4o-mini)",
);
assert(
!content.includes('"openai/gpt-4o-mini"'),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

MINOR — LiteLLM assertions cover gpt-4o-mini but not openai/gpt-4o, so the primary fallback regression above slips through.

This suite's LiteLLM test asserts !includes('"openai/gpt-4o-mini"') (this line) and the claude-haiku change, but never checks that getFallbackModels() no longer leads with "openai/gpt-4o". That is precisely why the litellm-fallback-gpt4o finding passes CI: the file can still contain the stale openai/gpt-4o fallback and this test stays green. Strengthen the guard:

Suggested change
!content.includes('"openai/gpt-4o-mini"'),
!content.includes('"openai/gpt-4o-mini"') &&
!content.includes('"openai/gpt-4o"'),
"litellm/client.ts must not contain the stale openai/gpt-4o-mini or openai/gpt-4o literal anywhere",

@Tara-ag

Tara-ag commented Sep 27, 2026

Copy link
Copy Markdown
Contributor

NEEDS_WORK — retirement done right, but one stale fallback and one surface drift remain

The core refactor is sound: FALLBACK_ANTHROPIC_MODEL/FALLBACK_LITELLM_MODEL constants move off the retired model ids, the health-check "try this instead" probes follow suit, and the new e2e suite locks the removal in. Two gaps still contradict the PR's own goal.

Severity File:line Finding
MAJOR src/lib/providers/litellm/client.ts:335-336 getFallbackModels() still leads with the retired openai/gpt-4o; align head with FALLBACK_LITELLM_MODEL
MAJOR src/lib/providers/openAI/client.ts:77 runtime default now gpt-5.4 but modelChoices.ts/providerHealth.ts still advertise gpt-4o; surface diverges
MINOR test/continuous-test-suite-retired-model-defaults.ts:129 litellm guard misses openai/gpt-4o, so the fallback regression above passes CI

Checked and clean

  • Thumb rules — getDefaultModel from modelAliases.ts uses alias, serializable enum, returns standard shapes; thumb model ids weren't part of the retirement surface.
  • Anthropic choice — anthropic/client.ts Maps reference ClaudeHaiku4_5_20251001, matching modelCache.ts currency map; consistent with constants.
  • engineDomainConfig.ts — model class + reason params bound to the new FALLBACK_ANTHROPIC_MODEL/FALLBACK_LITELLM_MODEL, no behavioural drift from the move of claudeHaikuClass.
  • Fake provider retries — killed in test via requestId (allowlist provider health) + keyed batches; no inflaction.
  • Graph/blast-radius — no callers of getFallbackModels/getOpenAIModel outside the providers; low blast radius; no hot-path (baseProvider/registry/types) touched in this diff.

Note for the author

Reply to the two MAJORs confirming whether the additional model-id surfaces (modelChoices.ts, providerHealth.ts) are intentionally out of scope — if so, a line in the PR description explaining why would suffice; otherwise align them.

@Tara-ag Tara-ag left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

NEEDS_WORK — two MAJOR findings remain (see inline comments): the LiteLLM fallback still leads with the retired openai/gpt-4o, and the OpenAI runtime default diverges from the modelChoices.ts/providerHealth.ts surfaces. Please address before merge.

@murdore
murdore merged commit ab2ab06 into release Sep 27, 2026
27 of 28 checks passed
@murdore
murdore deleted the fix/retired-model-defaults branch September 27, 2026 05:06
@github-actions

Copy link
Copy Markdown
Contributor

🎉 This PR is included in version 12.34.0 🎉

The release is available on:

Your semantic-release bot 📦🚀

murdore added a commit to juspay/neurolink-blog that referenced this pull request Sep 28, 2026
…PAA claims

Re-audit of the 173-post corpus (prompted by a hunch the last audit missed
some) found 10 posts still citing stale or retired model IDs and 4 flagged
for SOC2/HIPAA wording. On close read, only 5 files needed real changes:

- model-evaluation-scoring.md cited gemini-1.5-flash as NeuroLink's judge
  fallback -- that model is hard-retired (404s). Now gemini-2.5-flash,
  matching the SDK fix in juspay/neurolink#1823.
- openai-integration-guide.md and litellm-unified-routing.md cited the old
  gpt-4o/gpt-4o-mini/claude-3-haiku fallbacks -- also matching the same SDK
  fix; now gpt-5.4/gpt-5.4-mini/claude-haiku-4-5.
- provider-abstraction.md had one stale "legacy gpt-4o" clause.
- ai-ethics-responsible-use.md's two SOC2/HIPAA mentions implied NeuroLink
  could grant compliance; reworded to state plainly that NeuroLink holds no
  such certification and describe the real building blocks it provides.

The other 9 flagged files were read in full and left untouched: their model
mentions are legitimately historical/legacy-labeled, and 3 of the 4
SOC2/HIPAA flags were already honestly worded (attributing certification to
third-party providers or explicitly disclaiming NeuroLink's own).
murdore added a commit that referenced this pull request Oct 3, 2026
Each change below answers a review thread on an already-merged PR where the
reviewer's claim held on the current tree. Where behaviour changed, the suite
that guards it was made to fail without the change (reversed or mutated, then
restored) and pass with it.

Source
- litellm/client.ts: getFallbackModels() still offered the stale openai/gpt-4o;
  it now offers openai/gpt-5.4. retired-model-defaults asserts both the
  absence of the old literal and the presence of the new one.
  (F-test-misses-litellm-gpt4o, #1823)

Suites
- providers-mocked: the OpenAI-compat banner and header named seven providers
  and no Mistral although the loop runs every catalog entry plus Cohere
  (T3806479869, T3807355730, #1353: one defect raised twice). The detached-pump
  subprocess case now also requires at least one fetch call, exiting 5 when the
  right RateLimitError arrives without a request (PF-T3847750289, #1531). The
  pre-aborted Perplexity case now expects no request, see mockFetch below.
- realtime-unit: clearHandlers test reads the registry from inside
  disconnect(), so it fails if the registry is cleared first
  (T3813998725, #1354).
- stt-unit: captures the registry before the first test and restores it at the
  end, so the e2e-stt-suite-* handlers no longer outlive the file
  (T3813998716-b, #1354).
- vertex-loop-characterization: the stand-in now records headers and answers
  401 without the Express Mode key, the Express case asserts the key header,
  the tool round trip asserts the payload equals { result: { found: true } },
  and a new case covers a blank or whitespace baseURL falling through to
  GOOGLE_VERTEX_BASE_URL (T3827707056-a, T3827707069-a, T3827842925-a, #1408).
- video-no-ffprobe: the PATH is now only the ffmpeg link directory and node's
  directory; /usr/bin and /bin, where a distro ffprobe lives, no longer make
  both cases skip before asserting. The ffprobe preflight checks those two
  directories directly and the Skip stays (T4135201681, #1861).
- acceptance-gate: credentialFreeEnv now also strips names that carry a secret
  without saying key or token (service-account and private keys, speech keys,
  OTLP headers, REDIS_URL, auth config, SSH_AUTH_SOCK) and ambient
  *_BASE_URL, *_ENDPOINT, OTEL_EXPORTER_OTLP_* and AWS_PROFILE. Because the
  gate's own check reused the strip pattern, ambientSecretCanaries plants a
  literal list of dummy values before the strip and the gate fails if any
  survives (T4126861003-env-isolation, #1849).
- harness: withCaseTimeout now scales by NEUROLINK_TEST_TIMEOUT_SCALE like
  defineSuite does, so the JSDoc claim that the two cannot drift is true
  (T3837277828-timeout-scale-drift, #1487). Both go through scaleTimeoutMs,
  which floors a budget at 1ms so a tiny valid scale cannot round it to 0
  (T4053369060-1, #1732). harness-offline-timeout covers both.
- mockFetch: an already-aborted signal is rejected before the call is recorded,
  as real fetch does (F-mockfetch-aborted-records-call, #1357).

Docs and tooling
- model-not-found-retryable makes a real generate() on Anthropic and OpenAI and
  skips without both keys; it moves from test:unit to test:live in package.json,
  test/README.md and the CI comments (T3790920852, #1334).
- sse-bisection-findings.md no longer claims created, in_progress and
  output_item.done are required; only output_item.added before the deltas was
  isolated (F4-T4087478004-minimum-event-overclaim, #1783).
- acceptance-gate.md describes the wider strip and the canary check;
  docs-site search index regenerated.

Not changed
- T3813998753 (#1354): video-generation-unit sets OPENAI_API_KEY and never
  restores it. Deliberate (the file's own comment says why), runSuite() calls
  process.exit, CI runs each suite in its own process and nothing imports the
  file, so nothing can observe the leaked value.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants