Skip to content

fix(model-resolver): preserve @provider:model picks across cold catalogs - #3950

Closed
starship-s wants to merge 4 commits into
nesquena:masterfrom
starship-s:fix/at-provider-model-reset-to-default-pr
Closed

starship-s wants to merge 4 commits into
nesquena:masterfrom
starship-s:fix/at-provider-model-reset-to-default-pr

Conversation

@starship-s

Copy link
Copy Markdown
Contributor

Thinking Path

What Changed

  • Updated _resolve_compatible_session_model_state() so the @provider:model branch no longer defaults away an explicit user pick just because the provider is absent from the current catalog snapshot.
  • Preserved non-first-party provider-qualified selections, such as @ollama-cloud:minimax-m3, when the bare model does not look like a first-party family model (gpt*, claude*, gemini*).
  • Left the cached-catalog behavior unchanged; this PR does not force live catalog rebuilds back onto hot session-display paths.
  • Added regression coverage for:
    • a non-first-party @provider:model surviving a cold/partial catalog on explicit and non-explicit paths;
    • an explicit provider-qualified pick being honored even when the provider is not currently routable.

Why It Matters

  • Prevents WebUI from silently discarding an explicitly selected provider/model on later assistant turns or when switching chats.
  • Keeps live-discovered provider selections stable even when the fast cached catalog is intentionally incomplete.
  • Preserves the existing safety repair for genuinely stale first-party model/provider mismatches.
  • Avoids regressing the cached-catalog performance fix that prevents slow live provider discovery from blocking hot session reads.

Verification

Targeted resolver regression coverage:

python -m pytest tests/test_provider_mismatch.py -q
# 68 passed

Adjacent model/session resolver coverage:

python -m pytest \
  tests/test_provider_mismatch.py \
  tests/test_issue1855_resolve_model_provider_fast_path.py \
  tests/test_session_display_resolver_no_live_rebuild.py \
  tests/test_wakeup_model_resolve_hang.py \
  tests/test_chat_start_provider_fallback.py \
  tests/test_issue2518_active_provider_fallback.py \
  tests/test_issue3405_profile_provider_resolution.py -q
# 144 passed

Hygiene checks:

git diff --check origin/master...HEAD
python -m compileall -q api/routes.py
# passed

Risks / Follow-ups

  • If a user removes a non-first-party provider from config entirely, an old @provider:model session can now remain selected and fail at runtime instead of being silently replaced by the default. That is intentional here: the cached catalog cannot reliably distinguish “provider removed” from “provider configured but not present in this snapshot,” and preserving the explicit selection is less surprising than silently changing it.
  • A more precise follow-up could consult configured provider IDs from config, not the live catalog, before preserving missing non-first-party providers. That is intentionally left out to keep this fix narrow.

Model Used

  • Anthropic / Claude Opus 4.8 for implementation.
  • OpenAI / GPT-5.5 for review, verification, and PR description drafting.

… default on cold catalog

The @Provider:model branch of _resolve_compatible_session_model_state()
falls through to `return default_model` whenever the hinted provider is
absent from the catalog snapshot it was handed. For a non-first-party
provider (ollama-cloud, deepseek, xai, … — these normalize to "" and
discover their models live), the cached/minimal catalog used on the hot
GET /api/session path frequently lacks the group even though the provider
is configured. The result: an explicitly-picked model like
`@ollama-cloud:minimax-m3` silently snaps back to the global default on
the 2nd+ turn and on chat switch.

This was latent since v0.50.224 but became live on 2026-06-08 (v0.51.340,
26e133e), which switched the GET /api/session display resolver to
prefer_cached_catalog=True for performance — exposing the blind spot.

Fix: before the default-revert, preserve the selection when either
(a) it was an explicit user pick (mirrors the bare-model branch's nesquena#3737
guard), or (b) the provider hint is non-first-party (provider_normalized
== "") AND the bare id is not a first-party family name (gpt/claude/
gemini). This matches the slash-qualified branch, which already passes
models whose provider normalizes to "" through unchanged. A genuinely
stale misrouted first-party model (e.g. @copilot:claude-opus-4.6 after
switching to openai-codex) still falls through to the default-repair.

No change to catalog loading — the prefer_cached_catalog performance win
is fully retained; only the fallback decision is corrected.

Adds regression coverage in tests/test_provider_mismatch.py.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@starship-s
starship-s marked this pull request as ready for review June 10, 2026 20:47
@greptile-apps

greptile-apps Bot commented Jun 10, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

Adds _provider_is_known_or_configured() to api/config.py and threads it into _resolve_compatible_session_model_state() in api/routes.py to prevent @provider:model selections from being silently reverted to the default when the provider's catalog group is absent from a cold/cached snapshot.

  • api/config.py: New helper that decides provider "knowness" from the static _PROVIDER_DISPLAY/_PROVIDER_MODELS registries plus configured custom providers, deliberately bypassing any live catalog rebuild.
  • api/routes.py: Two new guard branches — an explicit-pick early return (above the family-match repair) and a non-explicit preservation block gated on not provider_normalized && not _bare_is_first_party_family && _provider_is_known_or_configured().
  • tests/test_provider_mismatch.py: Six new regression tests covering explicit-pick honor, cold-catalog survival, known-but-unconfigured built-ins, removed-provider reversion, and the documented gpt*-prefix false-positive limitation.

Confidence Score: 5/5

Narrow, targeted fix to the @Provider:model resolution path; the explicit-pick early return is unconditional and correct, and the non-explicit preservation block is gated on three independent conditions, all with test coverage.

The new helper consults only static registry data so it cannot trigger live catalog rebuilds. All edge cases that could cause silent model swaps have explicit regression tests pinning the behavior. No existing code paths are removed or reordered except for the insertion of the explicit-pick guard above the family-match repair, which is the intended fix.

No files require special attention. The known prefix-heuristic limitation in api/routes.py is documented in-code and pinned by a test.

Important Files Changed

Filename Overview
api/config.py New _provider_is_known_or_configured() helper is well-scoped: checks static registry + config state only, with clear rationale for excluding credential/auth evidence.
api/routes.py Two targeted additions to _resolve_compatible_session_model_state: explicit-pick early return and non-explicit preservation block, both with correct ordering and documented limitations.
tests/test_provider_mismatch.py Six new tests cover the full matrix of explicit vs. non-explicit, known vs. removed provider, family-name collision, and the documented prefix-heuristic limitation.

Flowchart

%%{init: {'theme': 'neutral'}}%%
flowchart TD
    A["@provider:model input"] --> B{provider_raw or bare_model empty?}
    B -- yes --> C[return model unchanged]
    B -- no --> D{explicit_model_pick?}
    D -- yes --> E["NEW: return model, provider_raw, False"]
    D -- no --> F{hint matches active provider?}
    F -- yes --> G[return model, provider_raw, False]
    F -- no --> H{catalog has provider?}
    H -- yes --> I[return model, provider_raw, False]
    H -- no --> J{bare model matches active family?}
    J -- yes --> K[strip prefix and reroute to active]
    J -- no --> L{"NEW: not provider_normalized AND not first-party name AND provider known/configured?"}
    L -- yes --> M["NEW: return model, provider_raw, False"]
    L -- no --> N{default_model available?}
    N -- yes --> O[revert to default]
    N -- no --> P[return model, provider_raw, False]
Loading

Reviews (4): Last reviewed commit: "review: honor explicit @provider picks a..." | Re-trigger Greptile

Comment thread api/routes.py
Comment thread tests/test_provider_mismatch.py Outdated
Addresses PR review (PR nesquena#3950):

- Rewrote the @Provider:model guard comment in
  _resolve_compatible_session_model_state to be self-contained and to call
  out the KNOWN LIMITATION of the bare-name prefix heuristic: a third-party
  model whose name starts with gpt/claude/gemini (e.g. "@ollama:gpt4all-mini")
  is still mis-classified as first-party and reverted on non-explicit paths.
  A name-only check can't disambiguate this; only configured-provider-aware
  resolution could. (No heuristic swap: the static-catalog helper is stale —
  it lacks current models like claude-opus-4.8 — so it would trade this
  false-positive for worse false-negatives on genuine first-party models.)

- Added test_at_provider_first_party_named_third_party_model_known_limitation
  to pin that boundary so the limitation is tracked, not silent (and to show
  an explicit pick still escapes the heuristic).

- The explicit-pick test now also asserts the returned provider ("copilot"),
  confirming the @-qualified hint is preserved rather than rewritten.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@starship-s

Copy link
Copy Markdown
Contributor Author

Thanks for the careful review — all fair. Pushed f6f1c50 addressing points 1–3; declining the heuristic swap, with rationale below.

The gpt4all-mini false positive is real (reproduced: @ollama:gpt4all-mini still reverts on non-explicit paths). Two notes for context: it's a strict improvement over the prior behavior (which reverted all non-first-party @provider models, not just the gpt/claude/gemini-named subset), and the prefix test is the existing convention in this file — _model_matches_active_provider_family uses the identical ("gpt","claude","gemini") startswith match and shares the same blind spot.

Why I didn't switch to the precise helper. I evaluated _is_first_party_model() / _PROVIDER_MODELS, and it's a net downgrade here — the static catalog is stale:

model startswith heuristic in static catalog
claude-opus-4.6 first-party ✓ first-party ✓
claude-opus-4.8 (current default) first-party ✓ not found ✗
gpt-5.5-codex first-party ✓ not found ✗
gpt4all-mini first-party (false +) third-party ✓

So swapping would fix gpt4all but start under-reverting genuinely-stale models newer than the catalog (@copilot:claude-opus-4.8 would be preserved instead of repaired), and would break the @copilot stale-repair test for current models. The root issue — copilot and ollama both normalize to "", and a model name can't distinguish "misrouted first-party" from "genuine third-party" — is only fully solvable by consulting the user's configured providers. That's a larger change; happy to file a follow-up issue for it.

What's in f6f1c50:

  1. Rewrote the guard comment to be self-contained and to call out the KNOWN LIMITATION explicitly, pointing at the pinning test.
  2. Added test_at_provider_first_party_named_third_party_model_known_limitation — documents the gpt4all-mini boundary so it's tracked, not silent, and verifies an explicit pick still escapes the heuristic.
  3. The explicit-pick test now also asserts provider == "copilot" (the @-hint is preserved, not rewritten).

No CHANGELOG entry included — leaving that to the release stamping process, but glad to add one if preferred.

@nesquena-hermes

Copy link
Copy Markdown
Collaborator

Thanks @starship-s — the concept is right (explicit @provider:model picks shouldn't snap back to the default on a cold catalog snapshot, mirroring #3737 / #1131 / #3867), and the surgical approach of narrowing only the fallback decision while keeping the cached-catalog hot path is exactly the right shape. I picked this up to absorb into a release and ran it through the full gate — it's almost there, but a regression review (Codex, GPT-5.5) caught one CORE issue I want fixed before it ships, plus I tightened one thing on the way in.

CORE — must fix: a genuinely-removed provider now routes to an unconfigured provider

The non-explicit guard preserves the selection for any provider whose group is absent from the catalog snapshot:

if explicit_model_pick or (not provider_normalized and not _bare_is_first_party_family):
    return model, provider_raw, False

But catalog-absence covers two distinct cases, and only one of them should be preserved:

  1. Cold / live-discovery provider — ollama-cloud is configured, its group just isn't folded into this snapshot yet → preserve ✓ (the bug you're fixing).
  2. Genuinely removed / unconfigured provider — the user's session points at @removed:mistral-large but that provider is no longer in config at all → this should still fall through to the default-repair, but the new guard now preserves it, so chat/start routes to a provider that isn't configured → broken turn.

I verified this empirically against the staged branch:

active_provider=anthropic, default_model=claude-opus-4.8, no "removed" group
_resolve_compatible_session_model_state("@removed:mistral-large", "removed", explicit_model_pick=False)
  master  → ("claude-opus-4.8", ..., changed=True)   # repairs to default ✓
  this PR → ("@removed:mistral-large", "removed", changed=False)   # ✗ routes to a gone provider

Suggested fix: split the guard so explicit_model_pick is always honored (the user just chose it), but the non-explicit not provider_normalized preservation only applies when provider_raw is still a configured/known provider — checked via config state (e.g. a configured-providers / custom_providers lookup), not via the catalog snapshot (which is exactly what's cold here, and re-deriving it live would defeat the perf win you're preserving). A genuinely-unconfigured missing provider should fall through to default_model.

Please also add a regression test pinning @removed:mistral-large (omitted catalog group, provider not configured) → reverts to the active default, alongside your existing @ollama-cloud:minimax-m3 cold-catalog preservation test, so the two cases are both locked down.

Already handled on my side (no action needed from you)

  • First-party-family detection tightened. Your bare_model.startswith("gpt"/"claude"/"gemini") would also catch legit third-party names like gptq-3b (and revert them on the non-explicit path — the exact loss this PR removes). I changed it to a token match (?:gpt|claude|gemini)(?:[-._]|\d|$) (using module-level re), which is a strict subset of your startswith — it can only convert revert-cases into preserve-cases, never weaken the stale first-party repair. gptq-3b is now correctly third-party; gpt-5.5/gpt4o/claude-opus-4.6 still classify first-party. (Residual: gpt4all/gpt-neo still match because they're lexically indistinguishable from gpt4/gpt- first-party shapes without a catalog lookup — same limitation as the existing _model_matches_active_provider_family helper; acceptable.) I'll fold this in once the CORE fix lands.

Reopen for re-gate whenever the configured-provider guard is in — happy to take it straight to release at that point. 🙏

@nesquena-hermes nesquena-hermes added the changes-requested Maintainer left detailed feedback requesting changes; PR is waiting on author to address label Jun 10, 2026
…nown/configured providers

Addresses the CORE review finding on PR nesquena#3950: the non-explicit guard
preserved an explicitly-qualified selection for ANY provider absent from the
catalog snapshot, including a genuinely removed/unconfigured one — so a stale
session pointing at e.g. "@removed:mistral-large" would route chat/start to a
provider that no longer exists instead of repairing to the default.

Split the guard:
  * explicit_model_pick is always honored on its own (a fresh, deliberate pick).
  * the non-explicit cold-catalog preservation now also requires the provider to
    be known or configured, via the new config-state helper
    _provider_is_known_or_configured() — decided from the static provider
    registry + custom_providers config, NOT from the cold catalog snapshot
    (re-deriving that live would defeat the prefer_cached_catalog hot-path win).

So a cold live-discovery provider (ollama-cloud configured, group missing from
this snapshot) is still preserved, while a genuinely removed/unknown provider
falls through to the default-repair.

Adds test_at_provider_removed_provider_still_reverts_to_default pinning both the
non-explicit revert and the explicit-pick escape hatch.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@starship-s

Copy link
Copy Markdown
Contributor Author

Fixed the CORE issue — pushed 86d8157. You're right: catalog-absence conflated "cold but configured" with "genuinely removed," and the guard preserved both.

Split the guard exactly as suggested:

# explicit pick is always honored on its own
if explicit_model_pick:
    return model, provider_raw, False
# non-explicit preservation now also requires the provider to be known/configured
if (not provider_normalized
        and not _bare_is_first_party_family
        and _provider_is_known_or_configured(provider_raw)):
    return model, provider_raw, False
# else → fall through to default-repair

New helper _provider_is_known_or_configured() (in api/config.py) decides this from config state only — named custom_providers slugs + the static provider registry (_PROVIDER_DISPLAY / _PROVIDER_MODELS, alias-resolved). It deliberately does not touch get_available_models() / catalog groups, since those are exactly what's cold here — re-deriving them live would defeat the prefer_cached_catalog win this preserves. So ollama-cloud (known, live-discovery) stays preserved while removed (unknown, unconfigured) reverts.

Verified against your empirical case:

active_provider=anthropic, default_model=claude-opus-4.8, no "removed"/"ollama-cloud" group
  @removed:mistral-large     non-explicit → claude-opus-4.8  (revert ✓)
  @removed:mistral-large     explicit     → @removed:mistral-large  (honored ✓)
  @ollama-cloud:minimax-m3   non-explicit → @ollama-cloud:minimax-m3  (preserve ✓)

Added test_at_provider_removed_provider_still_reverts_to_default pinning both the non-explicit revert and the explicit-pick escape hatch, alongside the existing @ollama-cloud cold-catalog preservation test.

I left the first-party-family startswith heuristic untouched since you said you're folding in the token-match regex on your side — happy to rebase onto that once it's in, or adjust if you'd rather I take it. Full test_provider_mismatch.py is green (70), and the broader provider/config/resolver sweep passes (990). Ready for re-gate. 🙏

@nesquena-hermes

Copy link
Copy Markdown
Collaborator

Thanks @starship-s — the guard split is exactly right and the new _provider_is_known_or_configured() helper resolves the CORE issue from the last round cleanly. I verified empirically: @removed:mistral-large non-explicit now reverts to default ✓, @ollama-cloud:minimax-m3 is preserved ✓, explicit removed-pick is honored ✓.

I staged this for release alongside another fix and ran the full gate. Both reviewers cleared the prior CORE finding; one residual edge came up that I'd like tightened before it ships (the other reviewer judged it acceptable-as-designed, so I'm flagging it as a should-fix, not a hard block — your call on the approach):

Should-fix: a built-in provider whose credentials were removed is still preserved

_provider_is_known_or_configured() returns True for any provider in the static _PROVIDER_DISPLAY / _PROVIDER_MODELS registry — which includes built-ins like deepseek / minimax / ollama-cloud even when the user hasn't configured a key for them. So with an Anthropic-only setup, a stale @deepseek:deepseek-v4-pro pick is preserved on a non-explicit resolve and would route to an unconfigured provider:

active=anthropic only, no deepseek config:
  @deepseek:deepseek-v4-pro  (non-explicit) -> PRESERVED  (arguably should revert)

The docstring deliberately chooses "known builtin → preserve so the user gets a clear runtime error rather than a silent model swap," and that's a defensible position (one reviewer accepted it as-designed). But the stricter reading is that catalog-absent preservation should require actual configured/authenticated provider evidence (an API key in .env, a config.yaml provider entry, or a custom-provider slug) — not static-registry membership alone — so a genuinely-unconfigured built-in falls back to the default like a removed one.

Two acceptable resolutions (your pick):

  1. Tighten _provider_is_known_or_configured() to require config/auth/env/custom-provider evidence for built-ins, not bare registry membership; add a regression test that @deepseek:... with no deepseek config reverts to the active default on non-explicit resolve.
  2. Document the deliberate choice in the helper + add a test pinning that a known-but-unconfigured builtin is intentionally preserved (user sees a runtime error, not a silent swap), so the behavior is locked and reviewed rather than incidental.

I lean toward (1) (revert is the safer default for a provider that demonstrably can't run), but (2) is legitimate if you'd rather surface a clear error.

Minor (non-blocking, pre-existing): explicit-pick guard sits below the family-repair branch

The new comment says explicit picks are "never second-guessed," but the if explicit_model_pick: return ... was inserted after the _model_matches_active_provider_family repair — so an explicit @ollama-cloud:gpt-oss-120b under active openai + cold catalog still gets rerouted to OpenAI. Narrow (needs a cold catalog at the instant of an explicit pick) and the branch order predates this PR, but since you're touching this path: consider moving if explicit_model_pick: return model, provider_raw, False above the family-match repair (a fresh pick is by definition not a stale artifact), or correcting the comment.

Reopen for re-gate whenever the config-evidence tightening is in — this is close. 🙏

…erate known-builtin preservation

Two follow-ups from PR nesquena#3950 review:

MINOR (branch order): the `if explicit_model_pick: return` guard sat *below* the
_model_matches_active_provider_family repair, so an explicit pick like
"@ollama-cloud:gpt-oss-120b" under an OpenAI-active agent was stripped to bare
"gpt-oss-120b" and rerouted to OpenAI (the family match fires on the "gpt"
prefix). Moved the explicit-pick guard to the top of the @Provider:model branch
— a fresh, deliberate pick is by definition not a stale artifact and must never
be rerouted. Adds test_at_provider_explicit_pick_not_rerouted_by_family_match.

SHOULD-FIX (known-but-unconfigured builtin, Option 2 — document + pin): the
reviewer noted a known builtin (deepseek/minimax/ollama-cloud) is preserved on a
cold catalog even when no key is configured for it. We keep this behavior on
purpose: the only fully-reliable "is this provider authenticated" signal is the
live auth store / catalog rebuild — exactly the cost this hot path avoids — and a
cheap env/config-only credential check would mis-classify OAuth/auth-store
providers (ollama-cloud among them) and re-introduce the original silent-revert
bug. A known-but-unconfigured pick is kept and surfaces a clear run-time auth
error rather than a silent swap. Documented the deliberate scope in the helper
docstring + guard comment, and pinned it with
test_at_provider_known_unconfigured_builtin_is_intentionally_preserved.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@starship-s

Copy link
Copy Markdown
Contributor Author

Confirmed both findings empirically and pushed ca58486.

Minor (explicit pick below family-repair) — fixed

Reproduced exactly:

active=openai, cold catalog, explicit pick "@ollama-cloud:gpt-oss-120b"
  before → ("gpt-oss-120b", None, changed=True)   # rerouted to OpenAI by the "gpt" family match

Moved if explicit_model_pick: return model, provider_raw, False to the top of the @Provider:model branch, above _model_matches_active_provider_family. A fresh pick is by definition not a stale artifact, so it's now honored verbatim before any repair. Comment corrected. Added test_at_provider_explicit_pick_not_rerouted_by_family_match.

Should-fix (known-but-unconfigured builtin) — went with Option 2 (document + pin)

Confirmed: @deepseek:deepseek-v4-pro under an Anthropic-only setup is preserved on non-explicit resolve. After digging into the alternative, I deliberately chose document + pin over tightening, because tightening has a real regression risk back into the original bug:

  • The only fully-reliable "is this provider authenticated" signal is the live auth store / catalog rebuild — exactly the cost this hot path avoids.
  • A cheap env/config-only credential check would mis-classify providers configured via OAuth / auth-store, and ollama-cloud is one of them: OLLAMA_API_KEY isn't in config.py's built-in env-detection list (it's detected via the auth-store path), so a config/env-only check would revert @ollama-cloud:minimax-m3 — re-breaking the exact report that started this PR.

So a known-but-unconfigured pick is kept on purpose; the user gets a clear run-time auth error rather than a silent swap to the default (the lesser evil, and consistent with respecting an explicit pick). Documented the deliberate scope in the _provider_is_known_or_configured() docstring + the guard comment, and pinned it with test_at_provider_known_unconfigured_builtin_is_intentionally_preserved.

If you'd rather take the stricter revert-when-unconfigured semantics despite that risk, I'm happy to do Option 1 + auth-store check (consult the auth store for OAuth providers so ollama-cloud stays detected) — it's more correct but reintroduces a cached auth lookup on the hot path, so I left it out unless you want it. @removed:mistral-large still reverts; full test_provider_mismatch.py green (72), broader sweep green (856). Ready for re-gate. 🙏

@nesquena-hermes

Copy link
Copy Markdown
Collaborator

Absorbed and shipped in v0.51.354 (Release LR, deployed live) via the batched release #3951 — rebased onto fresh master with attribution, full-suite (8588) + Codex + Opus gated, both SAFE.

Across three review rounds you nailed it: the guard split + _provider_is_known_or_configured() cleanly distinguishes a cold live-discovery provider (preserve) from a genuinely-unknown one (revert), and the documented decision to preserve a known-but-unconfigured builtin (clear runtime auth error vs silent swap, avoiding the OAuth/auth-store mis-classification a cheap credential check would cause) resolved the reviewer split well — and it's pinned with tests. Fixes the @ollama-cloud:minimax-m3 snap-to-default report.

Thanks @starship-s! 🙏

SysAdminDoc pushed a commit to SysAdminDoc/hermes-webui that referenced this pull request Jun 26, 2026
Release LR — v0.51.354 (nesquena#3950 preserve @Provider:model picks across cold catalogs)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

changes-requested Maintainer left detailed feedback requesting changes; PR is waiting on author to address

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants