Skip to content

fix(ollama): route models by advertised capability (#11087) — port of #11088 to the release line - #11271

Merged
diegosouzapw merged 3 commits into
diegosouzapw:release/v3.8.50from
yourspraveen:fix/11087-ollama-capabilities-release
Aug 23, 2026
Merged

diegosouzapw merged 3 commits into
diegosouzapw:release/v3.8.50from
yourspraveen:fix/11087-ollama-capabilities-release

Conversation

@yourspraveen

@yourspraveen yourspraveen commented Aug 23, 2026 •

Copy link
Copy Markdown
Contributor

⚠️ base-red inherited: #9985

What

Ports #11088 (merged) onto the active release line. Closes the gap you flagged on #11087:

it landed on main, not on the release/v3.8.50 branch, so it will NOT be in the v3.8.50 package about to publish; it rides the following release line (v3.8.51+)

Same precedent as #11165, which ported #11075 from main onto release/v3.8.50.

Why this is worth porting rather than waiting

release/v3.8.51 does not exist yet, and every PR currently open — 30 of them — is cut from release/v3.8.50. Until the next cycle branch appears, the fix is reachable only from main, which no active branch is based on. Meanwhile the defect is live on the release line: every Ollama Local model is flattened to chat, so /v1/embeddings and /v1/images/generations reject models the daemon advertises as capable.

If you would rather this wait for the v3.8.51 branch, say so and I will close it — the branch will keep.

⚠️ No longer a byte-identical port — #11088 carries two defects

This PR started as a byte-identical port and I described it that way below. CI proved that wrong, so the description is corrected here rather than quietly left standing. Two failures showed up that are not in the base-red set, reproduce on main (which already carries #11088), and do not reproduce on the pristine release tip — i.e. they are defects in the merged change, not in the port:

1. OpenAI import leaks image/video models into chat selections — real behavioural regression.

✖ OpenAI import excludes image and video generation models from chat selections
  actual:   [ 'gpt-5.6-sol', 'gpt-image-2', 'sora-2-pro', 'vendor-image-model' ]
  expected: [ 'gpt-5.6-sol' ]

#11088 dropped filterChatSelectableModels from importManagedModels for every provider, but the compensating read-time filter (resolveLocalSyncedEndpointRoute) is gated on isSelfHostedChatProvider. Non-local providers therefore lost the filter with nothing replacing it, and the models are persisted, not merely returned. Fixed in c018e3a45 by scoping the drop to self-hosted providers — the local behaviour #11087 asked for is untouched and the Ollama suite stays 3/3.

2. An unregistered credential site — hard-lease credential, executor, and connection-query inventory has no unclassified site (src/lib/embeddings/service.ts 2 → 3). The new site resolves through getProviderCredentials with the connection allowlist from resolveLocalSyncedEndpointRoute and handles allRateLimited, so it is fenced like its two siblings — the ledger had simply not been told. Registered with a comment saying why.

main still has both. This PR fixes them only on the release line; main needs the same two-line change. Flagging rather than opening a second PR unasked — say the word and it is up in minutes.

Fidelity (as originally written — see the correction above)

The change set is byte-identical to the merged #11088 — I diffed all 11 files against refs/pull/11088/head after applying, and every one matches:

open-sse/services/combo/autoStrategy.ts                     |   6 +-
src/app/api/providers/[id]/models/discovery/helpers.ts      |  38 ++++
src/app/api/providers/[id]/models/route.ts                  |   3 +
src/app/api/v1/images/generations/route.ts                  |  19 +-
src/lib/embeddings/service.ts                               |  59 +++++-
src/lib/providerModels/managedModelImport.ts                |   6 +-
src/lib/providerModels/modelDiscovery.ts                    |   9 +-
src/lib/providerModels/ollamaCapabilities.ts                |  98 ++++++++++
src/lib/providerModels/syncedEndpointRouting.ts             |  31 ++++
tests/unit/ollama-local-capabilities-routing.test.ts        | 205 +++++++++++++

The ported commit re-derived nothing: applied with git apply --3way at the release tip a7e09eda5, clean, no conflicts. The two-file fix above sits in a separate commit on top (c018e3a45), so the port and the correction stay reviewable independently.

Deliberately not carried over: docs/superpowers/plans/2026-08-23-qdrant-configuration-guidance.md and its specs/ sibling. Those are on main only because they rode in with #11088's squash — they came from #11213 and 855243ab1 already removed them from the release branch. They do not belong here (or on main — see the separate cleanup PR).

The changelog fragment is renumbered to this PR in a follow-up commit, matching how #11165 handled its own fragment.

Verification at the release tip

Check Result
tests/unit/ollama-local-capabilities-routing.test.ts 3/3 pass
npm run typecheck:core clean
npx eslint on all 10 changed source/test files clean¹

¹ One no-restricted-imports hit in src/app/api/providers/[id]/models/route.ts — the pre-existing @/lib/localDb import at line 14, already frozen in config/quality/eslint-suppressions.json with count: 1. Untouched by this port; npm run lint (which applies the suppressions, and is what CI runs) is clean.

Fixes #11087 on the release line.

Ports diegosouzapw#11088 onto the active release line. diegosouzapw#11088 merged to main only, so
release/v3.8.50 — and every branch cut from it — still flattens every Ollama
Local model to chat and rejects valid embedding/vision models on the specialty
routes.

The change set is byte-identical to the merged diegosouzapw#11088 (11 files, 462
insertions): the synced store now persists non-chat models, chat filtering
moves to read time, and /v1/embeddings + /v1/images/generations resolve
against advertised capabilities instead of the flattened chat list.

Verified at the release tip a7e09ed: ollama-local-capabilities-routing 3/3,
typecheck:core clean, eslint clean on all 10 changed files (the one
no-restricted-imports hit in models/route.ts is the pre-existing localDb
import already frozen in eslint-suppressions.json, count unchanged).
Hermes Developer added 2 commits August 23, 2026 13:28
…oviders

diegosouzapw#11088 removed filterChatSelectableModels from importManagedModels for EVERY
provider, but the compensating read-time filter (resolveLocalSyncedEndpointRoute)
is gated on isSelfHostedChatProvider. So non-local providers lost the filter with
nothing replacing it: an OpenAI sync now persists gpt-image-2, sora-2-pro and any
model declaring /v1/images/generations straight into chat selections.

Reproduced on main (which carries diegosouzapw#11088) as well as on this branch, and NOT on
the pristine release tip — so it is a defect in the merged change, not in the
port. Verified by tests/unit/managed-model-import.test.ts, which asserts both
result.discoveredModels and the persisted getSyncedAvailableModels("openai").

Self-hosted providers keep the new behaviour — non-chat models still persist and
are filtered at read time — so diegosouzapw#11087's goal is untouched: the Ollama capability
suite stays 3/3.

Also registers the third getProviderCredentials site diegosouzapw#11088 added in
embeddings/service.ts with the hard-lease inventory ledger (2 -> 3). The site
resolves through getProviderCredentials with the connection allowlist from
resolveLocalSyncedEndpointRoute and handles allRateLimited, so it is fenced like
the two existing sites; the ledger simply had not been told.
@yourspraveen

Copy link
Copy Markdown
Contributor Author

CI triage — 17 inherited, 1 flake, and 2 that were genuinely mine

The run came back red on all four shards. Breaking it down honestly, because two of these were not base-red:

Inherited (17 assertions). Identical to the set reproduced on the pristine tip and reported at #9985 — CLI_TOOLS ×5, Codex 872000 !== 272000 ×6, chatCore combo skip, search-route, i18n pt-BR ×2, settings-i18n-keys, check-deps. Also hits the maintainer's own #11262 unchanged.

Mine (2 assertions). managed-model-import and hard-session-lease-bypass-inventory. I verified these three ways before concluding anything:

pristine tip a7e09eda5 main (carries #11088) this branch
managed-model-import 9/9 8/9 8/9 → now 9/9
hard-session-lease-bypass-inventory 3/3 2/3 2/3 → now 3/3

They fail on main too, which means the port carried them faithfully — they are defects in the merged #11088, not port damage. Details and the main-side reproduction are in #11276; the corrected description is in the PR body above, and the fix is c018e3a45, a separate commit so the port and the correction review independently.

I have to correct something I wrote in the original body: "byte-identical to the merged #11088 … nothing was re-derived" was true of the port commit and is no longer true of the PR. Byte-identical turned out to be the wrong goal once the source change had a bug in it.

One flake, deliberately excluded: ttft() measures first-forwarded-chunk latency failed on one #11146 run and nowhere else — 8/8 on the pristine tip, 3× 8/8 on a branch. Not counted against either PR.

Where this leaves the PR

The two real defects are fixed here and the Ollama capability suite is still 3/3, so #11087's goal is intact. The remaining red is the base, which #9985 tracks and which no PR against this branch can currently escape.

main still carries both defects — #11271 cannot reach it from this base. Say the word and the main-side PR is up in minutes.

@diegosouzapw
diegosouzapw merged commit 6d4c484 into diegosouzapw:release/v3.8.50 Aug 23, 2026
11 of 16 checks passed
diegosouzapw pushed a commit that referenced this pull request Aug 24, 2026
…k the release PR

The living release PR #8875 was CONFLICTING, which makes GitHub skip EVERY
pull_request workflow silently (no ci.yml, no semgrep, no DAST). Back-merging
main restores a computable merge ref.

Strategy `-s ours`: main is a stale snapshot of the release line (PR #11088 was
merged into main from a release-tip base, dragging ~5094 files). All 7 main-only
commits were verified as already represented on this branch:

- #11088 ollama capability routing  -> ported here as #11271 (6d4c484)
- #11075 shared passthrough providers -> ported here as #11165 (92ef3c7)
- #10055 getModelsDevPricing memoization -> present (modelsDevSync.ts)
- #10026 hide health-check excluded models -> present and extended (catalog.ts)
- /_tasks anchored gitignore hardening -> present (.gitignore:288)
- nanoid/dompurify Dependabot bumps -> identical versions

main-only files intentionally NOT carried over:

- changelog.d/fixes/10286-gemini-3-5-flash-thinking.md + its regression test:
  the fix landed here as #10450 and was then deliberately superseded by
  2764812 "eliminate Gemini 3.5 Flash". The test fails on this branch by
  design.
- public/providers/hackclub.svg: provider removed here (migration 162).
- docs/superpowers/**/2026-08-23-qdrant-*: planning artifacts belong in _tasks/
  (AGENTS.md), never under docs/.
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
… — port of diegosouzapw#11088 to the release line (diegosouzapw#11271)

Validated on the combined 8-PR board: ollama-local-capabilities-routing 3/3, managed-model-import 9/9 (including the integration with the carried Gemini-3.5-Flash cleanup from diegosouzapw#11259), 88/88 across the board's focused suites, typecheck:core + dashboard-typecheck clean, gates within baseline. This brings diegosouzapw#11088 to the release line — it had squash-merged to main by base error (mine) — AND fixes the two defects the port caught: the global filter drop that leaked image/video models into OpenAI chat selections (now scoped to self-hosted providers) and the unregistered hard-lease credential site. Exemplary port discipline: byte-identical carry + the corrections in a separate reviewable commit + the superpowers docs deliberately left out. main still needs the same two-line fix. Thank you @yourspraveen!
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…k the release PR

The living release PR diegosouzapw#8875 was CONFLICTING, which makes GitHub skip EVERY
pull_request workflow silently (no ci.yml, no semgrep, no DAST). Back-merging
main restores a computable merge ref.

Strategy `-s ours`: main is a stale snapshot of the release line (PR diegosouzapw#11088 was
merged into main from a release-tip base, dragging ~5094 files). All 7 main-only
commits were verified as already represented on this branch:

- diegosouzapw#11088 ollama capability routing  -> ported here as diegosouzapw#11271 (ec959d8)
- diegosouzapw#11075 shared passthrough providers -> ported here as diegosouzapw#11165 (4813c32)
- diegosouzapw#10055 getModelsDevPricing memoization -> present (modelsDevSync.ts)
- diegosouzapw#10026 hide health-check excluded models -> present and extended (catalog.ts)
- /_tasks anchored gitignore hardening -> present (.gitignore:288)
- nanoid/dompurify Dependabot bumps -> identical versions

main-only files intentionally NOT carried over:

- changelog.d/fixes/10286-gemini-3-5-flash-thinking.md + its regression test:
  the fix landed here as diegosouzapw#10450 and was then deliberately superseded by
  bc825a1 "eliminate Gemini 3.5 Flash". The test fails on this branch by
  design.
- public/providers/hackclub.svg: provider removed here (migration 162).
- docs/superpowers/**/2026-08-23-qdrant-*: planning artifacts belong in _tasks/
  (AGENTS.md), never under docs/.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(providers): Ollama local capabilities are flattened to chat and specialty routes reject valid models

2 participants