Skip to content

[v3.8.50] fix: enforce OpenAI model lifecycle without silent reroutes - #8627

Merged
diegosouzapw merged 1 commit into
diegosouzapw:release/v3.8.50from
backryun:refactor/model-lifecycle
Aug 12, 2026
Merged

diegosouzapw merged 1 commit into
diegosouzapw:release/v3.8.50from
backryun:refactor/model-lifecycle

Conversation

@backryun

@backryun backryun commented Jul 26, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Add a provider-scoped OpenAI model lifecycle policy based on the official OpenAI API deprecations page.
  • Reject already-shut-down OpenAI model IDs before any upstream request with a deterministic HTTP 410 / model_shutdown error and replacement guidance.
  • Keep replacements advisory: lifecycle policy never silently rewrites a request to a different model.
  • Keep image/video generation models out of chat-only import and picker flows when upstream /models metadata is absent or historically persisted as synthetic chat metadata.
  • Constrain model-family fallback to the resolved provider's currently registered, lifecycle-selectable models; remove the obsolete global OpenAI GPT fallback chains.

Why

OmniRoute currently combines upstream discovery, local catalogs, aliases, and family fallback. Two failure modes can follow:

  1. Retired OpenAI IDs can remain selectable and then fail upstream, or be silently redirected by a legacy alias/fallback to a model with different behavior, capability, or cost.
  2. OpenAI's /models response does not reliably describe which endpoint owns each model. Image and video IDs can therefore be imported into chat selectors merely because OmniRoute historically defaulted missing endpoint metadata to chat.

This PR centralizes those decisions while keeping them provider-scoped and projection-scoped.

Compatibility / breaking-change assessment

This is an intentional correction of invalid routing state, not a public API or storage migration.

Surface Existing contract retained Intentional correction
Provider models API GET /api/providers/:id/models remains broad by default. No endpoint or response field is removed. chatOnly=true applies lifecycle and chat-endpoint filtering for chat consumers.
Chat import and picker Ordinary chat models and non-OpenAI providers keep their existing flow. Known OpenAI image/video IDs and deprecated/shut-down IDs are no longer offered for new chat selections.
Direct model requests Untracked models pass through. Pre-shutdown deprecated models remain directly callable and emit a warning. Explicit custom aliases remain honored. Already-shut-down OpenAI IDs return 410 model_shutdown before upstream I/O instead of an opaque provider failure or silent substitution.
Model fallback Registered Claude/Gemini same-provider family fallback remains available. Fallback now requires provider context and a candidate present in that provider's current registry; global OpenAI GPT-to-different-model chains are removed.
Media catalogs/endpoints Image and video models remain available to their dedicated media registries and handlers. Typed media entries such as openai/gpt-image-2 remain in the unified catalog. Stale media rows imported as chat models are suppressed only from chat projections.
Persistence/configuration No DB migration, destructive cleanup, environment-variable change, or configuration migration. Existing stored rows are not deleted. Stale rows are filtered at selection/catalog projection time and corrected on future managed imports.

Additional scope guards:

  • Lifecycle records and OpenAI model-ID heuristics apply only when the resolved provider is openai; a custom or compatible provider can use similarly named IDs without being reclassified.
  • Explicit multi-endpoint metadata can still make a specialty model chat/Responses-selectable.
  • Replacement IDs are returned as migration guidance only. There is no automatic provider or model switch.
  • gpt-4-1106-preview is deliberately not classified because the official deprecations page currently lists conflicting shutdown dates; this PR does not guess between them.

Behavior matrix

Case Result
OpenAI model not present in the lifecycle table Allowed unchanged
OpenAI model deprecated but before shutdown Direct request allowed with warning; hidden from new default selections/imports
OpenAI model after shutdown Rejected locally with HTTP 410, code model_shutdown, shutdown date, and replacement guidance when available
Non-OpenAI provider with the same model ID Unchanged
Known OpenAI gpt-image-*, dall-e-*, chatgpt-image-latest, or sora-* in a chat-only flow Excluded from that chat flow
The same image/video model in its dedicated media catalog Retained
Explicit user-defined alias Applied before lifecycle evaluation; the resolved target is then evaluated

Related Issues

  • No issue is automatically closed by this PR.

Validation

Run only the focused loop for what you changed — the full unit suite, Vitest, the
60% coverage gate, and the production build all run in CI on this PR (#8329):

  • Focused tests for lifecycle, endpoint classification, import, provider catalog, sync, and fallback: 183 passed
  • Full tests/unit/chatcore-translation-paths.test.ts: 63 passed
  • Latest-upstream cooldown regression suite: 4 passed
  • npm run lint
  • npm run typecheck:core
  • npm run check:cycles
  • npm run check:file-size
  • git diff --check
  • Production-code changes include new or updated automated tests in this PR
  • SonarQube PR analysis is green or any remaining issues are explicitly documented below (pending CI)

Focused command:

node --import tsx/esm --test \
  tests/unit/model-lifecycle.test.ts \
  tests/unit/model-lifecycle-integration.test.ts \
  tests/unit/model-deprecation.test.ts \
  tests/unit/model-endpoint-policy.test.ts \
  tests/unit/model-family-fallback-notation.test.ts \
  tests/unit/8134-github-t5-fallback-filter.test.ts \
  tests/unit/t30-kiro-400-model-unavailable.test.ts \
  tests/unit/kiro-claude-sonnet-5-2267.test.ts \
  tests/unit/repro-7268-401-model-not-supported-lockout.test.ts \
  tests/unit/managed-model-import.test.ts \
  tests/unit/provider-models-route.test.ts \
  tests/unit/models-catalog-route.test.ts \
  tests/unit/model-sync-route.test.ts

Latest-upstream regression command:

node --import tsx/esm --test \
  tests/unit/serial/combo-quota-share-cooldown-wait-timing.test.ts

Tests Added Or Updated

  • tests/unit/model-lifecycle.test.ts — provider scope, date transitions, replacement guidance, default catalog filtering, and the conflicted-date omission.
  • tests/unit/model-lifecycle-integration.test.ts — pre-upstream shutdown rejection, stale unified-catalog projection, typed media retention, and chat-only provider-catalog filtering.
  • tests/unit/model-endpoint-policy.test.ts — OpenAI image/video classification, synthetic metadata handling, explicit multi-endpoint override, and non-OpenAI isolation.
  • tests/unit/managed-model-import.test.ts — managed imports exclude retired and media-only OpenAI models from chat selections.
  • tests/unit/chatcore-translation-paths.test.ts — OpenAI failures no longer trigger cross-model substitution for unavailable/context/empty-content cases.
  • tests/unit/model-deprecation.test.ts — retired OpenAI IDs are no longer rewritten through legacy built-in aliases.
  • tests/unit/t30-kiro-400-model-unavailable.test.ts — provider-scoped fallback and Kiro-only malformed-request classification.

Coverage Notes

  • Lifecycle decision branches are covered directly and through handleChatCore.
  • Endpoint policy is covered directly and through import, provider-route, sync-route, and unified-catalog integration tests.
  • Family fallback is covered for provider notation, registry membership, Kiro/GitHub/Anthropic regressions, and removal of OpenAI substitution.
  • The full repository coverage gate and production build are intentionally left to the existing PR CI workflow.

Reviewer Notes

  • Highest-risk surface: model discovery and fallback. The compatibility table above documents the boundaries, and the focused regression suite exercises the existing provider-specific paths around them.
  • No migration or one-time cleanup is needed. Rollback is a code-only revert; stored provider/model data remains intact.
  • The lifecycle table is intentionally conservative and contains only unambiguous records verified from the official source on 2026-07-26.
  • The lifecycle/endpoint predicate shared by synced and custom unified-catalog sources was extracted into catalogModelPolicy.ts; the frozen file-size baseline was not raised.
  • Route-level lifecycle tests were consolidated into the new integration test file instead of growing the two already-frozen route suites.
  • The ESLint suppression count for chatcore-translation-paths.test.ts decreases from 34 to 31 because three any-using assertions were removed; no new warning is suppressed.

@backryun

backryun commented Jul 26, 2026 •

Copy link
Copy Markdown
Contributor Author
image confirmed these model deprecated from openAI, but not completly deleted from OpenAI endpoint server.

This patch is the first step toward hiding models that are clearly dead and improving the overly complex reroute logic.

@backryun
backryun marked this pull request as ready for review July 26, 2026 02:33
@backryun
backryun requested a review from diegosouzapw as a code owner July 26, 2026 02:33
@backryun
backryun force-pushed the refactor/model-lifecycle branch from b0a38a7 to 7fb8f5f Compare July 26, 2026 07:29
@backryun

Copy link
Copy Markdown
Contributor Author

Validation against a live /v1/models snapshot

The lifecycle table in this PR is hardcoded, so it is only as good as the day it was compiled. It has now been checked against an independent source: a direct GET https://api.openai.com/v1/models probe of a configured OpenAI project key, captured 2026-07-26 — the same day the table claims to have been verified. 129 models listed, 62 of them carrying an official shutdown date.

Summary: the snapshot corroborates every record this PR has, and shows the table stops covering announced shutdowns after 2026-08-10.

1. Where they overlap, they agree exactly

19 of the table's 24 records are for models the account still lists. For all 19, both the shutdown date and the replacement model ID match the snapshot exactly — zero mismatches, down to details like gpt-4o-mini-tts-2025-03-20 → gpt-4o-mini-tts-2025-12-15 and gpt-audio-mini-2025-10-06 → gpt-audio-1.5.

More importantly for merge risk: the 16 models this PR rejects with 410 today are exactly the 16 the snapshot marks as past their shutdown date. No model that is still alive is being cut off, and no rejection is contradicted by the live listing.

2. Coverage stops after 2026-08-10

Of the 62 API-listed models with an official shutdown date, the table covers 19. The 43 it omits are not randomly distributed:

Shutdown date In table Missing
2026-07-23 16 0
2026-08-10 2 0
2026-09-24 0 2
2026-09-28 0 5
2026-10-23 1 18
2026-12-01 0 3
2026-12-11 0 6
2027-01-20 0 9

Everything due on or before 2026-08-10 is covered 18/18. Everything from 2026-09-24 onward is covered 1 out of 44.

So in practice this is a near-term window, not the full deprecations page. That is a defensible scope — but the module comment says "Unambiguous shutdowns from the official OpenAI deprecations page, verified 2026-07-26", which reads as complete coverage. Nothing in the code states the cutoff.

The single exception makes the inconsistency concrete: gpt-3.5-turbo-0125 is in the table at 2026-10-23, while 18 models sharing that identical date are not — gpt-3.5-turbo, gpt-3.5-turbo-16k, gpt-4, gpt-4-0613, gpt-4-turbo, gpt-4-turbo-2024-04-09, gpt-4.1-nano, gpt-4.1-nano-2025-04-14, gpt-4o-2024-05-13, gpt-image-1, o1, o1-2024-12-17, o1-pro, o1-pro-2025-03-19, o3-mini, o3-mini-2025-01-31, o4-mini, o4-mini-2025-04-16.

Two consequences follow:

  • When the 2026-10-23 wave lands (~3 months out), exactly one of those 19 models will return a clean 410 model_shutdown and the other 18 will fail upstream opaquely — the failure mode this PR exists to eliminate.
  • 24 of the 43 omitted models are classified General LLM Import = Yes in the snapshot, including gpt-4, gpt-4-turbo, o1, o1-pro, o3-mini, o4-mini, o3-2025-04-16, o3-pro-2025-06-10, gpt-5-2025-08-07, gpt-5-mini-2025-08-07, gpt-5-nano-2025-08-07, gpt-5-pro-2025-10-06. They keep being offered in chat import and pickers with no deprecation warning, while gpt-3.5-turbo-0125 is hidden — inconsistent behaviour derived from the same upstream data.

The underlying issue is that a hardcoded table has no staleness signal. It will drift silently and there is nothing in CI that notices.

3. The cited source does not match the snapshot's attribution

Every record here carries source: OPENAI_MODEL_DEPRECATIONS_URL, and that string reaches the client in the 410 body. In the snapshot, the classification source recorded for all 19 matching models is the data-residency guide (/api/docs/guides/your-data#...), not the deprecations page — across the 62 dated rows, 56 cite data-residency and only 4 cite deprecations.

The dates themselves check out, so this is not a correctness problem. But if the deprecations page does not actually list these IDs, the 410 sends operators to a page where they cannot confirm what they were just told. Worth verifying the citation before merge.

4. Provider scoping verified — the Codex entries are safe

Worth recording, because it looks alarming at first glance. Five shut-down entries are Codex model IDs (gpt-5-codex, gpt-5.1-codex, gpt-5.1-codex-max, gpt-5.1-codex-mini, gpt-5.2-codex) and those IDs are referenced in this repo — open-sse/config/providers/registry/opencode/zen/index.ts, open-sse/executors/codex.ts, backgroundTaskDetector.ts, quotaAutoPing.ts and the CLI-code tool options.

None of them are affected. getModelLifecycleDecision keys on provider\0model and every record is provider: "openai", so codex/gpt-5.1-codex and opencode-zen/gpt-5.1-codex resolve to untracked → allow. The provider scoping the PR description promises is genuinely implemented, not just asserted.

5. Five records the snapshot cannot confirm

computer-use-preview, computer-use-preview-2025-03-11, gpt-4-0314, gpt-4-0125-preview and gpt-4-turbo-preview are all past their dates (so rejected with 410) and are absent from the account listing entirely. They cannot be cross-checked here — but absence from /v1/models is itself consistent with them being gone, so rejecting them costs nothing.

6. Limits of this validation

To be fair to the PR, the snapshot is not an oracle either:

  • It is one account's non-paginated /v1/models response, not a global historical catalog.
  • API listing does not imply inference availability — a model can be listed and still fail.
  • The source-attribution oddity in §3 applies to the snapshot as much as to the PR.

So this is not evidence that any record here is wrong. It is evidence that the table's coverage is narrower than its own description, and it independently confirms the 19 records that could be checked.

Recommendation

Nothing found here blocks the merge. Two follow-ups worth folding in:

  1. State the scope in the module comment — "already shut down, plus announced shutdowns through 2026-08-10" turns the current state from an omission into a decision. Then either drop gpt-3.5-turbo-0125 (which sits outside that window) or add the other 18 models sharing its date.
  2. Add a staleness gate — a test that fails when the newest record is older than some threshold, so the table cannot quietly rot between releases. Without one, the 2026-10-23 wave arrives with 18 uncovered models and nothing to flag it.

@backryun
backryun force-pushed the refactor/model-lifecycle branch from 7fb8f5f to ff98205 Compare July 26, 2026 20:51
@diegosouzapw diegosouzapw added the needs-vps PR requires Hard Rule #18 VPS smoke test on 192.168.0.15 before merge label Jul 26, 2026
@diegosouzapw

Copy link
Copy Markdown
Owner

Hi @backryun — high-quality refactor. The shared catalogModelPolicy.ts predicate is the right design (one source of truth across synced + custom catalogs), and the cross-model substitution removal is the right call — silent reroutes mask real provider problems.

Marked needs-vps (★4). I need a Hard Rule #18 VPS smoke on 192.168.0.15 covering:

  1. Normal OpenAI chat request, 200.
  2. OpenAI request with max_tokens above the model's output cap (should clamp, not 400).
  3. Image-only OpenAI model request (should be rejected as not-chat-eligible, not silent rerouted).
  4. Deprecated OpenAI model request (should be rejected as not-available, not silent rerouted).

Paste the exact curl commands + responses in the PR description. If the smoke finds a regression, we can patch in-place. If it passes, /merge-prs will merge.

Watching for the interaction with #8698 (max_tokens clamp) — merge order matters; recommend this lands first.

@backryun

Copy link
Copy Markdown
Contributor Author

Thanks — two things before the smoke, both of which should shrink it.

1. #8698 already landed, and this branch is already on top of it

It merged at 2026-07-26T19:30Z, about two hours before the note above. I then rebased this branch onto the current tip, so 6389c5b12 is an ancestor of the head. The recommended order is therefore already settled — in the other direction, which is arguably the better one: smoke #2 now exercises the combined behaviour rather than this PR alone. Flagging it only so the merge-order item is not still treated as open.

That rebase also resolved the single conflict it produced — an import collision in src/app/api/v1/models/catalog.ts against #8666's cc-discovery aliases — and CI is green; the PR is CLEAN.

One thing worth recording from that resolution, since it looked like a real interaction: #8666 appends claude/<id> mirror aliases and its own comment says they are "deliberately NOT filtered by model type". They are still safe here. appendCcDiscoveryAliases runs at catalog.ts:1458 against finalModels — i.e. after the isUnifiedChatSourceModelSelectable filters at 645 / 763 / 1143 — so a shut-down or media model never gets a mirror.

2. Two of the four scenarios need no VPS — they are already covered in this PR

Scenario 4 — deprecated model rejected, not silently rerouted.
tests/unit/model-lifecycle-integration.test.ts:53, "chatCore rejects a shutdown OpenAI model before an upstream request". It stubs globalThis.fetch and asserts:

assert.equal(result.status, 410);
assert.equal(result.errorCode, "model_shutdown");
assert.equal(upstreamCalls, 0);                                   // ← the no-reroute proof
assert.match(responseBody.error.message, /openai\/gpt-5\.2-codex/);
assert.match(responseBody.error.message, /openai\/gpt-5\.6-sol/);  // replacement is advisory text

upstreamCalls === 0 is precisely the property the smoke would be looking for — and it is a stronger check than a live call can give, because a live 410 cannot distinguish "we rejected locally" from "upstream rejected it".

Scenario 3 — image-only model not chat-eligible.
tests/unit/model-endpoint-policy.test.ts:7 covers gpt-image-2, gpt-image-1.5, gpt-image-1-mini, dall-e-3, chatgpt-image-latest; :23 covers sora-2-pro; :31 asserts the policy is provider-scoped (a custom provider with a gpt-image-shaped id is not reclassified); :38 and :48 pin the explicit-endpoint-metadata override in both directions. model-lifecycle-integration.test.ts:122 covers the chat-only catalog projection.

Neither path issues a network request, so a VPS run would exercise the same code with more setup and weaker assertions. Suggest narrowing the smoke to #1 and #2, which genuinely need a live key.

3. Ready-to-run commands for #1 and #2

I cannot run these — 192.168.0.15 is on your network, not reachable from mine (port 22 times out), so this needs your hands.

BASE=http://127.0.0.1:20128
KEY=<omniroute api key>

#1 — normal OpenAI chat request → expect 200.

curl -sS -i -X POST "$BASE/v1/chat/completions" \
  -H "Authorization: Bearer $KEY" -H "Content-Type: application/json" \
  -d '{"model":"openai/gpt-4o-mini","messages":[{"role":"user","content":"reply with the single word: ok"}],"max_tokens":16}'

Pass: 200 with a normal choices[0].message.content. Fail: any 4xx/5xx, or a response whose model is not the one requested (that would be a surviving silent reroute).

#2 — max_tokens above the model's output cap → expect a clamp, not 400.

The cap is resolved per install from synced capabilities / registry, so rather than guess a number near it, this uses one that exceeds any cap:

curl -sS -i -X POST "$BASE/v1/chat/completions" \
  -H "Authorization: Bearer $KEY" -H "Content-Type: application/json" \
  -d '{"model":"openai/gpt-4o-mini","messages":[{"role":"user","content":"count from 1 to 5"}],"max_tokens":999999}'

Pass: 200, with usage.completion_tokens at or below the model cap. Fail: 400 from upstream (that is the pre-#8698 behaviour leaking through), or a 5xx.

If it helps, I am happy to add a third local test asserting the clamp wiring end-to-end through handleChatCore with a stubbed upstream, so only #1 — a plain liveness check — would remain VPS-only. Say the word and I will push it.

@backryun
backryun force-pushed the refactor/model-lifecycle branch 3 times, most recently from afc173d to 9776245 Compare July 27, 2026 22:27
@diegosouzapw diegosouzapw added the deferred-v3.8.50 Adiada para o ciclo v3.8.50 (validacao VPS, refactor, ou escopo grande) label Jul 27, 2026
@diegosouzapw diegosouzapw changed the title fix: enforce OpenAI model lifecycle without silent reroutes [v3.8.50] fix: enforce OpenAI model lifecycle without silent reroutes Jul 27, 2026
@backryun
backryun force-pushed the refactor/model-lifecycle branch from 9776245 to f930d37 Compare July 28, 2026 02:11
@backryun

Copy link
Copy Markdown
Contributor Author

Root refresh (2026-07-28): rebased onto release/v3.8.49 at d6c0693 and force-pushed safely with lease. New head: f930d37. The rebase conflicts in modelFamilyFallback.ts and the T30 Kiro test were resolved by preserving both protections: Files API resource 404s do not trigger model fallback, while fallback and Kiro malformed-request classification remain provider-scoped. The new root also exhausted the frozen line budgets in chatCore.ts and the provider-models route, so lifecycle enforcement and no-auth model projection were extracted into focused helper modules without changing behavior. Validation: 250/250 focused regressions, core typecheck, dashboard frozen baseline, cycles, file-size, complexity ratchets, changed-file ESLint, and diff check all pass. The maintainer-requested VPS smoke remains the only external validation item.

@diegosouzapw

Copy link
Copy Markdown
Owner

Heads-up: this PR is running against a stale base

release/v3.8.49 has moved 4 commits ahead of the base this PR was last built against, and several gate failures that were red on the older base have since been repaired on the branch itself:

Gate class Repaired by
No new ESLint warnings (exit 2 — orphaned suppressions, not a lint violation) #8706, #8831
Fast Quality Gates / check:file-size / Quality Ratchet #8585, #8767, plus the post-merge-train re-pins
Docs Gates (fast-path) (missing Polish API_REFERENCE) #8831
backoff-clamp + compressionDetailNormalizers unit tests #8706

I verified this on a clean checkout of the current tip (3b515d90b3): lint:json --max-warnings 0, check:file-size, check:docs-sync, error-classification.test.ts (24/24) and check-db-rules.test.ts (22/22) all pass with no changes.

Currently failing here:

  • Build (advisory)
  • Docs Gates (fast-path)
  • No new ESLint warnings

Please update the branch (the Update branch button, or rebase onto origin/release/v3.8.49) and let CI re-run. That should clear any failure caused by the stale base.

Note this is not a promise that every red disappears — a failure specific to the changes in this PR will survive the rebase, and Build (advisory) is advisory and does not block. Rebasing just removes the noise so what is left is actually yours.

No action was taken on this PR beyond this comment.

@backryun
backryun force-pushed the refactor/model-lifecycle branch from ec5c098 to bddefc6 Compare July 28, 2026 07:15
@diegosouzapw
diegosouzapw changed the base branch from release/v3.8.49 to release/v3.8.50 July 28, 2026 18:42
@diegosouzapw

Copy link
Copy Markdown
Owner

Re-homed to release/v3.8.50: v3.8.49 entered its release freeze, so the branch now belongs to the release captain and development continues on the next cycle. Nothing is wrong with this PR — it just needed a live base. No action needed from you; CI will re-run against the new base.

@backryun
backryun force-pushed the refactor/model-lifecycle branch from bddefc6 to a161109 Compare July 28, 2026 21:07
@backryun
backryun force-pushed the refactor/model-lifecycle branch 5 times, most recently from 1842a03 to f0099c6 Compare August 1, 2026 08:00
@backryun
backryun force-pushed the refactor/model-lifecycle branch 21 times, most recently from bcc515a to 67277ac Compare August 9, 2026 23:57
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

deferred-v3.8.50 Adiada para o ciclo v3.8.50 (validacao VPS, refactor, ou escopo grande) needs-vps PR requires Hard Rule #18 VPS smoke test on 192.168.0.15 before merge

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants