Repository navigation
[v3.8.50] fix: enforce OpenAI model lifecycle without silent reroutes - #8627
Conversation
b0a38a7 to
7fb8f5f
Compare
Validation against a live
|
| Shutdown date | In table | Missing |
|---|---|---|
| 2026-07-23 | 16 | 0 |
| 2026-08-10 | 2 | 0 |
| 2026-09-24 | 0 | 2 |
| 2026-09-28 | 0 | 5 |
| 2026-10-23 | 1 | 18 |
| 2026-12-01 | 0 | 3 |
| 2026-12-11 | 0 | 6 |
| 2027-01-20 | 0 | 9 |
Everything due on or before 2026-08-10 is covered 18/18. Everything from 2026-09-24 onward is covered 1 out of 44.
So in practice this is a near-term window, not the full deprecations page. That is a defensible scope — but the module comment says "Unambiguous shutdowns from the official OpenAI deprecations page, verified 2026-07-26", which reads as complete coverage. Nothing in the code states the cutoff.
The single exception makes the inconsistency concrete: gpt-3.5-turbo-0125 is in the table at 2026-10-23, while 18 models sharing that identical date are not — gpt-3.5-turbo, gpt-3.5-turbo-16k, gpt-4, gpt-4-0613, gpt-4-turbo, gpt-4-turbo-2024-04-09, gpt-4.1-nano, gpt-4.1-nano-2025-04-14, gpt-4o-2024-05-13, gpt-image-1, o1, o1-2024-12-17, o1-pro, o1-pro-2025-03-19, o3-mini, o3-mini-2025-01-31, o4-mini, o4-mini-2025-04-16.
Two consequences follow:
- When the 2026-10-23 wave lands (~3 months out), exactly one of those 19 models will return a clean
410 model_shutdownand the other 18 will fail upstream opaquely — the failure mode this PR exists to eliminate. - 24 of the 43 omitted models are classified
General LLM Import = Yesin the snapshot, includinggpt-4,gpt-4-turbo,o1,o1-pro,o3-mini,o4-mini,o3-2025-04-16,o3-pro-2025-06-10,gpt-5-2025-08-07,gpt-5-mini-2025-08-07,gpt-5-nano-2025-08-07,gpt-5-pro-2025-10-06. They keep being offered in chat import and pickers with no deprecation warning, whilegpt-3.5-turbo-0125is hidden — inconsistent behaviour derived from the same upstream data.
The underlying issue is that a hardcoded table has no staleness signal. It will drift silently and there is nothing in CI that notices.
3. The cited source does not match the snapshot's attribution
Every record here carries source: OPENAI_MODEL_DEPRECATIONS_URL, and that string reaches the client in the 410 body. In the snapshot, the classification source recorded for all 19 matching models is the data-residency guide (/api/docs/guides/your-data#...), not the deprecations page — across the 62 dated rows, 56 cite data-residency and only 4 cite deprecations.
The dates themselves check out, so this is not a correctness problem. But if the deprecations page does not actually list these IDs, the 410 sends operators to a page where they cannot confirm what they were just told. Worth verifying the citation before merge.
4. Provider scoping verified — the Codex entries are safe
Worth recording, because it looks alarming at first glance. Five shut-down entries are Codex model IDs (gpt-5-codex, gpt-5.1-codex, gpt-5.1-codex-max, gpt-5.1-codex-mini, gpt-5.2-codex) and those IDs are referenced in this repo — open-sse/config/providers/registry/opencode/zen/index.ts, open-sse/executors/codex.ts, backgroundTaskDetector.ts, quotaAutoPing.ts and the CLI-code tool options.
None of them are affected. getModelLifecycleDecision keys on provider\0model and every record is provider: "openai", so codex/gpt-5.1-codex and opencode-zen/gpt-5.1-codex resolve to untracked → allow. The provider scoping the PR description promises is genuinely implemented, not just asserted.
5. Five records the snapshot cannot confirm
computer-use-preview, computer-use-preview-2025-03-11, gpt-4-0314, gpt-4-0125-preview and gpt-4-turbo-preview are all past their dates (so rejected with 410) and are absent from the account listing entirely. They cannot be cross-checked here — but absence from /v1/models is itself consistent with them being gone, so rejecting them costs nothing.
6. Limits of this validation
To be fair to the PR, the snapshot is not an oracle either:
- It is one account's non-paginated
/v1/modelsresponse, not a global historical catalog. - API listing does not imply inference availability — a model can be listed and still fail.
- The source-attribution oddity in §3 applies to the snapshot as much as to the PR.
So this is not evidence that any record here is wrong. It is evidence that the table's coverage is narrower than its own description, and it independently confirms the 19 records that could be checked.
Recommendation
Nothing found here blocks the merge. Two follow-ups worth folding in:
- State the scope in the module comment — "already shut down, plus announced shutdowns through 2026-08-10" turns the current state from an omission into a decision. Then either drop
gpt-3.5-turbo-0125(which sits outside that window) or add the other 18 models sharing its date. - Add a staleness gate — a test that fails when the newest record is older than some threshold, so the table cannot quietly rot between releases. Without one, the 2026-10-23 wave arrives with 18 uncovered models and nothing to flag it.
7fb8f5f to
ff98205
Compare
|
Hi @backryun — high-quality refactor. The shared Marked
Paste the exact curl commands + responses in the PR description. If the smoke finds a regression, we can patch in-place. If it passes, Watching for the interaction with #8698 (max_tokens clamp) — merge order matters; recommend this lands first. |
|
Thanks — two things before the smoke, both of which should shrink it. 1. #8698 already landed, and this branch is already on top of itIt merged at That rebase also resolved the single conflict it produced — an import collision in One thing worth recording from that resolution, since it looked like a real interaction: #8666 appends 2. Two of the four scenarios need no VPS — they are already covered in this PRScenario 4 — deprecated model rejected, not silently rerouted. assert.equal(result.status, 410);
assert.equal(result.errorCode, "model_shutdown");
assert.equal(upstreamCalls, 0); // ← the no-reroute proof
assert.match(responseBody.error.message, /openai\/gpt-5\.2-codex/);
assert.match(responseBody.error.message, /openai\/gpt-5\.6-sol/); // replacement is advisory text
Scenario 3 — image-only model not chat-eligible. Neither path issues a network request, so a VPS run would exercise the same code with more setup and weaker assertions. Suggest narrowing the smoke to #1 and #2, which genuinely need a live key. 3. Ready-to-run commands for #1 and #2I cannot run these — BASE=http://127.0.0.1:20128
KEY=<omniroute api key>#1 — normal OpenAI chat request → expect curl -sS -i -X POST "$BASE/v1/chat/completions" \
-H "Authorization: Bearer $KEY" -H "Content-Type: application/json" \
-d '{"model":"openai/gpt-4o-mini","messages":[{"role":"user","content":"reply with the single word: ok"}],"max_tokens":16}'Pass: #2 — The cap is resolved per install from synced capabilities / registry, so rather than guess a number near it, this uses one that exceeds any cap: curl -sS -i -X POST "$BASE/v1/chat/completions" \
-H "Authorization: Bearer $KEY" -H "Content-Type: application/json" \
-d '{"model":"openai/gpt-4o-mini","messages":[{"role":"user","content":"count from 1 to 5"}],"max_tokens":999999}'Pass: If it helps, I am happy to add a third local test asserting the clamp wiring end-to-end through |
afc173d to
9776245
Compare
9776245 to
f930d37
Compare
|
Root refresh (2026-07-28): rebased onto release/v3.8.49 at d6c0693 and force-pushed safely with lease. New head: f930d37. The rebase conflicts in modelFamilyFallback.ts and the T30 Kiro test were resolved by preserving both protections: Files API resource 404s do not trigger model fallback, while fallback and Kiro malformed-request classification remain provider-scoped. The new root also exhausted the frozen line budgets in chatCore.ts and the provider-models route, so lifecycle enforcement and no-auth model projection were extracted into focused helper modules without changing behavior. Validation: 250/250 focused regressions, core typecheck, dashboard frozen baseline, cycles, file-size, complexity ratchets, changed-file ESLint, and diff check all pass. The maintainer-requested VPS smoke remains the only external validation item. |
f930d37 to
5dca287
Compare
5dca287 to
ec5c098
Compare
Heads-up: this PR is running against a stale base
I verified this on a clean checkout of the current tip ( Currently failing here:
Please update the branch (the Update branch button, or rebase onto Note this is not a promise that every red disappears — a failure specific to the changes in this PR will survive the rebase, and No action was taken on this PR beyond this comment. |
ec5c098 to
bddefc6
Compare
|
Re-homed to |
bddefc6 to
a161109
Compare
1842a03 to
f0099c6
Compare
bcc515a to
67277ac
Compare

Summary
410/model_shutdownerror and replacement guidance./modelsmetadata is absent or historically persisted as syntheticchatmetadata.Why
OmniRoute currently combines upstream discovery, local catalogs, aliases, and family fallback. Two failure modes can follow:
/modelsresponse does not reliably describe which endpoint owns each model. Image and video IDs can therefore be imported into chat selectors merely because OmniRoute historically defaulted missing endpoint metadata tochat.This PR centralizes those decisions while keeping them provider-scoped and projection-scoped.
Compatibility / breaking-change assessment
This is an intentional correction of invalid routing state, not a public API or storage migration.
GET /api/providers/:id/modelsremains broad by default. No endpoint or response field is removed.chatOnly=trueapplies lifecycle and chat-endpoint filtering for chat consumers.410 model_shutdownbefore upstream I/O instead of an opaque provider failure or silent substitution.openai/gpt-image-2remain in the unified catalog.Additional scope guards:
openai; a custom or compatible provider can use similarly named IDs without being reclassified.gpt-4-1106-previewis deliberately not classified because the official deprecations page currently lists conflicting shutdown dates; this PR does not guess between them.Behavior matrix
410, codemodel_shutdown, shutdown date, and replacement guidance when availablegpt-image-*,dall-e-*,chatgpt-image-latest, orsora-*in a chat-only flowRelated Issues
Validation
Run only the focused loop for what you changed — the full unit suite, Vitest, the
60% coverage gate, and the production build all run in CI on this PR (#8329):
tests/unit/chatcore-translation-paths.test.ts: 63 passednpm run lintnpm run typecheck:corenpm run check:cyclesnpm run check:file-sizegit diff --checkFocused command:
Latest-upstream regression command:
Tests Added Or Updated
tests/unit/model-lifecycle.test.ts— provider scope, date transitions, replacement guidance, default catalog filtering, and the conflicted-date omission.tests/unit/model-lifecycle-integration.test.ts— pre-upstream shutdown rejection, stale unified-catalog projection, typed media retention, and chat-only provider-catalog filtering.tests/unit/model-endpoint-policy.test.ts— OpenAI image/video classification, synthetic metadata handling, explicit multi-endpoint override, and non-OpenAI isolation.tests/unit/managed-model-import.test.ts— managed imports exclude retired and media-only OpenAI models from chat selections.tests/unit/chatcore-translation-paths.test.ts— OpenAI failures no longer trigger cross-model substitution for unavailable/context/empty-content cases.tests/unit/model-deprecation.test.ts— retired OpenAI IDs are no longer rewritten through legacy built-in aliases.tests/unit/t30-kiro-400-model-unavailable.test.ts— provider-scoped fallback and Kiro-only malformed-request classification.Coverage Notes
handleChatCore.Reviewer Notes
catalogModelPolicy.ts; the frozen file-size baseline was not raised.chatcore-translation-paths.test.tsdecreases from 34 to 31 because threeany-using assertions were removed; no new warning is suppressed.