Skip to content

fix(autoCombo,sse): catalog hygiene — retire dead FITNESS_TABLE rows, fix 7 BUILT_IN_ALIASES targets, add model-lifecycle gate (#11503) - #11507

Merged
diegosouzapw merged 4 commits into
diegosouzapw:release/v3.8.51from
MumuTW:fix/catalog-hygiene-11503
Aug 25, 2026
Merged

diegosouzapw merged 4 commits into
diegosouzapw:release/v3.8.51from
MumuTW:fix/catalog-hygiene-11503

Conversation

@MumuTW

@MumuTW MumuTW commented Aug 25, 2026

Copy link
Copy Markdown
Contributor

Fixes #11503

Two hand-maintained tables decide routing and both had rotted against the vendors' own deprecation pages. This PR cleans them and adds an offline gate so the next drift is a red check instead of a silent ranking inversion.


Commit 1 — FITNESS_TABLE hygiene + segment-boundary matching

Problem. Layer 4 of the task-fitness chain was a table of version-less family patterns (codex, claude-sonnet, qwen, llama, gemini-pro, …) matched with String.includes. Recognition by the table, not availability, decided rank: openai/gpt-5.2-codex (OpenAI shut the codex family down 2026-07-23) scored 0.98 for coding via the codex row, while live flagships the table had never heard of (claude-fable-5-thinking-max, gpt-5.6-sol-xhigh) sat at the wildcard 0.50. Substring matching leaked further: o3 matched solar-pro3, and gpt-4o matched chatgpt-4o-latest — a different, retired model that inherited the live flagship's 0.9.

Fix.

  • Dropped every row that matches nothing in the catalog or names a retired family (o1, mixtral, grok-4-fast, codex) and every version-less family row (claude-sonnet, claude-opus, claude-haiku, qwen, llama, mistral, gemini-pro, gemini-flash, grok-4, kimi-k2, glm-5). Versioned rows naming a live model stay (gpt-4o, gpt-4o-mini, gpt-4-turbo, o3, o4-mini, gemini-2.5-*, gemini-3.1-pro, deepseek-*, grok-3, glm-5.1, minimax-*).
  • getStaticFitnessTableScore now matches on segment boundaries (- / . / / / start / end), longest pattern first as before. o3 still matches o3, o3-mini, openai/o3 but no longer solar-pro3; gpt-4o still matches gpt-4o-2024-05-13 but no longer chatgpt-4o-latest — that one now falls to 0.5, which is the intended outcome since OpenAI shut it down 2026-02-17.
  • The file header now documents that layer 4 is a small versioned table and that the wildcard 0.5 an unknown id falls to means "no evidence", never a quality claim. scoring.ts is untouched — the fix is removal, not reweighting.

Tests. New open-sse/services/autoCombo/__tests__/fitness-table-hygiene-11503.test.ts (vitest): loads the lifecycle snapshot and asserts no routable retired id scores through the table at all; solar-pro3, chatgpt-4o-latest, claude-sonnet-5, claude-fable-5-thinking-max, gpt-5.6-sol-xhigh → null; gpt-4o-mini still 0.8, o3-mini still 0.95 via o3, openai/o3 still resolves.

Before / after — every catalog id resolved through getTaskFitnessWithSource(id, "coding")

before (a179ffed5b) after
resolved via fitness_table 533 / 1327 73 / 1327
resolved via wildcard_boost 794 1254
retired ids scoring ≥ 0.88 via fitness_table 2 (openai/gpt-5.2-codex 0.98, chatgpt-4o-latest 0.90) 0

⚠️ The issue's audit quoted larger counts (12 ids ≥ 0.88). Those figures folded in ids the vendors mark retiring (shutdown scheduled) alongside retired, and counted bare ids the catalog does not actually route. Counting only ids that are both status: "retired" and present in REGISTRY, the reproducible numbers on this base are 8 routable retired ids, 2 of them ≥ 0.88. Both are now 0.

Commit 2 — BUILT_IN_ALIASES targets + provider guard

Problem. The alias table rewrites body.model on every request (modelLifecyclePolicy.ts), so a stale target is a guaranteed 404. Seven rows pointed at a retired or non-existent model — five Claude rows chained one retired id to another.

source old target status new target
claude-3-opus-20240229 claude-opus-4-20250514 retired 2026-06-15, absent claude-opus-4-8
claude-3-sonnet-20240229 claude-sonnet-4-20250514 retired 2026-06-15, absent claude-sonnet-4-6
claude-3-5-sonnet-latest claude-sonnet-4-20250514 retired 2026-06-15, absent claude-sonnet-4-6
claude-3-haiku-20240307 claude-3-5-sonnet-20241022 retired 2025-10-28 claude-haiku-4-5-20251001
claude-3-5-haiku-latest claude-3-5-sonnet-20241022 retired 2025-10-28 claude-haiku-4-5-20251001
gemini-3-pro-high gemini-3.1-pro-high not a catalog id gemini-3-1-pro-high (the catalog's spelling)
llama-3-8b llama3-8b-8192 Groq deprecated 2025-08-30, absent llama-3.1-8b-instant

Every replacement is the one the vendor publishes; the other 23 rows are untouched.

Provider guard. The table is global but the catalog is not — aggregators still serve kimi-k2, gemini-2.0-flash, mistral-large under their original ids, and the unconditional rewrite broke them. resolveModelAlias(modelId, provider?) now returns modelId unchanged when the given provider serves it as-is (hasKnownProviderModel, exported from open-sse/services/model.ts for this). Both modelLifecyclePolicy.ts call sites pass the provider they already have. getDeprecationNotice / isDeprecated keep their 1-arg behaviour, and a custom (operator-authored) alias still wins over a built-in one, provider or not.

Tests. New tests/unit/model-deprecation-aliases-11503.test.ts (node:test, 66 assertions): table-driven over every BUILT_IN_ALIASES entry — each target must be a catalog id (checked against REGISTRY, not a hardcoded list) and must not be retired in the snapshot; plus resolveModelAlias("kimi-k2", "t3-web") === "kimi-k2", resolveModelAlias("kimi-k2", "fireworks") === "moonshotai/Kimi-K2", resolveModelAlias("claude-3-opus-20240229") === "claude-opus-4-8", and custom-alias precedence.

Commit 3 — check:model-lifecycle gate + vendor snapshot

config/quality/model-lifecycle.json — 106 lifecycle entries (68 retired) curated from the five first-party deprecation pages listed in its sources. replacement is null wherever the vendor publishes none.

scripts/check/check-model-lifecycle.mjs — no network, three checks, all summed before exit:

  1. no FITNESS_TABLE pattern scores a routable retired id, under the new boundary rule;
  2. no BUILT_IN_ALIASES target is retired or absent from the catalog;
  3. every retired id the catalog still routes is either forwarded by BUILT_IN_ALIASES or listed in allowedRetiredInCatalog.

Check (1) is deliberately scoped to routable ids: legitimate versioned rows like gpt-4o also match retired ids the catalog never served (gpt-4o-audio-preview), and those cannot invert any routing decision. The gate sets DATA_DIR to a throwaway temp dir before importing production modules so it can never migrate the operator's database.

scripts/quality/refresh-model-lifecycle.mjs — network refresher, not wired into CI. It parses Anthropic's markdown table (| API model name | Current state | Deprecated | Tentative retirement date |); for the other vendors it keeps the existing entries and prints manual review needed rather than guessing at a shifting HTML layout. It aborts rather than write a snapshot with fewer entries than it started with — it must never silently drop one.

Wired into .github/workflows/quality.yml (gates=(…)) and .github/workflows/ci.yml's lint job — the latter is beyond the letter of the issue but is where the other pass/fail policy gates live, so a PR to main gets the same protection. Documented in docs/architecture/QUALITY_GATES.md. New scripts: check:model-lifecycle, quality:refresh-model-lifecycle.

Tests. tests/unit/check-model-lifecycle-gate.test.ts exercises each of the three checks plus the catalog/retired helpers against small fixtures.

allowedRetiredInCatalog — the burn-down list

Check (3) is red on the current catalog, which is the point. To keep this PR mergeable the 6 offenders are allowlisted with a TODO(#11503) note in the file; the maintainer's call is remove-from-catalog vs. add-a-forward for each:

id vendor status
chatgpt-4o-latest OpenAI, retired
claude-3-5-sonnet-20241022 Anthropic, retired 2025-10-28
claude-3-7-sonnet-20250219 Anthropic, retired 2026-02-19
google/gemini-2.0-flash Google, retired 2026-06-01 (the bare gemini-2.0-flash is forwarded; the prefixed form is not, since the alias table is exact-id keyed)
gpt-4-0125-preview OpenAI, retired
openai/gpt-5.2-codex OpenAI, retired

Adjacent but deliberately not changed here: llama-3.1-8b-instant (the new llama-3-8b target) is marked retiring by Groq with a 2026-08-16 date. It is the replacement Groq itself names and is live in the catalog, so it is the correct target today; the gate will flag it the moment the page flips it to retired.


Deviations from the issue's proposed scope

  • The lifecycle snapshot lands in commit 1, not commit 3, because commit 1's test consumes it — this keeps every commit green in isolation.
  • Existing suites that used the removed claude-sonnet / claude-opus rows as their "known static model" probe were retargeted to a surviving versioned row (o3 carries exactly the same coding 0.95 / review 0.92 numbers; gemini-2.5-pro covers one review case). The behaviour under test — layer-4 resolution, longest-pattern-first, default-table fallback, case-insensitivity — is unchanged, and the fix(backend): static fitness table substring match returns wrong row for longer model ids (gpt-4o-mini scores as gpt-4o) #8603 regression assertions are intact.
  • tests/unit/model-deprecation.test.ts had pinned the two stale Claude targets; updated to the corrected ones.

Verification

command result
npx vitest run --config vitest.mcp.config.ts open-sse/services/autoCombo/__tests__/ Test Files 7 passed (7) · Tests 106 passed (106)
node --test tests/unit/{task-fitness-missing-table-8603,model-deprecation-aliases-11503,check-model-lifecycle-gate}.test.ts # tests 77 · # pass 77 · # fail 0
node --test tests/unit/{model-deprecation,model-alias-seed,model-combo-mappings,model-lifecycle,model-lifecycle-integration,8676-monsterapi-deprecation,gemini-cli-deprecation}.test.ts 119 pass · 0 fail
npm run check:model-lifecycle PASS — snapshot 2026-08-25, 68 retired id(s), 1327 catalog id(s) (without the allowlist: 6 violation(s), listed above)
npm run check:cycles OK - no cycles detected across 429 files
npm run typecheck:core clean
npx eslint <touched files> --suppressions-location config/quality/eslint-suppressions.json 0 errors
npm run check:known-symbols / check:docs-sync / check:fabricated-docs / check:test-discovery all PASS
npm run check:docs-counts ✗ 3 stale "159 migrations" refs (README/AGENTS/llm.txt) — pre-existing on the base, unrelated

⚠️ base-red inherited: #11449 (the base tip is red for unrelated reasons; #11502 clears most).

@MumuTW
MumuTW requested a review from diegosouzapw as a code owner August 25, 2026 12:03
@diegosouzapw
diegosouzapw merged commit ae7a843 into diegosouzapw:release/v3.8.51 Aug 25, 2026
7 of 16 checks passed
@MumuTW
MumuTW deleted the fix/catalog-hygiene-11503 branch September 5, 2026 10:19
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
… fix 7 BUILT_IN_ALIASES targets, add model-lifecycle gate (diegosouzapw#11503) (diegosouzapw#11507)

Validated in a combined 4-PR batch worktree off release/v3.8.51 tip. This PR's diff overlapped taskFitness.ts and autoCombo.test.ts with the already-merged diegosouzapw#11492/diegosouzapw#11506 — git's merge auto-resolved both hunks cleanly (non-overlapping layers: diegosouzapw#11492/diegosouzapw#11506 touch layer 2 arena lookup, this PR touches layer 4 static-table hygiene); verified no conflict markers remained and re-ran the full suite after boarding.
- npm run check:model-lifecycle — PASS, 68 retired ids, 1327 catalog ids, 0 violations (re-ran with the correct `node --import tsx/esm` loader after an initial bare-node invocation mistakenly failed on path-alias resolution — that was my invocation error, not the gate)
- Focused tests: fitness-table-hygiene-11503.test.ts, taskFitness-pattern-order-8603.test.ts, model-deprecation-aliases-11503.test.ts, check-model-lifecycle-gate.test.ts, model-deprecation.test.ts, autoCombo.test.ts — part of batch's 126/126 vitest + 246/246 node:test runs
- typecheck:core, file-size, changelog-integrity, complexity, cognitive-complexity, check:cycles — all OK
- Full-repo lint: 228 problems remaining, all pre-existing dashboard react-hooks/* findings unrelated to this diff (zero errors in any file this PR touches)

Thanks for this — genuinely thorough methodology (segment-boundary matching, provider-scoped alias guard, offline lifecycle gate with a documented burn-down list for the 6 remaining catalog offenders).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(backend): catalog hygiene — FITNESS_TABLE ranks retired models at 0.98 and BUILT_IN_ALIASES forwards retired Claude ids to other retired ids

2 participants