Repository navigation
fix(autoCombo,sse): catalog hygiene — retire dead FITNESS_TABLE rows, fix 7 BUILT_IN_ALIASES targets, add model-lifecycle gate (#11503) - #11507
Merged
diegosouzapw merged 4 commits intoAug 25, 2026
Conversation
…; match patterns on segment boundaries (diegosouzapw#11503)
…bsent models; skip the rewrite when the target provider serves the id (diegosouzapw#11503)
…pshot and refresh script (diegosouzapw#11503)
diegosouzapw
merged commit Aug 25, 2026
ae7a843
into
diegosouzapw:release/v3.8.51
7 of 16 checks passed
4 of 5 tasks
muhamadgalihsaputra
pushed a commit
to niyatna/NiyatnaRoute
that referenced
this pull request
Sep 27, 2026
… fix 7 BUILT_IN_ALIASES targets, add model-lifecycle gate (diegosouzapw#11503) (diegosouzapw#11507) Validated in a combined 4-PR batch worktree off release/v3.8.51 tip. This PR's diff overlapped taskFitness.ts and autoCombo.test.ts with the already-merged diegosouzapw#11492/diegosouzapw#11506 — git's merge auto-resolved both hunks cleanly (non-overlapping layers: diegosouzapw#11492/diegosouzapw#11506 touch layer 2 arena lookup, this PR touches layer 4 static-table hygiene); verified no conflict markers remained and re-ran the full suite after boarding. - npm run check:model-lifecycle — PASS, 68 retired ids, 1327 catalog ids, 0 violations (re-ran with the correct `node --import tsx/esm` loader after an initial bare-node invocation mistakenly failed on path-alias resolution — that was my invocation error, not the gate) - Focused tests: fitness-table-hygiene-11503.test.ts, taskFitness-pattern-order-8603.test.ts, model-deprecation-aliases-11503.test.ts, check-model-lifecycle-gate.test.ts, model-deprecation.test.ts, autoCombo.test.ts — part of batch's 126/126 vitest + 246/246 node:test runs - typecheck:core, file-size, changelog-integrity, complexity, cognitive-complexity, check:cycles — all OK - Full-repo lint: 228 problems remaining, all pre-existing dashboard react-hooks/* findings unrelated to this diff (zero errors in any file this PR touches) Thanks for this — genuinely thorough methodology (segment-boundary matching, provider-scoped alias guard, offline lifecycle gate with a documented burn-down list for the 6 remaining catalog offenders).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #11503
Two hand-maintained tables decide routing and both had rotted against the vendors' own deprecation pages. This PR cleans them and adds an offline gate so the next drift is a red check instead of a silent ranking inversion.
Commit 1 —
FITNESS_TABLEhygiene + segment-boundary matchingProblem. Layer 4 of the task-fitness chain was a table of version-less family patterns (
codex,claude-sonnet,qwen,llama,gemini-pro, …) matched withString.includes. Recognition by the table, not availability, decided rank:openai/gpt-5.2-codex(OpenAI shut the codex family down 2026-07-23) scored 0.98 forcodingvia thecodexrow, while live flagships the table had never heard of (claude-fable-5-thinking-max,gpt-5.6-sol-xhigh) sat at the wildcard 0.50. Substring matching leaked further:o3matchedsolar-pro3, andgpt-4omatchedchatgpt-4o-latest— a different, retired model that inherited the live flagship's 0.9.Fix.
o1,mixtral,grok-4-fast,codex) and every version-less family row (claude-sonnet,claude-opus,claude-haiku,qwen,llama,mistral,gemini-pro,gemini-flash,grok-4,kimi-k2,glm-5). Versioned rows naming a live model stay (gpt-4o,gpt-4o-mini,gpt-4-turbo,o3,o4-mini,gemini-2.5-*,gemini-3.1-pro,deepseek-*,grok-3,glm-5.1,minimax-*).getStaticFitnessTableScorenow matches on segment boundaries (-/./// start / end), longest pattern first as before.o3still matcheso3,o3-mini,openai/o3but no longersolar-pro3;gpt-4ostill matchesgpt-4o-2024-05-13but no longerchatgpt-4o-latest— that one now falls to 0.5, which is the intended outcome since OpenAI shut it down 2026-02-17.scoring.tsis untouched — the fix is removal, not reweighting.Tests. New
open-sse/services/autoCombo/__tests__/fitness-table-hygiene-11503.test.ts(vitest): loads the lifecycle snapshot and asserts no routable retired id scores through the table at all;solar-pro3,chatgpt-4o-latest,claude-sonnet-5,claude-fable-5-thinking-max,gpt-5.6-sol-xhigh→null;gpt-4o-ministill 0.8,o3-ministill 0.95 viao3,openai/o3still resolves.Before / after — every catalog id resolved through
getTaskFitnessWithSource(id, "coding")a179ffed5b)fitness_tablewildcard_boostfitness_tableopenai/gpt-5.2-codex0.98,chatgpt-4o-latest0.90)Commit 2 —
BUILT_IN_ALIASEStargets + provider guardProblem. The alias table rewrites
body.modelon every request (modelLifecyclePolicy.ts), so a stale target is a guaranteed 404. Seven rows pointed at a retired or non-existent model — five Claude rows chained one retired id to another.claude-3-opus-20240229claude-opus-4-20250514claude-opus-4-8claude-3-sonnet-20240229claude-sonnet-4-20250514claude-sonnet-4-6claude-3-5-sonnet-latestclaude-sonnet-4-20250514claude-sonnet-4-6claude-3-haiku-20240307claude-3-5-sonnet-20241022claude-haiku-4-5-20251001claude-3-5-haiku-latestclaude-3-5-sonnet-20241022claude-haiku-4-5-20251001gemini-3-pro-highgemini-3.1-pro-highgemini-3-1-pro-high(the catalog's spelling)llama-3-8bllama3-8b-8192llama-3.1-8b-instantEvery replacement is the one the vendor publishes; the other 23 rows are untouched.
Provider guard. The table is global but the catalog is not — aggregators still serve
kimi-k2,gemini-2.0-flash,mistral-largeunder their original ids, and the unconditional rewrite broke them.resolveModelAlias(modelId, provider?)now returnsmodelIdunchanged when the given provider serves it as-is (hasKnownProviderModel, exported fromopen-sse/services/model.tsfor this). BothmodelLifecyclePolicy.tscall sites pass the provider they already have.getDeprecationNotice/isDeprecatedkeep their 1-arg behaviour, and a custom (operator-authored) alias still wins over a built-in one, provider or not.Tests. New
tests/unit/model-deprecation-aliases-11503.test.ts(node:test, 66 assertions): table-driven over everyBUILT_IN_ALIASESentry — each target must be a catalog id (checked againstREGISTRY, not a hardcoded list) and must not beretiredin the snapshot; plusresolveModelAlias("kimi-k2", "t3-web") === "kimi-k2",resolveModelAlias("kimi-k2", "fireworks") === "moonshotai/Kimi-K2",resolveModelAlias("claude-3-opus-20240229") === "claude-opus-4-8", and custom-alias precedence.Commit 3 —
check:model-lifecyclegate + vendor snapshotconfig/quality/model-lifecycle.json— 106 lifecycle entries (68retired) curated from the five first-party deprecation pages listed in itssources.replacementisnullwherever the vendor publishes none.scripts/check/check-model-lifecycle.mjs— no network, three checks, all summed before exit:FITNESS_TABLEpattern scores a routable retired id, under the new boundary rule;BUILT_IN_ALIASEStarget is retired or absent from the catalog;BUILT_IN_ALIASESor listed inallowedRetiredInCatalog.Check (1) is deliberately scoped to routable ids: legitimate versioned rows like
gpt-4oalso match retired ids the catalog never served (gpt-4o-audio-preview), and those cannot invert any routing decision. The gate setsDATA_DIRto a throwaway temp dir before importing production modules so it can never migrate the operator's database.scripts/quality/refresh-model-lifecycle.mjs— network refresher, not wired into CI. It parses Anthropic's markdown table (| API model name | Current state | Deprecated | Tentative retirement date |); for the other vendors it keeps the existing entries and printsmanual review neededrather than guessing at a shifting HTML layout. It aborts rather than write a snapshot with fewer entries than it started with — it must never silently drop one.Wired into
.github/workflows/quality.yml(gates=(…)) and.github/workflows/ci.yml'slintjob — the latter is beyond the letter of the issue but is where the other pass/fail policy gates live, so a PR tomaingets the same protection. Documented indocs/architecture/QUALITY_GATES.md. New scripts:check:model-lifecycle,quality:refresh-model-lifecycle.Tests.
tests/unit/check-model-lifecycle-gate.test.tsexercises each of the three checks plus the catalog/retired helpers against small fixtures.allowedRetiredInCatalog— the burn-down listCheck (3) is red on the current catalog, which is the point. To keep this PR mergeable the 6 offenders are allowlisted with a
TODO(#11503)note in the file; the maintainer's call is remove-from-catalog vs. add-a-forward for each:chatgpt-4o-latestclaude-3-5-sonnet-20241022claude-3-7-sonnet-20250219google/gemini-2.0-flashgemini-2.0-flashis forwarded; the prefixed form is not, since the alias table is exact-id keyed)gpt-4-0125-previewopenai/gpt-5.2-codexAdjacent but deliberately not changed here:
llama-3.1-8b-instant(the newllama-3-8btarget) is markedretiringby Groq with a 2026-08-16 date. It is the replacement Groq itself names and is live in the catalog, so it is the correct target today; the gate will flag it the moment the page flips it toretired.Deviations from the issue's proposed scope
claude-sonnet/claude-opusrows as their "known static model" probe were retargeted to a surviving versioned row (o3carries exactly the same coding 0.95 / review 0.92 numbers;gemini-2.5-procovers one review case). The behaviour under test — layer-4 resolution, longest-pattern-first, default-table fallback, case-insensitivity — is unchanged, and the fix(backend): static fitness table substring match returns wrong row for longer model ids (gpt-4o-mini scores as gpt-4o) #8603 regression assertions are intact.tests/unit/model-deprecation.test.tshad pinned the two stale Claude targets; updated to the corrected ones.Verification
npx vitest run --config vitest.mcp.config.ts open-sse/services/autoCombo/__tests__/Test Files 7 passed (7)·Tests 106 passed (106)node --test tests/unit/{task-fitness-missing-table-8603,model-deprecation-aliases-11503,check-model-lifecycle-gate}.test.ts# tests 77·# pass 77·# fail 0node --test tests/unit/{model-deprecation,model-alias-seed,model-combo-mappings,model-lifecycle,model-lifecycle-integration,8676-monsterapi-deprecation,gemini-cli-deprecation}.test.ts119 pass·0 failnpm run check:model-lifecyclePASS — snapshot 2026-08-25, 68 retired id(s), 1327 catalog id(s)(without the allowlist:6 violation(s), listed above)npm run check:cyclesOK - no cycles detected across 429 filesnpm run typecheck:corenpx eslint <touched files> --suppressions-location config/quality/eslint-suppressions.jsonnpm run check:known-symbols/check:docs-sync/check:fabricated-docs/check:test-discoverynpm run check:docs-counts