Repository navigation
fix(intelligence): synthesize base-model arena rows from effort/harness variants; drop MODEL_ALIAS_MAP (#11504) - #11506
Merged
diegosouzapw merged 5 commits intoAug 25, 2026
Conversation
… variants (diegosouzapw#11489) Auto-combo task fitness scored every catalog id by exact string match, but dispatch already resolves <model>-<effort> ids to a base model. A variant like gpt-5.6-sol-xhigh missed both DB layers and fell to the wildcard 0.5 while its base model was scored properly. Adds resolveScoresAs(), a catalog-anchored resolver with three tiers: an explicit scoresAs declared on the registry entry, a reasoning-effort suffix stripped by one of the EXISTING dispatch splitters, and a -free tier marker. Tiers 2-3 accept a stripped base only when it is itself a routable catalog id, which keeps qwen3.7-max (where -max is the model) and grok-4.6-fast-high (whose base does not exist) unresolved. No new suffix regex is introduced. getTaskFitnessWithSource retries user_override and arena_elo against the resolved base on a literal miss and reports <source>:inherited. The former -free arena_elo special case (diegosouzapw#4517) is folded into the same path and extended to user_override. Populates scoresAs for the relations stripping cannot express: the forward vendor alias gpt-5.6 -> gpt-5.6-sol, and cursor/agy's <version>-<family> spelling of the Claude ids. Fixes diegosouzapw#11489
…ss variants; drop MODEL_ALIAS_MAP (diegosouzapw#11504)
…llapse guards (diegosouzapw#11504) The two existing tests that asserted MODEL_ALIAS_MAP expansion encoded exactly the generation-collapsing behaviour diegosouzapw#11504 removes: a claude-opus-4-6-thinking score copied onto claude-opus-4. Rather than delete them, invert them — - "expands model aliases for known models" → "does not copy a variant's score onto a different generation": same fixture, now asserts claude-opus-4-6-thinking IS emitted and claude-opus-4 / anthropic/claude-opus-4 are NOT. - "model aliases are stored in DB alongside canonical names" → "stores no cross-generation alias rows in the DB": same assertion through syncArenaElo() against the DB rows. Every other pre-existing test is untouched.
diegosouzapw
merged commit Aug 25, 2026
a15af27
into
diegosouzapw:release/v3.8.51
7 of 16 checks passed
diegosouzapw
pushed a commit
that referenced
this pull request
Aug 25, 2026
… fix 7 BUILT_IN_ALIASES targets, add model-lifecycle gate (#11503) (#11507) Validated in a combined 4-PR batch worktree off release/v3.8.51 tip. This PR's diff overlapped taskFitness.ts and autoCombo.test.ts with the already-merged #11492/#11506 — git's merge auto-resolved both hunks cleanly (non-overlapping layers: #11492/#11506 touch layer 2 arena lookup, this PR touches layer 4 static-table hygiene); verified no conflict markers remained and re-ran the full suite after boarding. - npm run check:model-lifecycle — PASS, 68 retired ids, 1327 catalog ids, 0 violations (re-ran with the correct `node --import tsx/esm` loader after an initial bare-node invocation mistakenly failed on path-alias resolution — that was my invocation error, not the gate) - Focused tests: fitness-table-hygiene-11503.test.ts, taskFitness-pattern-order-8603.test.ts, model-deprecation-aliases-11503.test.ts, check-model-lifecycle-gate.test.ts, model-deprecation.test.ts, autoCombo.test.ts — part of batch's 126/126 vitest + 246/246 node:test runs - typecheck:core, file-size, changelog-integrity, complexity, cognitive-complexity, check:cycles — all OK - Full-repo lint: 228 problems remaining, all pre-existing dashboard react-hooks/* findings unrelated to this diff (zero errors in any file this PR touches) Thanks for this — genuinely thorough methodology (segment-boundary matching, provider-scoped alias guard, offline lifecycle gate with a documented burn-down list for the 6 remaining catalog offenders).
muhamadgalihsaputra
pushed a commit
to niyatna/NiyatnaRoute
that referenced
this pull request
Sep 27, 2026
…ss variants; drop MODEL_ALIAS_MAP (diegosouzapw#11504) (diegosouzapw#11506) Validated in a combined 4-PR batch worktree off release/v3.8.51 tip (stacked on diegosouzapw#11492, merged first). - TDD-first: the new base-model-synthesis describe block failed 5/5 pre-fix, passes now - Focused test: tests/unit/arena-elo-sync.test.ts — 56/56 pass, part of batch's 246/246 node:test + 126/126 vitest runs - typecheck:core, file-size, changelog-integrity, complexity, cognitive-complexity, check:cycles — all OK Thanks for closing the arena-lookup gap for harness/effort-annotated leaderboard rows and retiring MODEL_ALIAS_MAP's cross-generation score copying in favor of scoresAs + registry-owned aliases.
muhamadgalihsaputra
pushed a commit
to niyatna/NiyatnaRoute
that referenced
this pull request
Sep 27, 2026
… fix 7 BUILT_IN_ALIASES targets, add model-lifecycle gate (diegosouzapw#11503) (diegosouzapw#11507) Validated in a combined 4-PR batch worktree off release/v3.8.51 tip. This PR's diff overlapped taskFitness.ts and autoCombo.test.ts with the already-merged diegosouzapw#11492/diegosouzapw#11506 — git's merge auto-resolved both hunks cleanly (non-overlapping layers: diegosouzapw#11492/diegosouzapw#11506 touch layer 2 arena lookup, this PR touches layer 4 static-table hygiene); verified no conflict markers remained and re-ran the full suite after boarding. - npm run check:model-lifecycle — PASS, 68 retired ids, 1327 catalog ids, 0 violations (re-ran with the correct `node --import tsx/esm` loader after an initial bare-node invocation mistakenly failed on path-alias resolution — that was my invocation error, not the gate) - Focused tests: fitness-table-hygiene-11503.test.ts, taskFitness-pattern-order-8603.test.ts, model-deprecation-aliases-11503.test.ts, check-model-lifecycle-gate.test.ts, model-deprecation.test.ts, autoCombo.test.ts — part of batch's 126/126 vitest + 246/246 node:test runs - typecheck:core, file-size, changelog-integrity, complexity, cognitive-complexity, check:cycles — all OK - Full-repo lint: 228 problems remaining, all pre-existing dashboard react-hooks/* findings unrelated to this diff (zero errors in any file this PR touches) Thanks for this — genuinely thorough methodology (segment-boundary matching, provider-scoped alias guard, offline lifecycle gate with a documented burn-down list for the 6 remaining catalog offenders).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #11504
Stacked on #11492 — the diff includes #11492's commits until it merges; only the
src/lib/arenaEloSync.ts+tests/unit/arena-elo-sync.test.tschanges are this PR's.Problem
arenaEloSyncstores leaderboard rows under the raw Arena entry name, while task-fitness layer 2 looks rows up by the request's model id. Arena scores harness × effort combinations as separate entries, so the two never meet for a bare id: on a synced DB, 9 of the arena ids carry a(codex-harness)-style annotation and 24 effort variants have no bare-id twin at all. A request forgpt-5.6-sol,claude-opus-5,gemini-3.6-flashordeepseek-v4-flashmisses layer 2 with certainty and falls through to the static table.Same file, same class of defect:
MODEL_ALIAS_MAP(8 hand-maintained entries) copied a score under other names, and some of those collapse generations —gpt-5.5 → gpt-5,claude-opus-4-6-thinking → claude-opus-4,kimi-k2-thinking → moonshot/kimi-k2.Mechanism
normalizeModelName()only lowercased and stripped a vendor prefix, sogpt-5.6-sol-xhigh (codex-harness)was stored verbatim — even a request for the exact variant id missed. And nothing ever produced the base id, because the leaderboard never lists it.Fix
normalizeModelName()drops a trailing(…)harness annotation (after lowercasing, before the vendor-prefix strip). The exported contract is otherwise unchanged.withSynthesizedBaseRows) adds one synthesized base row per (base, category) for every entry thatresolveScoresAs()(fix(autoCombo): inherit task fitness from base model for effort/alias variants (#11489) #11492) resolves to a routable catalog id: score = the best task fit among that base's variants,eloRawandconfidencefrom the winning variant (no invented confidence label — the column's vocabulary stays high/medium/low). Rationale: the per-effort entries measure the same weights at different budgets, so the model's ceiling is the max, and cost/latency are already separate factors in the 12-factor score.MODEL_ALIAS_MAPand its expansion loop are deleted. Alias coverage is nowscoresAs+RegistryModel.aliases, which the registry owns. No replacement table.resolveScoresAsreturningvia: null(e.g.grok-4.6-fast-high→grok-4.6-fastis not a catalog id) synthesizes nothing.No schema change;
bulkUpsertModelIntelligenceisINSERT OR REPLACEon(model, source, category).Measured effect
Fixture built from the real distinct arena ids on a synced DB (
SELECT DISTINCT model, elo_raw FROM model_intelligence WHERE source='arena_elo' AND category='coding'), fed throughtransformToModelIntelligencein aDATA_DIR-isolated script:codingrows outNew bases now reachable by a bare request:
claude-opus-5,claude-sonnet-5,deepseek-v4-flash,gemini-3.5-flash,gemini-3.6-flash,gemini-3.7-flash,glm-5.2,glm-5.3,gpt-5.4,gpt-5.6-luna,gpt-5.6-sol,gpt-5.6-terra,grok-4.6,kimi-k3.Tests
TDD (Hard Rule #18): the new
describe("transformToModelIntelligence() — base-model synthesis")block was written first and failed 5/5 on the pre-fix code, then passed.normalizeModelName("gpt-5.6-sol-xhigh (codex-harness)")→gpt-5.6-sol-xhigh; same with a vendor prefix and mixed case.gpt-5.6-sol-xhigh (codex-harness)→ rows for the variant andgpt-5.6-sol; nogpt-5.6row (that is a vendor alias, resolved at lookup time).claude-opus-5-high(1663) +claude-opus-5-max(1691) → theclaude-opus-5row carries the-maxscore andeloRaw1691.claude-opus-5(1700) andclaude-opus-5-max(1691) → the base keeps 1700 (one row, explicit measurement not lowered).gpt-5.5never emits agpt-5row (regression guard for the deleted alias map).grok-4.6-fast-highsynthesizes nothing — asserted againstresolveScoresAs(...).via === nullrather than a frozen catalog fact, so the test stays true as the catalog moves.(model, category)keys are unique.Two pre-existing tests were deliberately inverted, not deleted. Both asserted
MODEL_ALIAS_MAPexpansion — i.e. exactly the generation collapse this PR removes (aclaude-opus-4-6-thinkingscore copied ontoclaude-opus-4). They now stand as anti-collapse regression guards on the same fixture:"expands model aliases for known models"— assertsclaude-opus-4/anthropic/claude-opus-4ARE emitted"does not copy a variant's score onto a different generation (MODEL_ALIAS_MAP removed, #11504)"— assertsclaude-opus-4-6-thinkingis emitted and neither alias is"model aliases are stored in DB alongside canonical names"— asserts aclaude-opus-4DB row exists aftersyncArenaElo()"stores no cross-generation alias rows in the DB (MODEL_ALIAS_MAP removed, #11504)"— same run, asserts that row is absentEvery other pre-existing test is untouched.
Verification commands