Skip to content

fix(intelligence): synthesize base-model arena rows from effort/harness variants; drop MODEL_ALIAS_MAP (#11504) - #11506

Merged
diegosouzapw merged 5 commits into
diegosouzapw:release/v3.8.51from
MumuTW:fix/arena-elo-sync-variant-base
Aug 25, 2026
Merged

diegosouzapw merged 5 commits into
diegosouzapw:release/v3.8.51from
MumuTW:fix/arena-elo-sync-variant-base

Conversation

@MumuTW

@MumuTW MumuTW commented Aug 25, 2026 •

Copy link
Copy Markdown
Contributor

Fixes #11504

Stacked on #11492 — the diff includes #11492's commits until it merges; only the src/lib/arenaEloSync.ts + tests/unit/arena-elo-sync.test.ts changes are this PR's.

Problem

arenaEloSync stores leaderboard rows under the raw Arena entry name, while task-fitness layer 2 looks rows up by the request's model id. Arena scores harness × effort combinations as separate entries, so the two never meet for a bare id: on a synced DB, 9 of the arena ids carry a (codex-harness)-style annotation and 24 effort variants have no bare-id twin at all. A request for gpt-5.6-sol, claude-opus-5, gemini-3.6-flash or deepseek-v4-flash misses layer 2 with certainty and falls through to the static table.

Same file, same class of defect: MODEL_ALIAS_MAP (8 hand-maintained entries) copied a score under other names, and some of those collapse generations — gpt-5.5 → gpt-5, claude-opus-4-6-thinking → claude-opus-4, kimi-k2-thinking → moonshot/kimi-k2.

Mechanism

normalizeModelName() only lowercased and stripped a vendor prefix, so gpt-5.6-sol-xhigh (codex-harness) was stored verbatim — even a request for the exact variant id missed. And nothing ever produced the base id, because the leaderboard never lists it.

Fix

  1. normalizeModelName() drops a trailing (…) harness annotation (after lowercasing, before the vendor-prefix strip). The exported contract is otherwise unchanged.
  2. The variant row is kept as-is — it is a real routable id for several providers.
  3. A new pass after the per-model loop (withSynthesizedBaseRows) adds one synthesized base row per (base, category) for every entry that resolveScoresAs() (fix(autoCombo): inherit task fitness from base model for effort/alias variants (#11489) #11492) resolves to a routable catalog id: score = the best task fit among that base's variants, eloRaw and confidence from the winning variant (no invented confidence label — the column's vocabulary stays high/medium/low). Rationale: the per-effort entries measure the same weights at different budgets, so the model's ceiling is the max, and cost/latency are already separate factors in the 12-factor score.
  4. A base the leaderboard measured directly is never lowered: the synthesized candidate replaces it only when strictly higher.
  5. MODEL_ALIAS_MAP and its expansion loop are deleted. Alias coverage is now scoresAs + RegistryModel.aliases, which the registry owns. No replacement table.
  6. Resolution never guesses: resolveScoresAs returning via: null (e.g. grok-4.6-fast-high → grok-4.6-fast is not a catalog id) synthesizes nothing.

No schema change; bulkUpsertModelIntelligence is INSERT OR REPLACE on (model, source, category).

Measured effect

Fixture built from the real distinct arena ids on a synced DB (SELECT DISTINCT model, elo_raw FROM model_intelligence WHERE source='arena_elo' AND category='coding'), fed through transformToModelIntelligence in a DATA_DIR-isolated script:

distinct arena ids in 50 (9 of them harness-annotated)
coding rows out 64
synthesized base rows 14

New bases now reachable by a bare request: claude-opus-5, claude-sonnet-5, deepseek-v4-flash, gemini-3.5-flash, gemini-3.6-flash, gemini-3.7-flash, glm-5.2, glm-5.3, gpt-5.4, gpt-5.6-luna, gpt-5.6-sol, gpt-5.6-terra, grok-4.6, kimi-k3.

Tests

TDD (Hard Rule #18): the new describe("transformToModelIntelligence() — base-model synthesis") block was written first and failed 5/5 on the pre-fix code, then passed.

  • normalizeModelName("gpt-5.6-sol-xhigh (codex-harness)") → gpt-5.6-sol-xhigh; same with a vendor prefix and mixed case.
  • gpt-5.6-sol-xhigh (codex-harness) → rows for the variant and gpt-5.6-sol; no gpt-5.6 row (that is a vendor alias, resolved at lookup time).
  • claude-opus-5-high (1663) + claude-opus-5-max (1691) → the claude-opus-5 row carries the -max score and eloRaw 1691.
  • Leaderboard listing both claude-opus-5 (1700) and claude-opus-5-max (1691) → the base keeps 1700 (one row, explicit measurement not lowered).
  • gpt-5.5 never emits a gpt-5 row (regression guard for the deleted alias map).
  • grok-4.6-fast-high synthesizes nothing — asserted against resolveScoresAs(...).via === null rather than a frozen catalog fact, so the test stays true as the catalog moves.
  • Idempotence: two calls on the same input are equal, and (model, category) keys are unique.

Two pre-existing tests were deliberately inverted, not deleted. Both asserted MODEL_ALIAS_MAP expansion — i.e. exactly the generation collapse this PR removes (a claude-opus-4-6-thinking score copied onto claude-opus-4). They now stand as anti-collapse regression guards on the same fixture:

was is now
"expands model aliases for known models" — asserts claude-opus-4 / anthropic/claude-opus-4 ARE emitted "does not copy a variant's score onto a different generation (MODEL_ALIAS_MAP removed, #11504)" — asserts claude-opus-4-6-thinking is emitted and neither alias is
"model aliases are stored in DB alongside canonical names" — asserts a claude-opus-4 DB row exists after syncArenaElo() "stores no cross-generation alias rows in the DB (MODEL_ALIAS_MAP removed, #11504)" — same run, asserts that row is absent

Every other pre-existing test is untouched.

Verification commands

DISABLE_SQLITE_AUTO_BACKUP=true node --import tsx/esm --import ./open-sse/utils/setupPolyfill.ts --import ./tests/_setup/isolateDataDir.ts --test tests/unit/arena-elo-sync.test.ts
# tests 56 | pass 56 | fail 0

npx eslint src/lib/arenaEloSync.ts tests/unit/arena-elo-sync.test.ts --suppressions-location config/quality/eslint-suppressions.json
# 0 errors

npm run typecheck:core
# clean

⚠️ base-red inherited: #11449 (the base tip is red for unrelated reasons; #11502 clears most of them).

MumuTW added 3 commits August 25, 2026 17:59
… variants (diegosouzapw#11489)

Auto-combo task fitness scored every catalog id by exact string match, but
dispatch already resolves <model>-<effort> ids to a base model. A variant like
gpt-5.6-sol-xhigh missed both DB layers and fell to the wildcard 0.5 while its
base model was scored properly.

Adds resolveScoresAs(), a catalog-anchored resolver with three tiers: an
explicit scoresAs declared on the registry entry, a reasoning-effort suffix
stripped by one of the EXISTING dispatch splitters, and a -free tier marker.
Tiers 2-3 accept a stripped base only when it is itself a routable catalog id,
which keeps qwen3.7-max (where -max is the model) and grok-4.6-fast-high (whose
base does not exist) unresolved. No new suffix regex is introduced.

getTaskFitnessWithSource retries user_override and arena_elo against the
resolved base on a literal miss and reports <source>:inherited. The former
-free arena_elo special case (diegosouzapw#4517) is folded into the same path and extended
to user_override.

Populates scoresAs for the relations stripping cannot express: the forward
vendor alias gpt-5.6 -> gpt-5.6-sol, and cursor/agy's <version>-<family>
spelling of the Claude ids.

Fixes diegosouzapw#11489
@MumuTW
MumuTW requested a review from diegosouzapw as a code owner August 25, 2026 11:28
MumuTW added 2 commits August 25, 2026 19:28
…llapse guards (diegosouzapw#11504)

The two existing tests that asserted MODEL_ALIAS_MAP expansion encoded exactly
the generation-collapsing behaviour diegosouzapw#11504 removes: a claude-opus-4-6-thinking
score copied onto claude-opus-4. Rather than delete them, invert them —

- "expands model aliases for known models" →
  "does not copy a variant's score onto a different generation": same fixture,
  now asserts claude-opus-4-6-thinking IS emitted and claude-opus-4 /
  anthropic/claude-opus-4 are NOT.
- "model aliases are stored in DB alongside canonical names" →
  "stores no cross-generation alias rows in the DB": same assertion through
  syncArenaElo() against the DB rows.

Every other pre-existing test is untouched.
@diegosouzapw
diegosouzapw merged commit a15af27 into diegosouzapw:release/v3.8.51 Aug 25, 2026
7 of 16 checks passed
diegosouzapw pushed a commit that referenced this pull request Aug 25, 2026
… fix 7 BUILT_IN_ALIASES targets, add model-lifecycle gate (#11503) (#11507)

Validated in a combined 4-PR batch worktree off release/v3.8.51 tip. This PR's diff overlapped taskFitness.ts and autoCombo.test.ts with the already-merged #11492/#11506 — git's merge auto-resolved both hunks cleanly (non-overlapping layers: #11492/#11506 touch layer 2 arena lookup, this PR touches layer 4 static-table hygiene); verified no conflict markers remained and re-ran the full suite after boarding.
- npm run check:model-lifecycle — PASS, 68 retired ids, 1327 catalog ids, 0 violations (re-ran with the correct `node --import tsx/esm` loader after an initial bare-node invocation mistakenly failed on path-alias resolution — that was my invocation error, not the gate)
- Focused tests: fitness-table-hygiene-11503.test.ts, taskFitness-pattern-order-8603.test.ts, model-deprecation-aliases-11503.test.ts, check-model-lifecycle-gate.test.ts, model-deprecation.test.ts, autoCombo.test.ts — part of batch's 126/126 vitest + 246/246 node:test runs
- typecheck:core, file-size, changelog-integrity, complexity, cognitive-complexity, check:cycles — all OK
- Full-repo lint: 228 problems remaining, all pre-existing dashboard react-hooks/* findings unrelated to this diff (zero errors in any file this PR touches)

Thanks for this — genuinely thorough methodology (segment-boundary matching, provider-scoped alias guard, offline lifecycle gate with a documented burn-down list for the 6 remaining catalog offenders).
@MumuTW
MumuTW deleted the fix/arena-elo-sync-variant-base branch August 27, 2026 01:47
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…ss variants; drop MODEL_ALIAS_MAP (diegosouzapw#11504) (diegosouzapw#11506)

Validated in a combined 4-PR batch worktree off release/v3.8.51 tip (stacked on diegosouzapw#11492, merged first).
- TDD-first: the new base-model-synthesis describe block failed 5/5 pre-fix, passes now
- Focused test: tests/unit/arena-elo-sync.test.ts — 56/56 pass, part of batch's 246/246 node:test + 126/126 vitest runs
- typecheck:core, file-size, changelog-integrity, complexity, cognitive-complexity, check:cycles — all OK

Thanks for closing the arena-lookup gap for harness/effort-annotated leaderboard rows and retiring MODEL_ALIAS_MAP's cross-generation score copying in favor of scoresAs + registry-owned aliases.
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
… fix 7 BUILT_IN_ALIASES targets, add model-lifecycle gate (diegosouzapw#11503) (diegosouzapw#11507)

Validated in a combined 4-PR batch worktree off release/v3.8.51 tip. This PR's diff overlapped taskFitness.ts and autoCombo.test.ts with the already-merged diegosouzapw#11492/diegosouzapw#11506 — git's merge auto-resolved both hunks cleanly (non-overlapping layers: diegosouzapw#11492/diegosouzapw#11506 touch layer 2 arena lookup, this PR touches layer 4 static-table hygiene); verified no conflict markers remained and re-ran the full suite after boarding.
- npm run check:model-lifecycle — PASS, 68 retired ids, 1327 catalog ids, 0 violations (re-ran with the correct `node --import tsx/esm` loader after an initial bare-node invocation mistakenly failed on path-alias resolution — that was my invocation error, not the gate)
- Focused tests: fitness-table-hygiene-11503.test.ts, taskFitness-pattern-order-8603.test.ts, model-deprecation-aliases-11503.test.ts, check-model-lifecycle-gate.test.ts, model-deprecation.test.ts, autoCombo.test.ts — part of batch's 126/126 vitest + 246/246 node:test runs
- typecheck:core, file-size, changelog-integrity, complexity, cognitive-complexity, check:cycles — all OK
- Full-repo lint: 228 problems remaining, all pre-existing dashboard react-hooks/* findings unrelated to this diff (zero errors in any file this PR touches)

Thanks for this — genuinely thorough methodology (segment-boundary matching, provider-scoped alias guard, offline lifecycle gate with a documented burn-down list for the 6 remaining catalog offenders).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(backend): arenaEloSync stores effort/harness variant names only — 24 of 54 arena rows have no bare-id twin, so bare requests never hit layer 2

2 participants