Skip to content

fix(sse): route task-aware defaults by intent, fix fitness pattern shadowing (#8602, #8603) - #8605

Merged
diegosouzapw merged 4 commits into
diegosouzapw:release/v3.8.49from
MumuTW:fix/auto-router-hardcoded-model-ids
Jul 27, 2026
Merged

diegosouzapw merged 4 commits into
diegosouzapw:release/v3.8.49from
MumuTW:fix/auto-router-hardcoded-model-ids

Conversation

@MumuTW

@MumuTW MumuTW commented Jul 25, 2026

Copy link
Copy Markdown
Contributor

Fixes #8602, fixes #8603.

Two defects in the hand-maintained model-quality tables that the auto-router leans on.

#8602 — task-aware defaults hardcoded stale model ids

coding: "deepseek/deepseek-chat",
analysis: "gemini/gemini-2.5-pro",
vision: "openai/gpt-4o",                    // 2024-05
summarization: "gemini/gemini-2.5-flash",
background: "gemini/gemini-2.5-flash-lite",

Wrong twice over:

  • The ids rot. They aged out by a generation or two, and every model release made them staler.
  • The shape bypasses the router. applyTaskAwareRouting overwrites body.model (src/sse/handlers/chat.ts:573), so a literal target skips auto-combo's 13-factor scoring (quota, circuit-breaker health, cost, latency, stability), connection cooldown and model lockout. An operator with no OpenAI connection got a vision request rewritten to openai/gpt-4o and a hard failure, where pass-through would have worked.

Refreshing the strings would only reset the rot clock, so the defaults now name intents:

task before after
coding deepseek/deepseek-chat auto/coding
analysis gemini/gemini-2.5-pro auto/reasoning
vision openai/gpt-4o auto/vision
summarization gemini/gemini-2.5-flash auto/chat:fast
background gemini/gemini-2.5-flash-lite auto/chat:cheap

creative and chat stay pass-through. These resolve on demand against the operator's actually-connected backends (suffixComposition.ts → virtualFactory.ts) and degrade gracefully as backends rotate. No provider/model literal remains in the module — the stale header comment is updated too.

Operators who pinned a specific model via PUT /api/settings/task-routing are unaffected; only the shipped defaults change. The feature remains off by default.

#8603 — fitness table matched the wrong row

lookupStaticFitnessTable returned the first String.includes hit in declaration order. FITNESS_TABLE.coding declares "gpt-4o": 0.9 before "gpt-4o-mini": 0.8, so an explicitly cheap model inherited the flagship's task fitness and its own row was unreachable:

gpt-4o-mini    -> matched "gpt-4o"      = 0.9   (own row = 0.8)
deepseek-v3.2  -> matched "deepseek-v3" = 0.85  (own row = 0.86)

Now matches longest-pattern-first, so the most specific row wins regardless of authoring order and the table can be extended without ordering hazards. Extracted as the exported getStaticFitnessTableScore — the surrounding resolution chain queries the DB (user_override / arena_elo / models.dev tier) before reaching layer 4, so testing through getTaskFitness would depend on DB fixture state.

Bounded but real: taskFit is weight 0.08 in DEFAULT_WEIGHTS, 0.37 in the quality-first mode pack, and layer 4 is what the long tail of the provider catalog actually lands on.

Validation (Hard Rule #18 — TDD)

Both written first and confirmed red.

tests/unit/task-router-auto-intent-8602.test.ts (node runner) — 4 cases, red before:

expected an auto intent id, got "deepseek/deepseek-chat"

open-sse/services/autoCombo/__tests__/taskFitness-pattern-order-8603.test.ts (vitest) — 7 cases, red before:

TypeError: getStaticFitnessTableScore is not a function

Green after, plus the full blocking vitest job:

node --import tsx/esm --test tests/unit/task-router-auto-intent-8602.test.ts   # 4/4
npm run test:vitest                                # Test Files 32 passed / Tests 281 passed

The #8602 guard is structural, not a value check — it asserts no default may be a concrete provider/model literal, so the rot cannot come back.

Gates

  • npm run typecheck:core — clean
  • npm run test:vitest — 32 files / 281 tests, all pass
  • npx eslint on changed files — 4 no-explicit-any in taskAwareRouter.ts, all pre-existing (eslint-suppressions.json records count: 4 for this file); taskFitness.ts clean. No new violations.

Not in scope

The detection patterns (taskAwareRouter.ts:46-155) are English-only literal substrings, so non-English prompts always fall through to chat. Noted in #8602, deliberately not addressed here. The static FITNESS_TABLE should eventually be deleted rather than curated once arena_elo/models.dev coverage is measured — noted in #8603.

@MumuTW
MumuTW requested a review from diegosouzapw as a code owner July 25, 2026 18:57
@MumuTW

MumuTW commented Jul 25, 2026

Copy link
Copy Markdown
Contributor Author

The two red checks are pre-existing base breakage, not this PR

Fast Quality Gates and Unit Tests fast-path (2/4) fail identically on the base commit itself. Verified by running both on a pristine worktree at 4053e2314 with no changes from this PR:

Check Pristine base 4053e2314 This PR's CI
check:complexity-ratchets complexity=2169 / cognitiveComplexity=956 complexity=2169 / cognitiveComplexity=956
tests/unit/serial/combo-quota-share-cooldown-wait-timing.test.ts 2 pass / 2 fail — 429 !== 200, 403 !== 429 same two assertions

Byte-identical. Unrelated PR #8599 fails on the same two jobs in the same CI window.

Already tracked, no new issue filed:

Per #8542, a green board can't be the evidence here — the fail-fast masks downstream gates. Local validation for this PR:

node --import tsx/esm --test tests/unit/task-router-auto-intent-8602.test.ts
# tests 4 / # pass 4 / # fail 0

npm run test:vitest        # Test Files 32 passed / Tests 281 passed
npm run typecheck:core     # clean
npx eslint <changed files>  # 4 no-explicit-any in taskAwareRouter.ts, all pre-existing
                            # (suppressions record count: 4); taskFitness.ts clean

Everything else on this PR is green: Vitest, Unit 3/4, ESLint, Docs Gates, Merge integrity, semgrep.

@MumuTW
MumuTW force-pushed the fix/auto-router-hardcoded-model-ids branch from aaa9446 to 17ed2fd Compare July 26, 2026 09:31
MumuTW added a commit to MumuTW/OmniRoute that referenced this pull request Jul 26, 2026
MumuTW added a commit to MumuTW/OmniRoute that referenced this pull request Jul 26, 2026
@MumuTW
MumuTW force-pushed the fix/auto-router-hardcoded-model-ids branch from 3bd8dbf to 1f6d998 Compare July 26, 2026 09:50
MumuTW added a commit to MumuTW/OmniRoute that referenced this pull request Jul 27, 2026
@MumuTW
MumuTW force-pushed the fix/auto-router-hardcoded-model-ids branch from 1f6d998 to d7cc7c7 Compare July 27, 2026 16:41
MumuTW added 4 commits July 28, 2026 00:56
…8601)

The T05 Task-Aware Smart Routing config was persisted to settings.taskRouting
by PUT /api/settings/task-routing but never read back, so it silently reverted
to enabled:false + the hardcoded default model map on every restart.

Two root causes, both fixed:

- No boot hydration existed. Adds hydrateTaskRoutingConfig(settings), wired into
  src/instrumentation-node.ts next to the Thinking-Budget restore (diegosouzapw#5312). It
  accepts either the JSON string the route persists or an already-parsed object,
  and fails open on malformed values. applyRuntimeSettings does not cover this
  key, same as the Global System Prompt (diegosouzapw#2470).

- The config lived in a plain module-level `let`, which is duplicated per module
  graph — a boot hydration would have landed on the instrumentation graph's copy
  and never reached the one src/sse/handlers/chat.ts reads. This is the exact
  break diegosouzapw#5312 fix-A hit on the VPS. Moves the store to the globalThis pattern
  already used by thinkingBudget.ts and systemPrompt.ts.

Runtime stats are never restored from the persisted blob.

Note the hydration is wired into instrumentation-node.ts, not the unused
src/server-init.ts.
…order (diegosouzapw#8602, diegosouzapw#8603)

Two related defects in the hand-maintained model-quality tables.

diegosouzapw#8602 — DEFAULT_TASK_MODEL_MAP hardcoded literal provider/model ids
(openai/gpt-4o, gemini/gemini-2.5-flash-lite, deepseek/deepseek-chat, ...).
Wrong twice over: the ids rotted by a generation or two, and applyTaskAwareRouting
overwrites body.model directly, so a literal target skipped auto-combo's 13-factor
scoring (quota, circuit-breaker health, cost, latency, stability), connection
cooldown and model lockout — hard-failing for any operator with no connection for
that provider. Refreshing the strings would only reset the rot clock, so the
defaults now name auto/* INTENTS that resolve against the operator's actually
connected backends:

  coding        -> auto/coding
  analysis      -> auto/reasoning
  vision        -> auto/vision
  summarization -> auto/chat:fast
  background    -> auto/chat:cheap

creative and chat stay pass-through. Operators can still pin a specific model via
PUT /api/settings/task-routing; only the shipped defaults change. No provider/model
literal remains in the module.

diegosouzapw#8603 — the pattern-shadowing fix LANDED UPSTREAM while this PR was open
(9f5be22, Train 1D). lookupStaticFitnessTable now ranks patterns longest-first,
so gpt-4o-mini no longer inherits gpt-4o's 0.9 and deepseek-v3.2 no longer inherits
deepseek-v3's 0.85. This PR therefore no longer changes that behaviour — the
upstream scan is kept verbatim.

What remains for diegosouzapw#8603 is the regression guard. The resolution chain hits the DB
(user_override / arena_elo / models.dev tier) before reaching layer 4, so asserting
the ordering through getTaskFitness would depend on DB fixture state. The layer is
exposed as getStaticFitnessTableScore and pinned directly by
taskFitness-pattern-order-8603.test.ts (7 cases), so the guarantee survives future
edits to FITNESS_TABLE. Those 7 cases were written against this PR's original
implementation and pass unchanged against the upstream one — independent
confirmation that the two are behaviourally equivalent.
@MumuTW
MumuTW force-pushed the fix/auto-router-hardcoded-model-ids branch from d7cc7c7 to 9041aa5 Compare July 27, 2026 16:56
@diegosouzapw
diegosouzapw merged commit 20aa796 into diegosouzapw:release/v3.8.49 Jul 27, 2026
15 checks passed
HouMinXi pushed a commit to HouMinXi/OmniRoute that referenced this pull request Aug 2, 2026
…adowing (diegosouzapw#8602, diegosouzapw#8603) (diegosouzapw#8605)

* fix(sse): restore task-aware routing config on restart (diegosouzapw#8601)

The T05 Task-Aware Smart Routing config was persisted to settings.taskRouting
by PUT /api/settings/task-routing but never read back, so it silently reverted
to enabled:false + the hardcoded default model map on every restart.

Two root causes, both fixed:

- No boot hydration existed. Adds hydrateTaskRoutingConfig(settings), wired into
  src/instrumentation-node.ts next to the Thinking-Budget restore (diegosouzapw#5312). It
  accepts either the JSON string the route persists or an already-parsed object,
  and fails open on malformed values. applyRuntimeSettings does not cover this
  key, same as the Global System Prompt (diegosouzapw#2470).

- The config lived in a plain module-level `let`, which is duplicated per module
  graph — a boot hydration would have landed on the instrumentation graph's copy
  and never reached the one src/sse/handlers/chat.ts reads. This is the exact
  break diegosouzapw#5312 fix-A hit on the VPS. Moves the store to the globalThis pattern
  already used by thinkingBudget.ts and systemPrompt.ts.

Runtime stats are never restored from the persisted blob.

Note the hydration is wired into instrumentation-node.ts, not the unused
src/server-init.ts.

* docs(changelog): add fragment for diegosouzapw#8604 task-routing boot restore

* fix(sse): route task-aware defaults by intent, guard fitness pattern order (diegosouzapw#8602, diegosouzapw#8603)

Two related defects in the hand-maintained model-quality tables.

diegosouzapw#8602 — DEFAULT_TASK_MODEL_MAP hardcoded literal provider/model ids
(openai/gpt-4o, gemini/gemini-2.5-flash-lite, deepseek/deepseek-chat, ...).
Wrong twice over: the ids rotted by a generation or two, and applyTaskAwareRouting
overwrites body.model directly, so a literal target skipped auto-combo's 13-factor
scoring (quota, circuit-breaker health, cost, latency, stability), connection
cooldown and model lockout — hard-failing for any operator with no connection for
that provider. Refreshing the strings would only reset the rot clock, so the
defaults now name auto/* INTENTS that resolve against the operator's actually
connected backends:

  coding        -> auto/coding
  analysis      -> auto/reasoning
  vision        -> auto/vision
  summarization -> auto/chat:fast
  background    -> auto/chat:cheap

creative and chat stay pass-through. Operators can still pin a specific model via
PUT /api/settings/task-routing; only the shipped defaults change. No provider/model
literal remains in the module.

diegosouzapw#8603 — the pattern-shadowing fix LANDED UPSTREAM while this PR was open
(f5df2cf, Train 1D). lookupStaticFitnessTable now ranks patterns longest-first,
so gpt-4o-mini no longer inherits gpt-4o's 0.9 and deepseek-v3.2 no longer inherits
deepseek-v3's 0.85. This PR therefore no longer changes that behaviour — the
upstream scan is kept verbatim.

What remains for diegosouzapw#8603 is the regression guard. The resolution chain hits the DB
(user_override / arena_elo / models.dev tier) before reaching layer 4, so asserting
the ordering through getTaskFitness would depend on DB fixture state. The layer is
exposed as getStaticFitnessTableScore and pinned directly by
taskFitness-pattern-order-8603.test.ts (7 cases), so the guarantee survives future
edits to FITNESS_TABLE. Those 7 cases were written against this PR's original
implementation and pass unchanged against the upstream one — independent
confirmation that the two are behaviourally equivalent.

* docs(changelog): add fragment for diegosouzapw#8605 task-routing intent + fitness order
@MumuTW
MumuTW deleted the fix/auto-router-hardcoded-model-ids branch September 5, 2026 10:19
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…adowing (diegosouzapw#8602, diegosouzapw#8603) (diegosouzapw#8605)

* fix(sse): restore task-aware routing config on restart (diegosouzapw#8601)

The T05 Task-Aware Smart Routing config was persisted to settings.taskRouting
by PUT /api/settings/task-routing but never read back, so it silently reverted
to enabled:false + the hardcoded default model map on every restart.

Two root causes, both fixed:

- No boot hydration existed. Adds hydrateTaskRoutingConfig(settings), wired into
  src/instrumentation-node.ts next to the Thinking-Budget restore (diegosouzapw#5312). It
  accepts either the JSON string the route persists or an already-parsed object,
  and fails open on malformed values. applyRuntimeSettings does not cover this
  key, same as the Global System Prompt (diegosouzapw#2470).

- The config lived in a plain module-level `let`, which is duplicated per module
  graph — a boot hydration would have landed on the instrumentation graph's copy
  and never reached the one src/sse/handlers/chat.ts reads. This is the exact
  break diegosouzapw#5312 fix-A hit on the VPS. Moves the store to the globalThis pattern
  already used by thinkingBudget.ts and systemPrompt.ts.

Runtime stats are never restored from the persisted blob.

Note the hydration is wired into instrumentation-node.ts, not the unused
src/server-init.ts.

* docs(changelog): add fragment for diegosouzapw#8604 task-routing boot restore

* fix(sse): route task-aware defaults by intent, guard fitness pattern order (diegosouzapw#8602, diegosouzapw#8603)

Two related defects in the hand-maintained model-quality tables.

diegosouzapw#8602 — DEFAULT_TASK_MODEL_MAP hardcoded literal provider/model ids
(openai/gpt-4o, gemini/gemini-2.5-flash-lite, deepseek/deepseek-chat, ...).
Wrong twice over: the ids rotted by a generation or two, and applyTaskAwareRouting
overwrites body.model directly, so a literal target skipped auto-combo's 13-factor
scoring (quota, circuit-breaker health, cost, latency, stability), connection
cooldown and model lockout — hard-failing for any operator with no connection for
that provider. Refreshing the strings would only reset the rot clock, so the
defaults now name auto/* INTENTS that resolve against the operator's actually
connected backends:

  coding        -> auto/coding
  analysis      -> auto/reasoning
  vision        -> auto/vision
  summarization -> auto/chat:fast
  background    -> auto/chat:cheap

creative and chat stay pass-through. Operators can still pin a specific model via
PUT /api/settings/task-routing; only the shipped defaults change. No provider/model
literal remains in the module.

diegosouzapw#8603 — the pattern-shadowing fix LANDED UPSTREAM while this PR was open
(74e2e45, Train 1D). lookupStaticFitnessTable now ranks patterns longest-first,
so gpt-4o-mini no longer inherits gpt-4o's 0.9 and deepseek-v3.2 no longer inherits
deepseek-v3's 0.85. This PR therefore no longer changes that behaviour — the
upstream scan is kept verbatim.

What remains for diegosouzapw#8603 is the regression guard. The resolution chain hits the DB
(user_override / arena_elo / models.dev tier) before reaching layer 4, so asserting
the ordering through getTaskFitness would depend on DB fixture state. The layer is
exposed as getStaticFitnessTableScore and pinned directly by
taskFitness-pattern-order-8603.test.ts (7 cases), so the guarantee survives future
edits to FITNESS_TABLE. Those 7 cases were written against this PR's original
implementation and pass unchanged against the upstream one — independent
confirmation that the two are behaviourally equivalent.

* docs(changelog): add fragment for diegosouzapw#8605 task-routing intent + fitness order
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

2 participants