Repository navigation
fix(sse): route task-aware defaults by intent, fix fitness pattern shadowing (#8602, #8603) - #8605
Merged
diegosouzapw merged 4 commits intoJul 27, 2026
Conversation
Contributor
Author
The two red checks are pre-existing base breakage, not this PR
Byte-identical. Unrelated PR #8599 fails on the same two jobs in the same CI window. Already tracked, no new issue filed:
Per #8542, a green board can't be the evidence here — the fail-fast masks downstream gates. Local validation for this PR: Everything else on this PR is green: Vitest, Unit 3/4, ESLint, Docs Gates, Merge integrity, semgrep. |
MumuTW
force-pushed
the
fix/auto-router-hardcoded-model-ids
branch
from
July 26, 2026 09:31
aaa9446 to
17ed2fd
Compare
MumuTW
added a commit
to MumuTW/OmniRoute
that referenced
this pull request
Jul 26, 2026
…nt + fitness order
MumuTW
added a commit
to MumuTW/OmniRoute
that referenced
this pull request
Jul 26, 2026
…nt + fitness order
MumuTW
force-pushed
the
fix/auto-router-hardcoded-model-ids
branch
from
July 26, 2026 09:50
3bd8dbf to
1f6d998
Compare
MumuTW
added a commit
to MumuTW/OmniRoute
that referenced
this pull request
Jul 27, 2026
…nt + fitness order
MumuTW
force-pushed
the
fix/auto-router-hardcoded-model-ids
branch
from
July 27, 2026 16:41
1f6d998 to
d7cc7c7
Compare
…8601) The T05 Task-Aware Smart Routing config was persisted to settings.taskRouting by PUT /api/settings/task-routing but never read back, so it silently reverted to enabled:false + the hardcoded default model map on every restart. Two root causes, both fixed: - No boot hydration existed. Adds hydrateTaskRoutingConfig(settings), wired into src/instrumentation-node.ts next to the Thinking-Budget restore (diegosouzapw#5312). It accepts either the JSON string the route persists or an already-parsed object, and fails open on malformed values. applyRuntimeSettings does not cover this key, same as the Global System Prompt (diegosouzapw#2470). - The config lived in a plain module-level `let`, which is duplicated per module graph — a boot hydration would have landed on the instrumentation graph's copy and never reached the one src/sse/handlers/chat.ts reads. This is the exact break diegosouzapw#5312 fix-A hit on the VPS. Moves the store to the globalThis pattern already used by thinkingBudget.ts and systemPrompt.ts. Runtime stats are never restored from the persisted blob. Note the hydration is wired into instrumentation-node.ts, not the unused src/server-init.ts.
…order (diegosouzapw#8602, diegosouzapw#8603) Two related defects in the hand-maintained model-quality tables. diegosouzapw#8602 — DEFAULT_TASK_MODEL_MAP hardcoded literal provider/model ids (openai/gpt-4o, gemini/gemini-2.5-flash-lite, deepseek/deepseek-chat, ...). Wrong twice over: the ids rotted by a generation or two, and applyTaskAwareRouting overwrites body.model directly, so a literal target skipped auto-combo's 13-factor scoring (quota, circuit-breaker health, cost, latency, stability), connection cooldown and model lockout — hard-failing for any operator with no connection for that provider. Refreshing the strings would only reset the rot clock, so the defaults now name auto/* INTENTS that resolve against the operator's actually connected backends: coding -> auto/coding analysis -> auto/reasoning vision -> auto/vision summarization -> auto/chat:fast background -> auto/chat:cheap creative and chat stay pass-through. Operators can still pin a specific model via PUT /api/settings/task-routing; only the shipped defaults change. No provider/model literal remains in the module. diegosouzapw#8603 — the pattern-shadowing fix LANDED UPSTREAM while this PR was open (9f5be22, Train 1D). lookupStaticFitnessTable now ranks patterns longest-first, so gpt-4o-mini no longer inherits gpt-4o's 0.9 and deepseek-v3.2 no longer inherits deepseek-v3's 0.85. This PR therefore no longer changes that behaviour — the upstream scan is kept verbatim. What remains for diegosouzapw#8603 is the regression guard. The resolution chain hits the DB (user_override / arena_elo / models.dev tier) before reaching layer 4, so asserting the ordering through getTaskFitness would depend on DB fixture state. The layer is exposed as getStaticFitnessTableScore and pinned directly by taskFitness-pattern-order-8603.test.ts (7 cases), so the guarantee survives future edits to FITNESS_TABLE. Those 7 cases were written against this PR's original implementation and pass unchanged against the upstream one — independent confirmation that the two are behaviourally equivalent.
…nt + fitness order
MumuTW
force-pushed
the
fix/auto-router-hardcoded-model-ids
branch
from
July 27, 2026 16:56
d7cc7c7 to
9041aa5
Compare
This was referenced Jul 28, 2026
Merged
HouMinXi
pushed a commit
to HouMinXi/OmniRoute
that referenced
this pull request
Aug 2, 2026
…adowing (diegosouzapw#8602, diegosouzapw#8603) (diegosouzapw#8605) * fix(sse): restore task-aware routing config on restart (diegosouzapw#8601) The T05 Task-Aware Smart Routing config was persisted to settings.taskRouting by PUT /api/settings/task-routing but never read back, so it silently reverted to enabled:false + the hardcoded default model map on every restart. Two root causes, both fixed: - No boot hydration existed. Adds hydrateTaskRoutingConfig(settings), wired into src/instrumentation-node.ts next to the Thinking-Budget restore (diegosouzapw#5312). It accepts either the JSON string the route persists or an already-parsed object, and fails open on malformed values. applyRuntimeSettings does not cover this key, same as the Global System Prompt (diegosouzapw#2470). - The config lived in a plain module-level `let`, which is duplicated per module graph — a boot hydration would have landed on the instrumentation graph's copy and never reached the one src/sse/handlers/chat.ts reads. This is the exact break diegosouzapw#5312 fix-A hit on the VPS. Moves the store to the globalThis pattern already used by thinkingBudget.ts and systemPrompt.ts. Runtime stats are never restored from the persisted blob. Note the hydration is wired into instrumentation-node.ts, not the unused src/server-init.ts. * docs(changelog): add fragment for diegosouzapw#8604 task-routing boot restore * fix(sse): route task-aware defaults by intent, guard fitness pattern order (diegosouzapw#8602, diegosouzapw#8603) Two related defects in the hand-maintained model-quality tables. diegosouzapw#8602 — DEFAULT_TASK_MODEL_MAP hardcoded literal provider/model ids (openai/gpt-4o, gemini/gemini-2.5-flash-lite, deepseek/deepseek-chat, ...). Wrong twice over: the ids rotted by a generation or two, and applyTaskAwareRouting overwrites body.model directly, so a literal target skipped auto-combo's 13-factor scoring (quota, circuit-breaker health, cost, latency, stability), connection cooldown and model lockout — hard-failing for any operator with no connection for that provider. Refreshing the strings would only reset the rot clock, so the defaults now name auto/* INTENTS that resolve against the operator's actually connected backends: coding -> auto/coding analysis -> auto/reasoning vision -> auto/vision summarization -> auto/chat:fast background -> auto/chat:cheap creative and chat stay pass-through. Operators can still pin a specific model via PUT /api/settings/task-routing; only the shipped defaults change. No provider/model literal remains in the module. diegosouzapw#8603 — the pattern-shadowing fix LANDED UPSTREAM while this PR was open (f5df2cf, Train 1D). lookupStaticFitnessTable now ranks patterns longest-first, so gpt-4o-mini no longer inherits gpt-4o's 0.9 and deepseek-v3.2 no longer inherits deepseek-v3's 0.85. This PR therefore no longer changes that behaviour — the upstream scan is kept verbatim. What remains for diegosouzapw#8603 is the regression guard. The resolution chain hits the DB (user_override / arena_elo / models.dev tier) before reaching layer 4, so asserting the ordering through getTaskFitness would depend on DB fixture state. The layer is exposed as getStaticFitnessTableScore and pinned directly by taskFitness-pattern-order-8603.test.ts (7 cases), so the guarantee survives future edits to FITNESS_TABLE. Those 7 cases were written against this PR's original implementation and pass unchanged against the upstream one — independent confirmation that the two are behaviourally equivalent. * docs(changelog): add fragment for diegosouzapw#8605 task-routing intent + fitness order
muhamadgalihsaputra
pushed a commit
to niyatna/NiyatnaRoute
that referenced
this pull request
Sep 27, 2026
…adowing (diegosouzapw#8602, diegosouzapw#8603) (diegosouzapw#8605) * fix(sse): restore task-aware routing config on restart (diegosouzapw#8601) The T05 Task-Aware Smart Routing config was persisted to settings.taskRouting by PUT /api/settings/task-routing but never read back, so it silently reverted to enabled:false + the hardcoded default model map on every restart. Two root causes, both fixed: - No boot hydration existed. Adds hydrateTaskRoutingConfig(settings), wired into src/instrumentation-node.ts next to the Thinking-Budget restore (diegosouzapw#5312). It accepts either the JSON string the route persists or an already-parsed object, and fails open on malformed values. applyRuntimeSettings does not cover this key, same as the Global System Prompt (diegosouzapw#2470). - The config lived in a plain module-level `let`, which is duplicated per module graph — a boot hydration would have landed on the instrumentation graph's copy and never reached the one src/sse/handlers/chat.ts reads. This is the exact break diegosouzapw#5312 fix-A hit on the VPS. Moves the store to the globalThis pattern already used by thinkingBudget.ts and systemPrompt.ts. Runtime stats are never restored from the persisted blob. Note the hydration is wired into instrumentation-node.ts, not the unused src/server-init.ts. * docs(changelog): add fragment for diegosouzapw#8604 task-routing boot restore * fix(sse): route task-aware defaults by intent, guard fitness pattern order (diegosouzapw#8602, diegosouzapw#8603) Two related defects in the hand-maintained model-quality tables. diegosouzapw#8602 — DEFAULT_TASK_MODEL_MAP hardcoded literal provider/model ids (openai/gpt-4o, gemini/gemini-2.5-flash-lite, deepseek/deepseek-chat, ...). Wrong twice over: the ids rotted by a generation or two, and applyTaskAwareRouting overwrites body.model directly, so a literal target skipped auto-combo's 13-factor scoring (quota, circuit-breaker health, cost, latency, stability), connection cooldown and model lockout — hard-failing for any operator with no connection for that provider. Refreshing the strings would only reset the rot clock, so the defaults now name auto/* INTENTS that resolve against the operator's actually connected backends: coding -> auto/coding analysis -> auto/reasoning vision -> auto/vision summarization -> auto/chat:fast background -> auto/chat:cheap creative and chat stay pass-through. Operators can still pin a specific model via PUT /api/settings/task-routing; only the shipped defaults change. No provider/model literal remains in the module. diegosouzapw#8603 — the pattern-shadowing fix LANDED UPSTREAM while this PR was open (74e2e45, Train 1D). lookupStaticFitnessTable now ranks patterns longest-first, so gpt-4o-mini no longer inherits gpt-4o's 0.9 and deepseek-v3.2 no longer inherits deepseek-v3's 0.85. This PR therefore no longer changes that behaviour — the upstream scan is kept verbatim. What remains for diegosouzapw#8603 is the regression guard. The resolution chain hits the DB (user_override / arena_elo / models.dev tier) before reaching layer 4, so asserting the ordering through getTaskFitness would depend on DB fixture state. The layer is exposed as getStaticFitnessTableScore and pinned directly by taskFitness-pattern-order-8603.test.ts (7 cases), so the guarantee survives future edits to FITNESS_TABLE. Those 7 cases were written against this PR's original implementation and pass unchanged against the upstream one — independent confirmation that the two are behaviourally equivalent. * docs(changelog): add fragment for diegosouzapw#8605 task-routing intent + fitness order
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #8602, fixes #8603.
Two defects in the hand-maintained model-quality tables that the auto-router leans on.
#8602 — task-aware defaults hardcoded stale model ids
Wrong twice over:
applyTaskAwareRoutingoverwritesbody.model(src/sse/handlers/chat.ts:573), so a literal target skips auto-combo's 13-factor scoring (quota, circuit-breaker health, cost, latency, stability), connection cooldown and model lockout. An operator with no OpenAI connection got avisionrequest rewritten toopenai/gpt-4oand a hard failure, where pass-through would have worked.Refreshing the strings would only reset the rot clock, so the defaults now name intents:
codingdeepseek/deepseek-chatauto/codinganalysisgemini/gemini-2.5-proauto/reasoningvisionopenai/gpt-4oauto/visionsummarizationgemini/gemini-2.5-flashauto/chat:fastbackgroundgemini/gemini-2.5-flash-liteauto/chat:cheapcreativeandchatstay pass-through. These resolve on demand against the operator's actually-connected backends (suffixComposition.ts→virtualFactory.ts) and degrade gracefully as backends rotate. No provider/model literal remains in the module — the stale header comment is updated too.Operators who pinned a specific model via
PUT /api/settings/task-routingare unaffected; only the shipped defaults change. The feature remains off by default.#8603 — fitness table matched the wrong row
lookupStaticFitnessTablereturned the firstString.includeshit in declaration order.FITNESS_TABLE.codingdeclares"gpt-4o": 0.9before"gpt-4o-mini": 0.8, so an explicitly cheap model inherited the flagship's task fitness and its own row was unreachable:Now matches longest-pattern-first, so the most specific row wins regardless of authoring order and the table can be extended without ordering hazards. Extracted as the exported
getStaticFitnessTableScore— the surrounding resolution chain queries the DB (user_override / arena_elo / models.dev tier) before reaching layer 4, so testing throughgetTaskFitnesswould depend on DB fixture state.Bounded but real:
taskFitis weight0.08inDEFAULT_WEIGHTS,0.37in thequality-firstmode pack, and layer 4 is what the long tail of the provider catalog actually lands on.Validation (Hard Rule #18 — TDD)
Both written first and confirmed red.
tests/unit/task-router-auto-intent-8602.test.ts(node runner) — 4 cases, red before:open-sse/services/autoCombo/__tests__/taskFitness-pattern-order-8603.test.ts(vitest) — 7 cases, red before:Green after, plus the full blocking vitest job:
The #8602 guard is structural, not a value check — it asserts no default may be a concrete provider/model literal, so the rot cannot come back.
Gates
npm run typecheck:core— cleannpm run test:vitest— 32 files / 281 tests, all passnpx eslinton changed files — 4no-explicit-anyintaskAwareRouter.ts, all pre-existing (eslint-suppressions.jsonrecordscount: 4for this file);taskFitness.tsclean. No new violations.Not in scope
The detection patterns (
taskAwareRouter.ts:46-155) are English-only literal substrings, so non-English prompts always fall through tochat. Noted in #8602, deliberately not addressed here. The staticFITNESS_TABLEshould eventually be deleted rather than curated once arena_elo/models.dev coverage is measured — noted in #8603.