fix(api): enumerate tiered auto combo endpoints in /api/combos/auto - #7662
diegosouzapw merged 6 commits into
Conversation
The backend already supports auto/<category>[:<tier>] routing via suffixComposition.ts + virtualFactory.ts, but GET /api/combos/auto only exposed 6 flat variants. This adds a second loop enumerating the 10 curated AUTO_SUFFIX_VARIANTS (auto/coding:free, auto/coding:cheap, auto/coding:pro, auto/reasoning, auto/vision, etc.). Fixes #7619
|
Caution The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased. |
|
When I retrived model list from omniroute 3.8.48 (with another omniroute 3.8.48 server), I got these endpoints but they didn't work. I found that they are not in this 10 name list. Is there anything possible missing? auto/glm |
…combos/auto The endpoint was missing 27 auto variants that /v1/models already advertises, causing 404s when clients tried to use them: - 20 template variants (auto/best-coding, auto/pro-*, auto/claude-*, auto/best-free, etc.) - 7 family variants (auto/glm, auto/minimax, auto/mimo, auto/zai, auto/gemma, auto/llama, auto/gemini) Fixes #7619 Refs #6453
|
Great catch, thanks for reporting this! You're right — those I've pushed a fix that adds three additional enumeration phases to the combos endpoint:
The endpoint now exposes all 36 auto variants that Thanks again for the thorough testing! |
Template variants (Phase C) now enumerate before suffix variants (Phase B) so that overlapping ids like auto/reasoning and auto/vision use template resolution (variant-based) rather than suffix resolution (category-based), matching the behavior in catalog.ts.
|
Thanks for wiring the template/suffix/family enumeration through — the logic itself (AUTO_TEMPLATE_VARIANTS/AUTO_SUFFIX_VARIANTS/AUTO_FAMILY_IDS + parseAutoSuffix + createVirtualAutoCombo(variant, spec)) is correct and matches builtinCatalog.ts's own resolution order. However, running the existing suite against this branch turns up a real regression: Root cause: |
|
(internal) Still fix-in-place: same defect as before, target files untouched by tip drift since last analysis. Needs the null max_output_tokens guard before merge. |
… (owner-approved) The /fix-prs validation-train sweep surfaced a cluster of otherwise-clean contributor feature PRs (#6973/#7683/#7662/#7672/#7633/#7767) whose per-PR +1/+2 own-growth collectively exceeded the tip's 3-unit complexity slack (2056 vs 2059). This was the 4th such block of the day (#7695/#7747/#7768 each needed helper extraction earlier). Owner approved raising both ceilings to give new-feature PRs breathing room: complexity to 2072 (combined-cluster 2068 + 4 headroom), cognitive to 900 (combined 896 + 4). Structural shrink stays debt (#3501); tighten via --update next cycle.
1a61ad5 to
9d4cf45
Compare
|
Added a direct test asserting the new tiered, family, and template auto combo IDs show up in |
computeAdvertisedLimits() has no generic default for maxOutputTokens the way getTokenLimit() does for context length: when a combo's candidate pool is entirely unregistered models (e.g. a no-auth provider's model like duckduckgo-web/llama-4-scout), it legitimately returns null. The new template/suffix/family enumeration surfaces exactly that case (e.g. auto/llama), so /api/combos/auto advertised max_output_tokens: null and broke tests/unit/auto-combo-context-advertising.test.ts. Mirror the existing fallback already used by src/app/api/v1/models/catalog.ts (advertisedContextLength || 128000, advertisedMaxOutputTokens || 8192) at all 4 combo-push sites in this route so clients never see a disabling null/0 for a non-empty candidate pool. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
|
Validated in local merge-train on 192.168.0.113 @ 9d2b92a63acdf6915dbcb5d9ae7a4b18caf79d6d (FAST gates green: static + changed tests + vitest; today's full-suite parity ran in trains 4+5) |
5dc9a8a
into
diegosouzapw:release/v3.8.49
… (owner-approved) The /fix-prs validation-train sweep surfaced a cluster of otherwise-clean contributor feature PRs (diegosouzapw#6973/diegosouzapw#7683/diegosouzapw#7662/diegosouzapw#7672/diegosouzapw#7633/diegosouzapw#7767) whose per-PR +1/+2 own-growth collectively exceeded the tip's 3-unit complexity slack (2056 vs 2059). This was the 4th such block of the day (diegosouzapw#7695/diegosouzapw#7747/diegosouzapw#7768 each needed helper extraction earlier). Owner approved raising both ceilings to give new-feature PRs breathing room: complexity to 2072 (combined-cluster 2068 + 4 headroom), cognitive to 900 (combined 896 + 4). Structural shrink stays debt (diegosouzapw#3501); tighten via --update next cycle.
…cognitive 950 Tip was at 2069/2072 and 900/900 (zero slack) after the day's 17 merges; the remaining queue (diegosouzapw#6973, diegosouzapw#7662, diegosouzapw#7719, diegosouzapw#7744, diegosouzapw#7779 reworks) was collectively blocked. Owner picked the wide margin in chat (2026-07-20).
…iegosouzapw#7662) * fix(api): enumerate tiered auto combo endpoints in /api/combos/auto The backend already supports auto/<category>[:<tier>] routing via suffixComposition.ts + virtualFactory.ts, but GET /api/combos/auto only exposed 6 flat variants. This adds a second loop enumerating the 10 curated AUTO_SUFFIX_VARIANTS (auto/coding:free, auto/coding:cheap, auto/coding:pro, auto/reasoning, auto/vision, etc.). Fixes diegosouzapw#7619 * fix(combos): enumerate template and family auto variants in GET /api/combos/auto The endpoint was missing 27 auto variants that /v1/models already advertises, causing 404s when clients tried to use them: - 20 template variants (auto/best-coding, auto/pro-*, auto/claude-*, auto/best-free, etc.) - 7 family variants (auto/glm, auto/minimax, auto/mimo, auto/zai, auto/gemma, auto/llama, auto/gemini) Fixes diegosouzapw#7619 Refs diegosouzapw#6453 * fix(combos): swap Phase B/C ordering to match catalog.ts Template variants (Phase C) now enumerate before suffix variants (Phase B) so that overlapping ids like auto/reasoning and auto/vision use template resolution (variant-based) rather than suffix resolution (category-based), matching the behavior in catalog.ts. * fix(combos): fix comment labels and redundant as const * fix(api): fall back to a positive max_output_tokens for /api/combos/auto computeAdvertisedLimits() has no generic default for maxOutputTokens the way getTokenLimit() does for context length: when a combo's candidate pool is entirely unregistered models (e.g. a no-auth provider's model like duckduckgo-web/llama-4-scout), it legitimately returns null. The new template/suffix/family enumeration surfaces exactly that case (e.g. auto/llama), so /api/combos/auto advertised max_output_tokens: null and broke tests/unit/auto-combo-context-advertising.test.ts. Mirror the existing fallback already used by src/app/api/v1/models/catalog.ts (advertisedContextLength || 128000, advertisedMaxOutputTokens || 8192) at all 4 combo-push sites in this route so clients never see a disabling null/0 for a non-empty candidate pool. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * chore(changelog): prefix diegosouzapw#7662 fragment with markdown bullet Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: Erick Kinnee <erick@ekinnee.dev> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
… (owner-approved) The /fix-prs validation-train sweep surfaced a cluster of otherwise-clean contributor feature PRs (diegosouzapw#6973/diegosouzapw#7683/diegosouzapw#7662/diegosouzapw#7672/diegosouzapw#7633/diegosouzapw#7767) whose per-PR +1/+2 own-growth collectively exceeded the tip's 3-unit complexity slack (2056 vs 2059). This was the 4th such block of the day (diegosouzapw#7695/diegosouzapw#7747/diegosouzapw#7768 each needed helper extraction earlier). Owner approved raising both ceilings to give new-feature PRs breathing room: complexity to 2072 (combined-cluster 2068 + 4 headroom), cognitive to 900 (combined 896 + 4). Structural shrink stays debt (diegosouzapw#3501); tighten via --update next cycle.
…cognitive 950 Tip was at 2069/2072 and 900/900 (zero slack) after the day's 17 merges; the remaining queue (diegosouzapw#6973, diegosouzapw#7662, diegosouzapw#7719, diegosouzapw#7744, diegosouzapw#7779 reworks) was collectively blocked. Owner picked the wide margin in chat (2026-07-20).
…iegosouzapw#7662) * fix(api): enumerate tiered auto combo endpoints in /api/combos/auto The backend already supports auto/<category>[:<tier>] routing via suffixComposition.ts + virtualFactory.ts, but GET /api/combos/auto only exposed 6 flat variants. This adds a second loop enumerating the 10 curated AUTO_SUFFIX_VARIANTS (auto/coding:free, auto/coding:cheap, auto/coding:pro, auto/reasoning, auto/vision, etc.). Fixes diegosouzapw#7619 * fix(combos): enumerate template and family auto variants in GET /api/combos/auto The endpoint was missing 27 auto variants that /v1/models already advertises, causing 404s when clients tried to use them: - 20 template variants (auto/best-coding, auto/pro-*, auto/claude-*, auto/best-free, etc.) - 7 family variants (auto/glm, auto/minimax, auto/mimo, auto/zai, auto/gemma, auto/llama, auto/gemini) Fixes diegosouzapw#7619 Refs diegosouzapw#6453 * fix(combos): swap Phase B/C ordering to match catalog.ts Template variants (Phase C) now enumerate before suffix variants (Phase B) so that overlapping ids like auto/reasoning and auto/vision use template resolution (variant-based) rather than suffix resolution (category-based), matching the behavior in catalog.ts. * fix(combos): fix comment labels and redundant as const * fix(api): fall back to a positive max_output_tokens for /api/combos/auto computeAdvertisedLimits() has no generic default for maxOutputTokens the way getTokenLimit() does for context length: when a combo's candidate pool is entirely unregistered models (e.g. a no-auth provider's model like duckduckgo-web/llama-4-scout), it legitimately returns null. The new template/suffix/family enumeration surfaces exactly that case (e.g. auto/llama), so /api/combos/auto advertised max_output_tokens: null and broke tests/unit/auto-combo-context-advertising.test.ts. Mirror the existing fallback already used by src/app/api/v1/models/catalog.ts (advertisedContextLength || 128000, advertisedMaxOutputTokens || 8192) at all 4 combo-push sites in this route so clients never see a disabling null/0 for a non-empty candidate pool. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * chore(changelog): prefix diegosouzapw#7662 fragment with markdown bullet Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: Erick Kinnee <erick@ekinnee.dev> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
Problem
The OmniRoute server's backend already supports
auto/<category>[:<tier>]routing (e.g.auto/coding:free,auto/reasoning:pro,auto/vision) viasuffixComposition.ts+virtualFactory.ts, and the chat handler routes them correctly. ButGET /api/combos/autoonly enumerates 6 flat variants (coding, fast, cheap, smart, offline, lkgp), so clients can't discover the tiered variants.Root Cause
src/app/api/combos/auto/route.tsonly iteratesVALID_VARIANTS(the 6 flat variants). The 10 curatedAUTO_SUFFIX_VARIANTSfrombuiltinCatalog.tswere never exposed via this endpoint.Solution
Added a second loop after the flat-variant loop that iterates
AUTO_SUFFIX_VARIANTS, parses each viaparseAutoSuffix, and materializes the combo viacreateVirtualAutoCombo(undefined, { category, tier })— the same pattern used bycreateBuiltinAutoComboinbuiltinCatalog.ts.Changes
src/app/api/combos/auto/route.tsNewly exposed endpoints
auto/coding:fast,auto/coding:cheap,auto/coding:free,auto/coding:pro,auto/coding:reliableauto/reasoning,auto/reasoning:proauto/vision,auto/multimodalVerification
node --checkpassesRelated Issues
Fixes #7619