Skip to content

fix(api): enumerate tiered auto combo endpoints in /api/combos/auto - #7662

Merged
diegosouzapw merged 6 commits into
diegosouzapw:release/v3.8.49from
ekinnee:fix/tiered-auto-combos-7619
Jul 20, 2026
Merged

diegosouzapw merged 6 commits into
diegosouzapw:release/v3.8.49from
ekinnee:fix/tiered-auto-combos-7619

Conversation

@ekinnee

@ekinnee ekinnee commented Jul 17, 2026

Copy link
Copy Markdown
Contributor

Problem

The OmniRoute server's backend already supports auto/<category>[:<tier>] routing (e.g. auto/coding:free, auto/reasoning:pro, auto/vision) via suffixComposition.ts + virtualFactory.ts, and the chat handler routes them correctly. But GET /api/combos/auto only enumerates 6 flat variants (coding, fast, cheap, smart, offline, lkgp), so clients can't discover the tiered variants.

Root Cause

src/app/api/combos/auto/route.ts only iterates VALID_VARIANTS (the 6 flat variants). The 10 curated AUTO_SUFFIX_VARIANTS from builtinCatalog.ts were never exposed via this endpoint.

Solution

Added a second loop after the flat-variant loop that iterates AUTO_SUFFIX_VARIANTS, parses each via parseAutoSuffix, and materializes the combo via createVirtualAutoCombo(undefined, { category, tier }) — the same pattern used by createBuiltinAutoCombo in builtinCatalog.ts.

Changes

File Change
src/app/api/combos/auto/route.ts +43 lines — new loop enumerating 10 tiered auto combo variants

Newly exposed endpoints

  • auto/coding:fast, auto/coding:cheap, auto/coding:free, auto/coding:pro, auto/coding:reliable
  • auto/reasoning, auto/reasoning:pro
  • auto/vision, auto/multimodal

Verification

  • Syntax check: node --check passes
  • Only 1 file changed
  • Individual try/catch per variant; failures skip cleanly without breaking the list
  • Existing flat variants unchanged

Related Issues

Fixes #7619

The backend already supports auto/<category>[:<tier>] routing via
suffixComposition.ts + virtualFactory.ts, but GET /api/combos/auto only
exposed 6 flat variants. This adds a second loop enumerating the 10
curated AUTO_SUFFIX_VARIANTS (auto/coding:free, auto/coding:cheap,
auto/coding:pro, auto/reasoning, auto/vision, etc.).

Fixes #7619
@ekinnee
ekinnee requested a review from diegosouzapw as a code owner July 17, 2026 23:00
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

@manchairwang

Copy link
Copy Markdown

When I retrived model list from omniroute 3.8.48 (with another omniroute 3.8.48 server), I got these endpoints but they didn't work. I found that they are not in this 10 name list. Is there anything possible missing?

auto/glm
auto/minimax
auto/mimo
auto/zai
auto/gemma
auto/llama
auto/gemini

…combos/auto

The endpoint was missing 27 auto variants that /v1/models already
advertises, causing 404s when clients tried to use them:

- 20 template variants (auto/best-coding, auto/pro-*, auto/claude-*,
  auto/best-free, etc.)
- 7 family variants (auto/glm, auto/minimax, auto/mimo, auto/zai,
  auto/gemma, auto/llama, auto/gemini)

Fixes #7619
Refs #6453
@ekinnee

ekinnee commented Jul 18, 2026

Copy link
Copy Markdown
Contributor Author

Great catch, thanks for reporting this! You're right — those auto/glm, auto/llama, auto/gemini etc. endpoints are legitimate family variants from the #6453 feature. They were being advertised by /v1/models but GET /api/combos/auto was missing them entirely, along with the template variants (auto/best-coding, auto/pro-*, auto/claude-*, etc.).

I've pushed a fix that adds three additional enumeration phases to the combos endpoint:

  1. Template variants (20) — auto/best-coding, auto/pro-reasoning, auto/claude-opus, auto/best-free, etc.
  2. Family variants (7) — auto/glm, auto/minimax, auto/mimo, auto/zai, auto/gemma, auto/llama, auto/gemini
  3. Suffix variants (9) — auto/coding:fast, auto/reasoning:pro, etc. (already in the original PR)

The endpoint now exposes all 36 auto variants that /v1/models advertises. A seenIds set prevents duplicates across the overlapping phases.

Thanks again for the thorough testing!

Erick Kinnee added 2 commits July 18, 2026 15:33
Template variants (Phase C) now enumerate before suffix variants
(Phase B) so that overlapping ids like auto/reasoning and auto/vision
use template resolution (variant-based) rather than suffix resolution
(category-based), matching the behavior in catalog.ts.
@diegosouzapw

Copy link
Copy Markdown
Owner

Thanks for wiring the template/suffix/family enumeration through — the logic itself (AUTO_TEMPLATE_VARIANTS/AUTO_SUFFIX_VARIANTS/AUTO_FAMILY_IDS + parseAutoSuffix + createVirtualAutoCombo(variant, spec)) is correct and matches builtinCatalog.ts's own resolution order.

However, running the existing suite against this branch turns up a real regression:

$ node --import tsx/esm --test tests/unit/auto-combo-context-advertising.test.ts
✖ GET /api/combos/auto includes positive context_length for combos with candidates
  AssertionError: combo auto/llama must advertise a positive max_output_tokens, got null

Root cause: computeAdvertisedLimits() in virtualFactory.ts has a guaranteed-positive fallback for contextLength (via getTokenLimit()) but not for maxOutputTokens (via getResolvedModelCapabilities()), which can legitimately return null for some family models. That gap was invisible while only 6 flat variants were exposed — your new family enumeration (Phase D, auto/llama etc.) is what surfaces it. Could you either give maxOutputTokens the same positive fallback as contextLength, or gate Phase D enumeration on every candidate having a resolved max_output_tokens? Also worth adding a direct test asserting the new ids show up in the response (right now the PR relies entirely on pre-existing, unrelated test coverage to catch regressions).

@diegosouzapw

Copy link
Copy Markdown
Owner

(internal) Still fix-in-place: same defect as before, target files untouched by tip drift since last analysis. Needs the null max_output_tokens guard before merge.

diegosouzapw added a commit that referenced this pull request Jul 19, 2026
… (owner-approved)

The /fix-prs validation-train sweep surfaced a cluster of otherwise-clean
contributor feature PRs (#6973/#7683/#7662/#7672/#7633/#7767) whose per-PR
+1/+2 own-growth collectively exceeded the tip's 3-unit complexity slack
(2056 vs 2059). This was the 4th such block of the day (#7695/#7747/#7768
each needed helper extraction earlier). Owner approved raising both ceilings
to give new-feature PRs breathing room: complexity to 2072 (combined-cluster
2068 + 4 headroom), cognitive to 900 (combined 896 + 4). Structural shrink
stays debt (#3501); tighten via --update next cycle.
@ekinnee
ekinnee force-pushed the fix/tiered-auto-combos-7619 branch from 1a61ad5 to 9d4cf45 Compare July 19, 2026 22:08
@ekinnee

ekinnee commented Jul 19, 2026

Copy link
Copy Markdown
Contributor Author

Added a direct test asserting the new tiered, family, and template auto combo IDs show up in GET /api/combos/auto — covers the specific gap Diego called out. 9/9 tests passing locally.

computeAdvertisedLimits() has no generic default for maxOutputTokens the
way getTokenLimit() does for context length: when a combo's candidate
pool is entirely unregistered models (e.g. a no-auth provider's model
like duckduckgo-web/llama-4-scout), it legitimately returns null. The
new template/suffix/family enumeration surfaces exactly that case (e.g.
auto/llama), so /api/combos/auto advertised max_output_tokens: null and
broke tests/unit/auto-combo-context-advertising.test.ts.

Mirror the existing fallback already used by
src/app/api/v1/models/catalog.ts (advertisedContextLength || 128000,
advertisedMaxOutputTokens || 8192) at all 4 combo-push sites in this
route so clients never see a disabling null/0 for a non-empty candidate
pool.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
diegosouzapw added a commit that referenced this pull request Jul 20, 2026
…cognitive 950

Tip was at 2069/2072 and 900/900 (zero slack) after the day's 17 merges; the
remaining queue (#6973, #7662, #7719, #7744, #7779 reworks) was collectively
blocked. Owner picked the wide margin in chat (2026-07-20).
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
@diegosouzapw

Copy link
Copy Markdown
Owner

Validated in local merge-train on 192.168.0.113 @ 9d2b92a63acdf6915dbcb5d9ae7a4b18caf79d6d (FAST gates green: static + changed tests + vitest; today's full-suite parity ran in trains 4+5)

@diegosouzapw
diegosouzapw merged commit 5dc9a8a into diegosouzapw:release/v3.8.49 Jul 20, 2026
4 of 5 checks passed
@diegosouzapw diegosouzapw mentioned this pull request Jul 23, 2026
HouMinXi pushed a commit to HouMinXi/OmniRoute that referenced this pull request Aug 2, 2026
… (owner-approved)

The /fix-prs validation-train sweep surfaced a cluster of otherwise-clean
contributor feature PRs (diegosouzapw#6973/diegosouzapw#7683/diegosouzapw#7662/diegosouzapw#7672/diegosouzapw#7633/diegosouzapw#7767) whose per-PR
+1/+2 own-growth collectively exceeded the tip's 3-unit complexity slack
(2056 vs 2059). This was the 4th such block of the day (diegosouzapw#7695/diegosouzapw#7747/diegosouzapw#7768
each needed helper extraction earlier). Owner approved raising both ceilings
to give new-feature PRs breathing room: complexity to 2072 (combined-cluster
2068 + 4 headroom), cognitive to 900 (combined 896 + 4). Structural shrink
stays debt (diegosouzapw#3501); tighten via --update next cycle.
HouMinXi pushed a commit to HouMinXi/OmniRoute that referenced this pull request Aug 2, 2026
…cognitive 950

Tip was at 2069/2072 and 900/900 (zero slack) after the day's 17 merges; the
remaining queue (diegosouzapw#6973, diegosouzapw#7662, diegosouzapw#7719, diegosouzapw#7744, diegosouzapw#7779 reworks) was collectively
blocked. Owner picked the wide margin in chat (2026-07-20).
HouMinXi pushed a commit to HouMinXi/OmniRoute that referenced this pull request Aug 2, 2026
…iegosouzapw#7662)

* fix(api): enumerate tiered auto combo endpoints in /api/combos/auto

The backend already supports auto/<category>[:<tier>] routing via
suffixComposition.ts + virtualFactory.ts, but GET /api/combos/auto only
exposed 6 flat variants. This adds a second loop enumerating the 10
curated AUTO_SUFFIX_VARIANTS (auto/coding:free, auto/coding:cheap,
auto/coding:pro, auto/reasoning, auto/vision, etc.).

Fixes diegosouzapw#7619

* fix(combos): enumerate template and family auto variants in GET /api/combos/auto

The endpoint was missing 27 auto variants that /v1/models already
advertises, causing 404s when clients tried to use them:

- 20 template variants (auto/best-coding, auto/pro-*, auto/claude-*,
  auto/best-free, etc.)
- 7 family variants (auto/glm, auto/minimax, auto/mimo, auto/zai,
  auto/gemma, auto/llama, auto/gemini)

Fixes diegosouzapw#7619
Refs diegosouzapw#6453

* fix(combos): swap Phase B/C ordering to match catalog.ts

Template variants (Phase C) now enumerate before suffix variants
(Phase B) so that overlapping ids like auto/reasoning and auto/vision
use template resolution (variant-based) rather than suffix resolution
(category-based), matching the behavior in catalog.ts.

* fix(combos): fix comment labels and redundant as const

* fix(api): fall back to a positive max_output_tokens for /api/combos/auto

computeAdvertisedLimits() has no generic default for maxOutputTokens the
way getTokenLimit() does for context length: when a combo's candidate
pool is entirely unregistered models (e.g. a no-auth provider's model
like duckduckgo-web/llama-4-scout), it legitimately returns null. The
new template/suffix/family enumeration surfaces exactly that case (e.g.
auto/llama), so /api/combos/auto advertised max_output_tokens: null and
broke tests/unit/auto-combo-context-advertising.test.ts.

Mirror the existing fallback already used by
src/app/api/v1/models/catalog.ts (advertisedContextLength || 128000,
advertisedMaxOutputTokens || 8192) at all 4 combo-push sites in this
route so clients never see a disabling null/0 for a non-empty candidate
pool.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

* chore(changelog): prefix diegosouzapw#7662 fragment with markdown bullet

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: Erick Kinnee <erick@ekinnee.dev>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
… (owner-approved)

The /fix-prs validation-train sweep surfaced a cluster of otherwise-clean
contributor feature PRs (diegosouzapw#6973/diegosouzapw#7683/diegosouzapw#7662/diegosouzapw#7672/diegosouzapw#7633/diegosouzapw#7767) whose per-PR
+1/+2 own-growth collectively exceeded the tip's 3-unit complexity slack
(2056 vs 2059). This was the 4th such block of the day (diegosouzapw#7695/diegosouzapw#7747/diegosouzapw#7768
each needed helper extraction earlier). Owner approved raising both ceilings
to give new-feature PRs breathing room: complexity to 2072 (combined-cluster
2068 + 4 headroom), cognitive to 900 (combined 896 + 4). Structural shrink
stays debt (diegosouzapw#3501); tighten via --update next cycle.
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…cognitive 950

Tip was at 2069/2072 and 900/900 (zero slack) after the day's 17 merges; the
remaining queue (diegosouzapw#6973, diegosouzapw#7662, diegosouzapw#7719, diegosouzapw#7744, diegosouzapw#7779 reworks) was collectively
blocked. Owner picked the wide margin in chat (2026-07-20).
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…iegosouzapw#7662)

* fix(api): enumerate tiered auto combo endpoints in /api/combos/auto

The backend already supports auto/<category>[:<tier>] routing via
suffixComposition.ts + virtualFactory.ts, but GET /api/combos/auto only
exposed 6 flat variants. This adds a second loop enumerating the 10
curated AUTO_SUFFIX_VARIANTS (auto/coding:free, auto/coding:cheap,
auto/coding:pro, auto/reasoning, auto/vision, etc.).

Fixes diegosouzapw#7619

* fix(combos): enumerate template and family auto variants in GET /api/combos/auto

The endpoint was missing 27 auto variants that /v1/models already
advertises, causing 404s when clients tried to use them:

- 20 template variants (auto/best-coding, auto/pro-*, auto/claude-*,
  auto/best-free, etc.)
- 7 family variants (auto/glm, auto/minimax, auto/mimo, auto/zai,
  auto/gemma, auto/llama, auto/gemini)

Fixes diegosouzapw#7619
Refs diegosouzapw#6453

* fix(combos): swap Phase B/C ordering to match catalog.ts

Template variants (Phase C) now enumerate before suffix variants
(Phase B) so that overlapping ids like auto/reasoning and auto/vision
use template resolution (variant-based) rather than suffix resolution
(category-based), matching the behavior in catalog.ts.

* fix(combos): fix comment labels and redundant as const

* fix(api): fall back to a positive max_output_tokens for /api/combos/auto

computeAdvertisedLimits() has no generic default for maxOutputTokens the
way getTokenLimit() does for context length: when a combo's candidate
pool is entirely unregistered models (e.g. a no-auth provider's model
like duckduckgo-web/llama-4-scout), it legitimately returns null. The
new template/suffix/family enumeration surfaces exactly that case (e.g.
auto/llama), so /api/combos/auto advertised max_output_tokens: null and
broke tests/unit/auto-combo-context-advertising.test.ts.

Mirror the existing fallback already used by
src/app/api/v1/models/catalog.ts (advertisedContextLength || 128000,
advertisedMaxOutputTokens || 8192) at all 4 combo-push sites in this
route so clients never see a disabling null/0 for a non-empty candidate
pool.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

* chore(changelog): prefix diegosouzapw#7662 fragment with markdown bullet

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: Erick Kinnee <erick@ekinnee.dev>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(api): list tiered auto combo variants (:free/:cheap/:pro) in /api/combos/auto

3 participants