fix(compression): apply compression combo assignments to routing combos - #7779
Conversation
The backend already supports auto/<category>[:<tier>] routing via suffixComposition.ts + virtualFactory.ts, but GET /api/combos/auto only exposed 6 flat variants. This adds a second loop enumerating the 10 curated AUTO_SUFFIX_VARIANTS (auto/coding:free, auto/coding:cheap, auto/coding:pro, auto/reasoning, auto/vision, etc.). Fixes #7619
…combos/auto The endpoint was missing 27 auto variants that /v1/models already advertises, causing 404s when clients tried to use them: - 20 template variants (auto/best-coding, auto/pro-*, auto/claude-*, auto/best-free, etc.) - 7 family variants (auto/glm, auto/minimax, auto/mimo, auto/zai, auto/gemma, auto/llama, auto/gemini) Fixes #7619 Refs #6453
Template variants (Phase C) now enumerate before suffix variants (Phase B) so that overlapping ids like auto/reasoning and auto/vision use template resolution (variant-based) rather than suffix resolution (category-based), matching the behavior in catalog.ts.
Routing combos (e.g. codex, free-only, or-free) use provider-prefixed model strings like codex/gpt-5.5 and go through handleSingleModelChat, which passes comboName: null, isCombo: false. The compression combo assignment lookup in chatCore.ts was gated behind if (isCombo && comboName), so routing combos never had their compression combos applied. Fix: - Add routingComboId parameter threaded through handleSingleModelChat → executeChatWithBreaker → handleChatCore - In handleChat(), resolve the routing combo UUID from the model string's provider prefix via getComboByName - In chatCore.ts, change the gate to (isCombo && comboName) || routingComboId and add routingComboId to the lookup key array Fixes #7771
There was a problem hiding this comment.
Code Review
This pull request introduces numerous features, bug fixes, and maintenance updates, including a new Mergify merge queue configuration, a CLI fast-path for version queries, and an improved native binary compatibility check for better-sqlite3. However, a critical issue was identified in bin/cli/runtime/nativeDeps.mjs where spawnSync is used within the new probeNativeBinaryLoadable function but is not imported, which will cause a silent ReferenceError and force unnecessary binary rebuilds on every startup.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
| */ | ||
| function probeNativeBinaryLoadable(binary) { | ||
| try { | ||
| const res = spawnSync( |
There was a problem hiding this comment.
The function probeNativeBinaryLoadable uses spawnSync to probe the native binary in a subprocess. However, spawnSync is not imported in this file (bin/cli/runtime/nativeDeps.mjs). This will cause a ReferenceError when spawnSync is called, which is caught by the try-catch block and silently returns false. As a result, isBetterSqliteBinaryValid() will always return false, triggering an unnecessary rebuild of the better-sqlite3 binary on every startup.\n\nPlease add the missing import at the top of bin/cli/runtime/nativeDeps.mjs:\njavascript\nimport { spawnSync } from \"node:child_process\";\n
|
Thanks for the clear root-cause writeup — this matches the code exactly: the compression-combo lookup in Two things before this can land:
|
Whatever is easiest for you, I can do the work if you'd like. |
…d + fix changelog fragment bullet Resolves merge conflicts from release/v3.8.49: keeps both the tip's providerId param on hasBlockingProxyAssignment and #7779's new routingComboId threading; merges both new fields (routingComboId, reasoningDecision/reasoningIntent/reasoningRequestTags) into the handleSingleModelChat runtime options. Fixes a malformed changelog.d fragment (missing leading bullet) and rebaselines chatHelpers.ts file-size by the PR's own +1 net growth (routingComboId passthrough). Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
…eAdvertisedLimits GET /api/combos/auto now enumerates auto/<family> variants (auto/llama, auto/glm, etc). computeAdvertisedLimits() already guaranteed a positive contextLength for any non-empty candidate pool via getTokenLimit()'s fallback chain, but had no equivalent fallback for maxOutputTokens — candidates whose registry entry and models.dev sync data both lack that field (common for no-auth/free-tier providers matching a family filter, e.g. llama-* on groq/bazaarlink/etc) left maxOutputTokens null, which tests/unit/auto-combo-context-advertising.test.ts catches as a contract violation of the endpoint (opencode disables smart auto-compaction when a limit is falsy — the same bug class this module's docstring already describes for contextLength). Fall back to a conservative generic default (4096) when no candidate in the pool resolves a known maxOutputTokens, mirroring the existing contextLength guarantee. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
…catalog convention (8192) Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
# Conflicts: # config/quality/file-size-baseline.json # src/app/api/combos/auto/route.ts
…ead (876->877) Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
|
Validated in local merge-train on 192.168.0.113 @ 4c7058069292fa601621cb616d233be21e096e06 (FAST gates green: static + changed tests + vitest; today's full-suite parity ran in trains 4+5) |
eba6eca
into
diegosouzapw:release/v3.8.49
…cognitive 950 Tip was at 2069/2072 and 900/900 (zero slack) after the day's 17 merges; the remaining queue (diegosouzapw#6973, diegosouzapw#7662, diegosouzapw#7719, diegosouzapw#7744, diegosouzapw#7779 reworks) was collectively blocked. Owner picked the wide margin in chat (2026-07-20).
…os (diegosouzapw#7779) * fix(api): enumerate tiered auto combo endpoints in /api/combos/auto The backend already supports auto/<category>[:<tier>] routing via suffixComposition.ts + virtualFactory.ts, but GET /api/combos/auto only exposed 6 flat variants. This adds a second loop enumerating the 10 curated AUTO_SUFFIX_VARIANTS (auto/coding:free, auto/coding:cheap, auto/coding:pro, auto/reasoning, auto/vision, etc.). Fixes diegosouzapw#7619 * fix(combos): enumerate template and family auto variants in GET /api/combos/auto The endpoint was missing 27 auto variants that /v1/models already advertises, causing 404s when clients tried to use them: - 20 template variants (auto/best-coding, auto/pro-*, auto/claude-*, auto/best-free, etc.) - 7 family variants (auto/glm, auto/minimax, auto/mimo, auto/zai, auto/gemma, auto/llama, auto/gemini) Fixes diegosouzapw#7619 Refs diegosouzapw#6453 * fix(combos): swap Phase B/C ordering to match catalog.ts Template variants (Phase C) now enumerate before suffix variants (Phase B) so that overlapping ids like auto/reasoning and auto/vision use template resolution (variant-based) rather than suffix resolution (category-based), matching the behavior in catalog.ts. * fix(combos): fix comment labels and redundant as const * fix(compression): apply compression combo assignments to routing combos Routing combos (e.g. codex, free-only, or-free) use provider-prefixed model strings like codex/gpt-5.5 and go through handleSingleModelChat, which passes comboName: null, isCombo: false. The compression combo assignment lookup in chatCore.ts was gated behind if (isCombo && comboName), so routing combos never had their compression combos applied. Fix: - Add routingComboId parameter threaded through handleSingleModelChat → executeChatWithBreaker → handleChatCore - In handleChat(), resolve the routing combo UUID from the model string's provider prefix via getComboByName - In chatCore.ts, change the gate to (isCombo && comboName) || routingComboId and add routingComboId to the lookup key array Fixes diegosouzapw#7771 * fix(autoCombo): guarantee positive maxOutputTokens fallback in computeAdvertisedLimits GET /api/combos/auto now enumerates auto/<family> variants (auto/llama, auto/glm, etc). computeAdvertisedLimits() already guaranteed a positive contextLength for any non-empty candidate pool via getTokenLimit()'s fallback chain, but had no equivalent fallback for maxOutputTokens — candidates whose registry entry and models.dev sync data both lack that field (common for no-auth/free-tier providers matching a family filter, e.g. llama-* on groq/bazaarlink/etc) left maxOutputTokens null, which tests/unit/auto-combo-context-advertising.test.ts catches as a contract violation of the endpoint (opencode disables smart auto-compaction when a limit is falsy — the same bug class this module's docstring already describes for contextLength). Fall back to a conservative generic default (4096) when no candidate in the pool resolves a known maxOutputTokens, mirroring the existing contextLength guarantee. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * fix(autoCombo): align advertised max_output_tokens fallback with the catalog convention (8192) Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * chore(quality): file-size baseline for chatHelpers routingComboId thread (876->877) Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: Erick Kinnee <erick@ekinnee.dev> Co-authored-by: Diego Rodrigues de Sa e Souza <diegosouza.pw@gmail.com> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
…cognitive 950 Tip was at 2069/2072 and 900/900 (zero slack) after the day's 17 merges; the remaining queue (diegosouzapw#6973, diegosouzapw#7662, diegosouzapw#7719, diegosouzapw#7744, diegosouzapw#7779 reworks) was collectively blocked. Owner picked the wide margin in chat (2026-07-20).
…os (diegosouzapw#7779) * fix(api): enumerate tiered auto combo endpoints in /api/combos/auto The backend already supports auto/<category>[:<tier>] routing via suffixComposition.ts + virtualFactory.ts, but GET /api/combos/auto only exposed 6 flat variants. This adds a second loop enumerating the 10 curated AUTO_SUFFIX_VARIANTS (auto/coding:free, auto/coding:cheap, auto/coding:pro, auto/reasoning, auto/vision, etc.). Fixes diegosouzapw#7619 * fix(combos): enumerate template and family auto variants in GET /api/combos/auto The endpoint was missing 27 auto variants that /v1/models already advertises, causing 404s when clients tried to use them: - 20 template variants (auto/best-coding, auto/pro-*, auto/claude-*, auto/best-free, etc.) - 7 family variants (auto/glm, auto/minimax, auto/mimo, auto/zai, auto/gemma, auto/llama, auto/gemini) Fixes diegosouzapw#7619 Refs diegosouzapw#6453 * fix(combos): swap Phase B/C ordering to match catalog.ts Template variants (Phase C) now enumerate before suffix variants (Phase B) so that overlapping ids like auto/reasoning and auto/vision use template resolution (variant-based) rather than suffix resolution (category-based), matching the behavior in catalog.ts. * fix(combos): fix comment labels and redundant as const * fix(compression): apply compression combo assignments to routing combos Routing combos (e.g. codex, free-only, or-free) use provider-prefixed model strings like codex/gpt-5.5 and go through handleSingleModelChat, which passes comboName: null, isCombo: false. The compression combo assignment lookup in chatCore.ts was gated behind if (isCombo && comboName), so routing combos never had their compression combos applied. Fix: - Add routingComboId parameter threaded through handleSingleModelChat → executeChatWithBreaker → handleChatCore - In handleChat(), resolve the routing combo UUID from the model string's provider prefix via getComboByName - In chatCore.ts, change the gate to (isCombo && comboName) || routingComboId and add routingComboId to the lookup key array Fixes diegosouzapw#7771 * fix(autoCombo): guarantee positive maxOutputTokens fallback in computeAdvertisedLimits GET /api/combos/auto now enumerates auto/<family> variants (auto/llama, auto/glm, etc). computeAdvertisedLimits() already guaranteed a positive contextLength for any non-empty candidate pool via getTokenLimit()'s fallback chain, but had no equivalent fallback for maxOutputTokens — candidates whose registry entry and models.dev sync data both lack that field (common for no-auth/free-tier providers matching a family filter, e.g. llama-* on groq/bazaarlink/etc) left maxOutputTokens null, which tests/unit/auto-combo-context-advertising.test.ts catches as a contract violation of the endpoint (opencode disables smart auto-compaction when a limit is falsy — the same bug class this module's docstring already describes for contextLength). Fall back to a conservative generic default (4096) when no candidate in the pool resolves a known maxOutputTokens, mirroring the existing contextLength guarantee. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * fix(autoCombo): align advertised max_output_tokens fallback with the catalog convention (8192) Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * chore(quality): file-size baseline for chatHelpers routingComboId thread (876->877) Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: Erick Kinnee <erick@ekinnee.dev> Co-authored-by: Diego Rodrigues de Sa e Souza <diegosouza.pw@gmail.com> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
Summary
Compression combo assignments (output mode, pipeline overrides) are correctly saved in the database but never applied to routing combos like
codex,free-only,or-free.Root Cause
The compression combo assignment lookup in
open-sse/handlers/chatCore.tswas gated behindif (isCombo && comboName). Routing combos use provider-prefixed model strings (e.g.codex/gpt-5.5) and go throughhandleSingleModelChat, which passescomboName: null, isCombo: false— so the gate always blocked the lookup.Additionally, even when the gate was passed, the lookup passed human-readable combo names (like
"codex") togetCompressionComboForRoutingCombo(), but the DB stores assignments keyed by routing combo UUID. The name-based lookup always returned null.Fix
Three changes across three files:
open-sse/handlers/chatCore.ts— AddedroutingComboIdparameter. Changed the gate fromif (isCombo && comboName)toif ((isCombo && comboName) || routingComboId). AddedroutingComboIdto the lookup key array so the DB query uses the UUID directly.src/sse/handlers/chat.ts— InhandleChat(), when no explicit combo is matched, resolve the routing combo UUID from the model string's provider prefix viagetComboByName(). Thread the UUID throughhandleSingleModelChat→executeChatWithBreaker→handleChatCore.src/sse/handlers/chatHelpers.ts— AddedroutingComboIdparameter toexecuteChatWithBreaker()and threaded it to thehandleChatCorecall.Testing
pnpm typecheck:corepassesisCombo && comboNamepath still works identically)Related