Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions changelog.d/fixes/13720-suffix-effort-propagation.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
- fix(chat): preserve suffix-model reasoning effort across model attempts so a replacement model no longer inherits or drops the original suffix, and keep explicit reasoning choices in request dedup hashes (#13720)
4 changes: 3 additions & 1 deletion config/quality/file-size-baseline.json
Original file line number Diff line number Diff line change
Expand Up @@ -357,8 +357,10 @@
"_rebaseline_2026_07_27_3850_relax_filesize_cap": "OWNER-APPROVED TEMPORARY relax for v3.8.50-3.8.54 PREPARE phase (docs/ROADMAP.md). cap 800->900 (+100), testCap 800->900 (+100). Targets: decompose-existing-frozen unchanged (frozen still only-shrink); this only relaxes the cap for NEW files in the decompose/extract-while-PREPARE phase (.51='executor registry in-place' and .52='combo.ts decomposition' create new leaf modules above 800). RE-TIGHTENING MANDATORY in v3.8.51: cap target 850 = 850 once decomposition wave stabilizes. SUPERSEDED by _rebaseline_2026_07_27_3850_relax_filesize_cap_v2_20pct (v1 +20% buffer) — retained for audit. Tracked via same roadmap issue.",
"_rebaseline_2026_07_27_v3849_train1h": "Merge-train 1H (31 PRs) — owner-approved 2026-07-27. Two distinct causes, kept separate on purpose: (1) GENUINE irreducible growth at existing chokepoints — providerLimits/auth (#8632 Kimi quota-reset recovery), rateLimitManager (#8616 idle wedged limiters), models-catalog-route.test (#8610 OpenCode Go effort aliases); (2) COLLISION with #8585, which banked shrinks measured on the pre-train release tip while 30 sibling PRs in the SAME train grew those files again — chat/accountFallback (#8628), chatCore (#8613), videoGeneration (#8581), imageGeneration. The zero-headroom frozen entries cannot absorb either. Ceilings re-pinned to the post-merge tip; #8612 (also in this train) automates shrink-banking so this self-inflicted drift stops recurring. Detail: src/lib/usage/providerLimits.ts 1006->1013 (#8632); src/sse/services/auth.ts 2492->2508 (#8632); open-sse/services/rateLimitManager.ts 1014->1060 (#8616); src/sse/handlers/chat.ts 1842->1845 (#8628); open-sse/handlers/chatCore.ts 4939->4955 (#8613); open-sse/handlers/imageGeneration.ts 3100->3101 ((sem PR — teto do #8585)); open-sse/handlers/videoGeneration.ts 1038->1063 (#8581); open-sse/services/accountFallback.ts 1965->1966 (#8628); tests/unit/models-catalog-route.test.ts 1608->1636 (#8610)",
"frozen": {
"src/sse/handlers/chatHelpers.ts": 1253,
"_rebaseline_2026_09_17_13720_merge_release_v3851": "Merge de release/v3.8.51 na #13720 (2026-09-17). src/sse/handlers/chatHelpers.ts 1246 -> 1253, decomposto: 1246 -> 1250 e crescimento INHERITED do tip (base-red ja presente em origin/release/v3.8.51 no commit 9688032451fc, arquivo com 1250 linhas contra cap 1246 — nao e desta PR e nao foi introduzido por este merge); 1250 -> 1253 sao as MESMAS +3 linhas da propria #13720 ja auditadas e aprovadas pelo dono na entrada _rebaseline_2026_09_16_13720_suffix_effort_propagation abaixo (threading de resolvedThinkingEffort). Nenhum outro teto foi tocado por este merge; tests/unit/chatcore-translation-paths.test.ts (3449 > 3447) permanece vermelho de proposito — e base-red herdado e a PR nao toca o arquivo.",
"_rebaseline_2026_09_16_13720_suffix_effort_propagation": "OWNER-APPROVED 2026-09-16 (explicit exception for this unit only, chatHelpers.ts only). PR #13720 (HouMinXi, suffix-effort propagation across model attempts): src/sse/handlers/chatHelpers.ts merge-base (before PR's own commit) was 1164; the release tip independently grew it to 1213 (+49, unrelated merged PRs) while the frozen cap sat at 1214 to cover exactly that tip growth. The PR's own diff on this file is +3 lines only (threading resolvedThinkingEffort: one field on resolveModelOrError's return object, one destructured param and one passthrough call-site argument in executeChatWithBreaker — see commit 4fdb0c5851b7f645efbe5627cd0babb0c3d230c3), taking the merged result to 1216 (1217 per check-file-size.mjs's countLines, which counts the trailing newline as an extra split segment). All 3 added lines are single-property additions inside existing multi-line object literals/signatures; there is no redundant or duplicated line in the PR's own hunks to trim, and none of the +3 lines are outside the PR's own diff. Covered by tests/unit/suffix-effort-propagation.test.ts (27/27), tests/unit/chatcore-upstream-body.test.ts + tests/unit/request-dedup-tenant-isolation.test.ts (57/57), all green against this exact head.",
"_rebaseline_2026_09_17_13947_tip_growth": "Base-red drain da PR #13947 (Refs #13866) — crescimento de PRODUCAO que chegou pelo tip e nunca foi rebaselinado; nenhum destes arquivos e tocado por esta PR. #12906 (d70f43d4, retry empty_response 502 + timeout de inicio de resposta ciente de reasoning): src/sse/handlers/chat.ts 2498->2500, src/sse/handlers/chatHelpers.ts 1231->1245, open-sse/utils/proxyFetch.ts 1275->1276, open-sse/utils/stream.ts 3098->3123. #12904 (f3acf4f8, injecao unica do system prompt global pos-traducao) + #12910 (051576fd, finalizacao de cache semantico por request id exato): open-sse/handlers/chatCore.ts 6181->6203. Anteriores ao lote, ja acima do cap na base 3d5baf13: open-sse/handlers/imageGeneration.ts 3293->3304 (#13748, b97338a8) e open-sse/services/combo/roundRobinCombo.ts 1213->1221 (#13776, aeba6b1a). Registrado contra o estado mergeado; nenhum outro cap e tocado.",
"src/sse/handlers/chatHelpers.ts": 1246,
"_rebaseline_2026_09_15_13609_mistral_ambiguous_401": "PR #13609 rework (maxmad64bis, bare Mistral 401 soft lockout behind MISTRAL_AMBIGUOUS_401_SOFT_LOCKOUT, default off). open-sse/services/accountFallback.ts 2469->2501 (+32): +14 are the change itself (shared-predicate + flag imports, the documented ambiguousAuth field on the checkFallbackError return type, and the flag-gated 401 branch formatted normally instead of the PR's 139-char squeezed configuredRule line); +18 are the lint-staged prettier pass normalizing lines that were already unformatted on the release tip (multi-import, ISO_RETRY_RE, two regex arrays, persistAntigravityFamilyCooldownIfQuota call, applyErrorState guard, trailing commas) — pure formatting, no logic. src/sse/services/auth.ts 3556->3557 (+1): markAccountUnavailable passes connectionId to resolveTerminalConnectionStatus so the soft-strike bound is per connection. The predicate and strike tracker live in the leaf open-sse/services/accountFallback/mistralAmbiguousAuth.ts (under cap). Covered by tests/unit/provider-401-ambiguous-runtime.test.ts (flag off/on, end-to-end through markAccountUnavailable).",
"_rebaseline_2026_06_22_4644_deepseek_web_tools": "PR #4644 (BugsBag/robust deepseek-web tool-call parsing): open-sse/executors/deepseek-web.ts 1117->1125 (+8). The new agentic tool-call path emits surrounding text + reasoning before tool_calls and swaps to the dedicated deepseekWebTools.ts parser; the +8 lines are cohesive wiring at the existing transformSSE chokepoint (the parser itself lives in the new deepseekWebTools.ts file, already under cap). The PR's own fast-gate (PR->release) does not run check:file-size, so this surfaced only at release reconcile. Covered by tests/unit/deepseek-web-tools-variants.test.ts + deepseek-web-tools-execute.test.ts.",
"_rebaseline_2026_06_23_4712_deepseek_web_tool_results": "PR for #4712 (deepseek-web drops role:tool): open-sse/executors/deepseek-web.ts 1125->1148 (+23). messagesToPrompt() now folds role:\"tool\" results into the single-prompt transcript (recovering the tool name from the preceding assistant tool_calls by tool_call_id) instead of silently dropping them; the lines are cohesive wiring inside the existing function. Covered by tests/unit/deepseek-web-tool-result-prompt-4712.test.ts.",
Expand Down
163 changes: 34 additions & 129 deletions open-sse/handlers/chatCore.ts
Original file line number Diff line number Diff line change
Expand Up @@ -165,26 +165,9 @@ import {
getStripTypesForProviderModel,
stripIncompatibleMessageContent,
} from "../services/modelStrip.ts";
import { normalizeMimoThinking } from "../services/mimoThinking.ts";
import {
isOpencodeGoProvider,
stripBooleanReasoning,
} from "../services/opencodeReasoningSanitizer.ts";
import {
normalizeClaudeAdaptiveThinking,
normalizeClaudeDisabledThinkingEffort,
} from "../services/claudeAdaptiveThinking.ts";
import { shouldUseMidConversationSystem } from "../executors/claudeIdentity.ts";
import { normalizeClaudeHaikuConstraints } from "../services/claudeHaikuConstraints.ts";
import { applyDefaultReasoningEffort } from "../services/defaultReasoningEffort.ts";
import { wireAdaptiveEffort } from "./chatCore/adaptiveEffortWiring.ts";
import { echoModelInObject } from "../services/responseModelEcho.ts";
import {
stripGpt5SamplingWhenReasoning,
stripGpt5ReasoningWhenTools,
} from "../services/gpt5SamplingGuard.ts";
import { getUnsupportedParams, REGISTRY } from "../config/providerRegistry.ts";
import { stripUnsupportedParams } from "./chatCore/unsupportedParamsStrip.ts";
import { checkToolCallingRequiredButUnsupported } from "./chatCore/toolCallingRequiredCheck.ts";
import {
supportsMaxTokens,
Expand All @@ -210,7 +193,6 @@ import {
isTinyBudgetReasoningProbe,
toPositiveInteger,
} from "../services/reasoningTokenBuffer.ts";
import { normalizeThinkingForModel } from "@/shared/constants/modelSpecs.ts";
import {
buildErrorBody,
createErrorResult,
Expand Down Expand Up @@ -512,6 +494,19 @@ export async function handleChatCore({
videoBridgeLog = undefined,
fallbackAttempts = undefined,
}) {
const {
model: originModel,
resolvedThinkingEffort,
defaultThinkingEffort,
} = modelInfo as typeof modelInfo & {
resolvedThinkingEffort?: string | null;
defaultThinkingEffort?: string | null;
};
const trustedEffortContext = Object.freeze({
originModel,
resolvedThinkingEffort,
defaultThinkingEffort,
});
let { provider, model, extendedContext } = modelInfo;
// Keep the selected rule across format conversion, retries and refreshed credentials.
// Each combo leg gets its own execution context; nothing is written to shared accounts.
Expand Down Expand Up @@ -2734,77 +2729,6 @@ export async function handleChatCore({
}
translatedBody.model = finalModelToUpstream;

// #3554: a combo/route may substitute the upstream model AFTER the client chose its
// `thinking` value. Claude Code sends `thinking:{type:"disabled"}` for internal calls,
// which claude-fable-5 (adaptive-only) rejects with a 400. Drop the now-invalid value
// when the resolved target model rejects it; models that accept `disabled` are untouched.
if (typeof finalModelToUpstream === "string") {
translatedBody = normalizeThinkingForModel(translatedBody, finalModelToUpstream);
// Claude Opus 4.7+/Fable 5 removed manual extended thinking: `thinking.type:"enabled"`
// or any `thinking.budget_tokens` is a hard 400. Collapse any manual thinking that
// reached this point (passthrough legacy shape, reasoning_effort buckets, per-model
// defaults) to `{type:"adaptive"}` — effort stays on `output_config.effort`. Keyed on
// the resolved upstream model, so it covers every routing mode. See claudeAdaptiveThinking.ts.
translatedBody = normalizeClaudeAdaptiveThinking(translatedBody, finalModelToUpstream);
// Opus 5 allows disabled thinking only through high effort on Anthropic's direct
// Messages API. The helper scopes this constraint to `anthropic` and `claude`;
// GitHub Copilot and Claude Web use separate upstream contracts.
translatedBody = normalizeClaudeDisabledThinkingEffort(
translatedBody,
finalModelToUpstream,
provider
);
// Claude Haiku rejects `thinking.type:"adaptive"` and `output_config.effort`
// (both Sonnet 4.6 / Opus 4.5+ only). Several paths can still emit those
// shapes on a Haiku target — native passthrough, reasoning_effort buckets,
// per-model defaults — so collapse them to a Haiku-valid shape here, after
// model substitution. Mirrors upstream 9router 401d93bd5. See
// services/claudeHaikuConstraints.ts.
translatedBody = normalizeClaudeHaikuConstraints(translatedBody, finalModelToUpstream);
// #6879: per-model default reasoning_effort, injected only when the request
// carries no reasoning field of any shape — an explicit client/combo-leg value
// always wins. Scoped to the OpenAI Chat Completions dispatch shape (the shape
// `reasoning_effort` is native to); unset ModelSpec.defaultReasoningEffort is a
// no-op. #7694: `modelInfo.resolvedThinkingEffort` — set when the request's model
// id carried a `<prefix>/<model>-{effort}` synced-model alias suffix
// (`src/sse/services/model.ts`) — takes priority over the static per-model default.
// The synced catalog's vendor-declared `defaultThinkingEffort` (OpenRouter
// `reasoning.default_effort`, captured by `detectDefaultThinkingEffort`) is the
// lowest-priority default: it only fires when neither the suffix alias nor a
// static operator default exists. See open-sse/services/defaultReasoningEffort.ts.
if (targetFormat === FORMATS.OPENAI) {
translatedBody = applyDefaultReasoningEffort(
translatedBody,
finalModelToUpstream,
(modelInfo as { resolvedThinkingEffort?: string })?.resolvedThinkingEffort,
(modelInfo as { defaultThinkingEffort?: string })?.defaultThinkingEffort
);
}
translatedBody = wireAdaptiveEffort(translatedBody, {
rawBody: body,
clientRawRequest,
targetFormat,
});
}

// Xiaomi MiMo controls reasoning ONLY via `thinking:{type:"enabled"|"disabled"}` and
// rejects unknown/extra params with a strict "400 Param Incorrect". Map OmniRoute's
// OpenAI reasoning signals onto that native shape: reduce any thinking object to
// `{type}` and drop `reasoning_effort`/`reasoning`. See services/mimoThinking.ts.
if (provider === "xiaomi-mimo") {
translatedBody = normalizeMimoThinking(translatedBody);
}

// opencode-go backed providers (ollama-cloud, opencode-go, opencode,
// opencode-zen) use a Go ChatCompletionRequest struct where `reasoning`
// is typed as openai.Reasoning (a structured type). A boolean
// `reasoning: true/false` — valid per the OpenAI API — causes a 400
// "json: cannot unmarshal bool into Go struct field" on the Go side.
// Strip the boolean before forwarding. See opencodeReasoningSanitizer.ts.
if (isOpencodeGoProvider(provider)) {
translatedBody = stripBooleanReasoning(translatedBody);
}

const previousResponseIdPolicy = applyResponsesPreviousResponseIdPolicy(translatedBody, {
mode: settings.responsesPreviousResponseIdMode,
provider,
Expand Down Expand Up @@ -2862,43 +2786,6 @@ export async function handleChatCore({
return createErrorResult(400, toolCallingCheck.message!, null, "tool_calling_not_supported");
}

if (unsupported.length > 0) {
const { strippedParams } = stripUnsupportedParams(translatedBody, unsupported);
if (strippedParams.length > 0) {
log?.warn?.(
"PARAMS",
`Stripped unsupported params for ${model}: ${strippedParams.join(", ")}`
);
}
}

// GPT-5 reasoning models (openai Chat Completions) reject temperature/top_p with a 400
// whenever a reasoning effort is active, yet accept them under reasoning_effort=none (the
// GPT-5.1+ default). A static unsupportedParams list can't express that, so strip sampling
// conditionally here. The codex Responses path is already covered by the executor allowlist.
translatedBody = stripGpt5SamplingWhenReasoning(
translatedBody,
provider,
finalModelToUpstream,
log
);

// GPT-5.x reasoning models on the raw openai Chat Completions surface reject function
// `tools` combined with an active `reasoning_effort`: HTTP 400 "Function tools with
// reasoning_effort are not supported ... Please use /v1/responses instead." This used to
// be true for every GPT-5.x model on the plain `openai` provider, but #7242 (targetFormat
// "openai-responses" on GPT_5_6_API_CAPABILITIES) now routes the GPT-5.6 family to
// /v1/responses instead, which accepts tools + reasoning natively — so the strip must not
// fire there. Pass the already-resolved `targetFormat` so the guard gates on the actual
// upstream surface for this request instead of a model-name list. Port of 9router#2540.
translatedBody = stripGpt5ReasoningWhenTools(
translatedBody,
provider,
finalModelToUpstream,
targetFormat,
log
);

// Rename max_tokens to max_completion_tokens if not supported (#1961)
if (!supportsMaxTokens({ provider, model })) {
if (translatedBody.max_tokens !== undefined) {
Expand Down Expand Up @@ -3111,7 +2998,9 @@ export async function handleChatCore({
// Namespaced by the calling API key: dedup hands the SAME response object to
// every joiner, so a shared hash across keys is a cross-principal response
// leak (GHSA-6c7w-56xp-wpc6).
const dedupHash = dedupEnabled ? computeRequestHash(dedupRequestBody, apiKeyInfo?.id) : null;
const dedupHash = dedupEnabled
? computeRequestHash(dedupRequestBody, apiKeyInfo?.id, trustedEffortContext)
: null;

const executeProviderRequest = async (modelToCall = effectiveModel, allowDedup = false) => {
const execute = async () => {
Expand All @@ -3121,12 +3010,15 @@ export async function handleChatCore({
let bodyToSend = await prepareUpstreamBody({
translatedBody,
modelToCall,
...trustedEffortContext,
provider,
targetFormat,
credentials,
credentials: getExecutionCredentials(),
log,
bypassDefaultToolLimit: isOpencodeClient,
isOpencodeClient,
rawBody: body,
clientRawRequest,
});

// Global System Prompt — SINGLE injection point (post-translation) for
Expand Down Expand Up @@ -4524,12 +4416,25 @@ export async function handleChatCore({
// stay aligned if this block ever runs after a path that mutates body.model (e.g. fallback).
try {
const retryModelId = String(translatedBody.model || effectiveModel);
const retryBody = await prepareUpstreamBody({
translatedBody,
modelToCall: retryModelId,
...trustedEffortContext,
provider,
targetFormat,
credentials: getExecutionCredentials(),
log,
bypassDefaultToolLimit: isOpencodeClient,
isOpencodeClient,
rawBody: body,
clientRawRequest,
});
assertManagedLeaseFence(getExecutionConnectionId(getExecutionCredentials()));
const retryResult = normalizeExecutorResult(
await runWithCapture(providerRequestCapture, () =>
executor.execute({
model: retryModelId,
body: translatedBody,
body: retryBody,
stream: upstreamStream,
credentials: getExecutionCredentials(),
signal: streamController.signal,
Expand Down
Loading
Loading