Skip to content

feat: adaptive reasoning effort (auto) — gateway-resolved, per-turn pinned, all harnesses - #13448

Merged
diegosouzapw merged 16 commits into
diegosouzapw:release/v3.8.51from
patrykkopycinski:feat/adaptive-effort
Sep 16, 2026
Merged

diegosouzapw merged 16 commits into
diegosouzapw:release/v3.8.51from
patrykkopycinski:feat/adaptive-effort

Conversation

@patrykkopycinski

Copy link
Copy Markdown
Contributor

What

Opt-in adaptive reasoning effort at the gateway: reasoning_effort: "auto" (or an X-OmniRoute-Effort: auto header) is resolved to a concrete level from request-shape signals before the request goes upstream. One implementation covers every harness (Claude Code, Cursor, Codex, opencode, Hermes) with zero client changes.

Why

Harnesses either pin a static effort (wastes tokens on trivial turns, under-thinks heavy ones) or omit it (model-side default). The gateway sees every request body regardless of harness, so it can pick the level per turn.

How — two levers, both opt-in

Resolution lives in open-sse/services/adaptiveEffort.ts:

  • Deterministic data signals, no LLM classifier: last-user-message boundary, user-message size, conversation context tokens (reuses estimateMessageTokens), tool-result volume, conversation depth.
  • Bands: trivial ask → low, heavy turn → high, mid → medium; empty/degenerate body → medium.
  • Stateless per-turn pin: signals are computed from the last user message and everything before it — tool results after the last user message are ignored — so every request in a turn resolves to the same level (no cache-busting effort flips mid-turn) with zero stored state.
  • Explicit client effort always wins; auto never overrides a concrete value.

Tests

  • 9 new unit tests (tests/unit/adaptive-effort.test.ts): parse/isAdaptiveEffort, bands, pin invariance across tool-loop growth, explicit-wins, both levers.
  • Mutation-proofed: breaking the pin, the empty→medium guard, and explicit-wins each fail ≥1 test; restored → green.
  • Related suites green: default-reasoning-effort-6879, vendor-default-thinking-effort, sync-reasoning-supported-efforts-7694 (55/55 total).
  • typecheck:core clean, check:cycles clean, no new eslint errors (chatCore.ts carries 27 pre-existing on base — identical count verified).

Docs

docs/routing/AUTO-COMBO.md: request-controls table row + usage section.

Notes

@diegosouzapw

Copy link
Copy Markdown
Owner

The adaptive-effort logic itself is solid — clean, isolated service
(open-sse/services/adaptiveEffort.ts), stateless per-turn pin, explicit-wins semantics, and
9/9 new unit tests pass at this PR's head. However, the PR's actual diff against the release
tip bundles 15 commits, and only the adaptive-effort one is described here — the rest
duplicate your still-open #12742 (min output budget floor) and #12723 (Kimi tool-call
narration) plus several unrelated Cursor executor fixes and ops/m1max/Dockerfile.local
changes. As a result the .env.example diff actually reverts documentation clarifications
already on the release tip (the branch is stale relative to several merged changes, not purely
additive). Could you rebase this PR to contain only the adaptive-effort commit(s), so it can be
reviewed and merged independently of #12742/#12723? Also worth a joint look with #13556 since
both resolve reasoning effort in the same chatCore.ts area (they look complementary by
priority order, but should be verified together).

…faultReasoningEffort:auto

Gateway-side counterpart of hermes-agent#109044: resolve the thinking budget
per user turn from deterministic request-shape signals, covering every harness
(Claude Code, Cursor, Codex, opencode, Hermes) in one place, pre-translation.

- open-sse/services/adaptiveEffort.ts: pure resolver + stateless per-turn pin
  (signals scoped to the last user message; post-boundary tool traffic ignored,
  so mid-tool-loop requests resolve identically without stored state)
- chatCore.ts: wired beside applyDefaultReasoningEffort (diegosouzapw#6879) — a literal
  reasoning_effort:'auto' injected by the diegosouzapw#6879 default is treated as an opt-in
  marker and resolved, never forwarded verbatim; explicit client reasoning
  fields always win; X-OmniRoute-Effort header opt-in mirrors the diegosouzapw#6023-25
  request-controls pattern
- thresholds mirror the Hermes resolver (three coarse bands; near-miss costs a
  slightly over/under-thought answer, not a wrong route)
- docs/routing/AUTO-COMBO.md: X-OmniRoute-Effort row in the request-controls table

Gates: typecheck:core clean; 9/9 new + 46/46 related suites; cycles OK; lint
parity with base (27 pre-existing chatCore errors, 0 new); mutation-proofed
(pin/empty-band/explicit-wins all bite).
@patrykkopycinski

patrykkopycinski commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor Author

Rebuilt the branch surgically onto release/v3.8.51 tip (c3945a724). Old head 89321f587 is preserved at backup/adaptive-effort-pre-rebuild for reference; new head is e307f29e5.

Now 1 commit, containing only the adaptive-effort change (cherry-picked verbatim from the old branch, original message kept):

  • feat(reasoning): adaptive effort via X-OmniRoute-Effort header and defaultReasoningEffort:auto

Everything else has been dropped from this branch — those live in their own PRs:

  • #12742 min-output-budget commit (0738ad5f5)
  • #12723 Kimi narration-recovery commits (df6b3602d + related cursor executor fixes)
  • ops/m1max rung-4 cleanup commits
  • Dockerfile.local / Dockerfile.local.dockerignore build-context commits
  • unrelated cursor executor fixes (inline resume, frame ceiling, conversation key, etc.)

New diff --stat vs release tip:

changelog.d/features/13448-adaptive-reasoning-effort.md |   1 +
docs/routing/AUTO-COMBO.md                               |  11 +-
open-sse/handlers/chatCore.ts                            |  35 ++++++
open-sse/services/adaptiveEffort.ts                      | 137 +++++++++++++++++++++
tests/unit/adaptive-effort.test.ts                       |  89 +++++++++++++
5 files changed, 268 insertions(+), 5 deletions(-)

.env.example regression: resolved. git diff origin/release/v3.8.51..HEAD -- .env.example is now empty — this feature doesn't touch .env.example at all (header-based opt-in, X-OmniRoute-Effort, documented in docs/routing/AUTO-COMBO.md instead). The doc-clarification reverts you flagged came entirely from the cursor-fix commits that are no longer on this branch.

Tests, exact commands + counts:

DISABLE_SQLITE_AUTO_BACKUP=true node --import tsx/esm --import ./open-sse/utils/setupPolyfill.ts \
  --import ./tests/_setup/isolateDataDir.ts --test tests/unit/adaptive-effort.test.ts
# tests 9, pass 9, fail 0

DISABLE_SQLITE_AUTO_BACKUP=true node --import tsx/esm --import ./open-sse/utils/setupPolyfill.ts \
  --import ./tests/_setup/isolateDataDir.ts --test \
  tests/unit/chatcore-claude-effort-variant.test.ts tests/unit/chatcore-translation-paths.test.ts \
  tests/unit/chatcore-upstream-body.test.ts tests/unit/chatCore-reasoning-cache-guard.test.ts \
  tests/unit/chatcore-imports-cleanly.test.ts
# tests 113, pass 113, fail 0

DISABLE_SQLITE_AUTO_BACKUP=true node --import tsx/esm --import ./open-sse/utils/setupPolyfill.ts \
  --import ./tests/_setup/isolateDataDir.ts --test tests/unit/default-reasoning-effort-6879.test.ts
# tests 13, pass 13, fail 0 (the #6879 default-effort suite this wiring sits beside)

9/9 on the adaptive-effort suite, matching what was observed before. 126 additional related chatCore/effort-wiring tests green, 0 regressions.

Changelog fragment: none existed on the old head for #13448 (verified via ls changelog.d/{features,fixes}/ | grep 13448, git grep 13448 -- changelog.d/ on the old head — no hits). Added changelog.d/features/13448-adaptive-reasoning-effort.md. node scripts/check/check-changelog-integrity.mjs passes for this fragment (the one integrity failure it reports, changelog.d/fixes/reset-aware-model-family.md, is a pre-existing fragment already broken on the release/v3.8.51 tip, unrelated to this PR).

Interplay with #13556 (fix(routing): preserve reasoning overrides across transports and fallbacks): complementary, no conflict.

  • fix(routing): preserve reasoning overrides across transports and fallbacks #13556 adds a server-side forced rule pipeline: _omnirouteReasoningRule → applyReasoningRuleDirective() (src/lib/reasoningRouting/policy.ts) sets body.reasoning_effort / body.reasoning.effort / body.output_config.effort directly in chatCore.ts (new block right after handleChatCore starts, open-sse/handlers/chatCore.ts ~L1217-1233), and separately threads a getForcedReasoningEffort(credentials) symbol through resolveExecutionCredentialsFor via withReasoningRuleContext so the Codex executor (open-sse/executors/codex.ts L822-836, L1399-1403) can force reasoning.effort even on native/passthrough paths that skip chatCore's translated-body branch entirely.
  • This PR's adaptive-effort resolver runs later, guarded by hasExplicitReasoningField(translatedBody) (chatCore.ts L2738, our new block right after applyDefaultReasoningEffort at L2723). Since applyReasoningRuleDirective in fix(routing): preserve reasoning overrides across transports and fallbacks #13556 already sets reasoning_effort/reasoning.effort on body/translatedBody before our block runs, hasExplicitReasoningField sees a non-undefined reasoning_effort and short-circuits — our resolver never overwrites a forced rule. Different keys aren't actually in play; it's priority-by-execution-order, and fix(routing): preserve reasoning overrides across transports and fallbacks #13556's directive always lands first in the pipeline (attached at request-entry / retained across translation), so a forced rule always wins over auto, consistent with the intended "explicit client reasoning fields always win" contract this PR already documents.
  • The one shared file is open-sse/handlers/chatCore.ts — different regions (their directive-application block near the top of handleChatCore + native-passthrough branches around L2286-2439; our two additions are the header-read near L1164 and the hasExplicitReasoningField gate near L2738). A straightforward textual merge; no logical clash. fix(routing): preserve reasoning overrides across transports and fallbacks #13556's own description also confirms #13554/#13555 are the ones it expects file-overlap friction with, not this PR.

No code changes needed on this branch for the #13556 interplay — flagging it here as requested for reviewer visibility.

@patrykkopycinski

Copy link
Copy Markdown
Contributor Author

CI attribution (cross-posted to #13355 #13448 #13617 #13627 #13359): the red Fast Quality Gates on these heads are pre-existing base debt, visible on the base branch's own runs with no PR diff applied:

  • mutation-test-coverage: run 34970965604 (base head 1e1af28) fails "2 covering unit test(s) across 2 module(s) are missing from stryker.conf.json" (incl. tests/unit/noauth-model-lockout.test.ts) — same signature on our PR runs.
  • secrets ratchet: "REGRESSÃO — 1 secret findings > baseline 0" reproduces across PRs with disjoint diffs.
  • forgotten-sibling-tests: "Invalid string length" (RangeError, string overflow) crashes the checker on the agyModels→shared→tests graph — hit 13355 and 13627 hours apart on unrelated diffs.
    Base tip FQG itself: last completed run 34970965604 = failure. Happy to PR the stryker testFiles additions if useful.

The file-size gate freezes chatCore.ts (cap 5984) and the inline wiring pushed
it to 6006. Move the call-site adapter to open-sse/handlers/chatCore/
adaptiveEffortWiring.ts, matching the in-flight chatCore/ decomposition on the
release line; chatCore.ts keeps only an import and the call.

Extraction exposed a real bug in the inline version: hasExplicitReasoningField()
returns true for the literal "auto" that applyDefaultReasoningEffort injects from
ModelSpec.defaultReasoningEffort, so the early return fired and the model-default
path was dead code -- "auto" would ship upstream verbatim instead of resolving to
a concrete level. The marker is now checked before the explicit-field guard.

Tests: 8 new wiring cases (precedence, marker resolution, header opt-in, identity
on no opt-in). 30/30 across the three effort suites. Mutation-verified: reverting
the guard reorder turns 3 of 8 red.
@patrykkopycinski

Copy link
Copy Markdown
Contributor Author

Pushed 8867c1ca8 — fixes the one gate failure that was genuinely this PR's.

file-size gate. chatCore.ts is frozen at 5984 lines; the inline wiring took it to 6006. Extracted the call-site adapter to open-sse/handlers/chatCore/adaptiveEffortWiring.ts, matching the in-flight chatCore/ decomposition already on the release line. chatCore.ts now carries only an import plus the call, and node scripts/check/check-file-size.mjs --base-ref c3945a724 reports OK (prettier clean too — the one-line-call variant that also passed the cap was reformatted back over it, so the comment block moved into the module instead).

The extraction exposed a real bug, which is the part worth your attention. hasExplicitReasoningField() treats the literal "auto" as an explicit client field, but applyDefaultReasoningEffort injects exactly that string when ModelSpec.defaultReasoningEffort opts in. The inline guard ran the explicit-field check first, so the model-default branch below it was dead code and "auto" would have gone upstream verbatim instead of resolving to a concrete level. The header path masked it — that is the one the tests covered. Fixed by checking the marker before the explicit-field guard.

Tests. 8 new cases in tests/unit/adaptive-effort-wiring.test.ts (precedence, marker resolution, header opt-in, identity when nothing opts in, missing rawBody). 30/30 across adaptive-effort, default-reasoning-effort-6879 and the wiring suite. Mutation-verified: reverting the guard reorder turns 3 of 8 red — and the pre-fix code fails those same 3, which is how the bug surfaced.

Remaining red on this head (Fast Quality Gates mutation-test-coverage/secrets, unit shards, forgotten-sibling-tests) is base debt, not this diff — evidence in my other comment: the base branch's own run 34970965604 fails the identical mutation-test-coverage gate with zero PR diff.

…faultReasoningEffort:auto

Gateway-side counterpart of hermes-agent#109044: resolve the thinking budget
per user turn from deterministic request-shape signals, covering every harness
(Claude Code, Cursor, Codex, opencode, Hermes) in one place, pre-translation.

- open-sse/services/adaptiveEffort.ts: pure resolver + stateless per-turn pin
  (signals scoped to the last user message; post-boundary tool traffic ignored,
  so mid-tool-loop requests resolve identically without stored state)
- chatCore.ts: wired beside applyDefaultReasoningEffort (diegosouzapw#6879) — a literal
  reasoning_effort:'auto' injected by the diegosouzapw#6879 default is treated as an opt-in
  marker and resolved, never forwarded verbatim; explicit client reasoning
  fields always win; X-OmniRoute-Effort header opt-in mirrors the diegosouzapw#6023-25
  request-controls pattern
- thresholds mirror the Hermes resolver (three coarse bands; near-miss costs a
  slightly over/under-thought answer, not a wrong route)
- docs/routing/AUTO-COMBO.md: X-OmniRoute-Effort row in the request-controls table

Gates: typecheck:core clean; 9/9 new + 46/46 related suites; cycles OK; lint
parity with base (27 pre-existing chatCore errors, 0 new); mutation-proofed
(pin/empty-band/explicit-wins all bite).
The file-size gate freezes chatCore.ts (cap 5984) and the inline wiring pushed
it to 6006. Move the call-site adapter to open-sse/handlers/chatCore/
adaptiveEffortWiring.ts, matching the in-flight chatCore/ decomposition on the
release line; chatCore.ts keeps only an import and the call.

Extraction exposed a real bug in the inline version: hasExplicitReasoningField()
returns true for the literal "auto" that applyDefaultReasoningEffort injects from
ModelSpec.defaultReasoningEffort, so the early return fired and the model-default
path was dead code -- "auto" would ship upstream verbatim instead of resolving to
a concrete level. The marker is now checked before the explicit-field guard.

Tests: 8 new wiring cases (precedence, marker resolution, header opt-in, identity
on no opt-in). 30/30 across the three effort suites. Mutation-verified: reverting
the guard reorder turns 3 of 8 red.
chatCore.ts is frozen at 6146 lines; rebasing onto the new release tip put
the branch's wiring call-site 2 lines over the cap. Move the header read into
wireAdaptiveEffort (reusing the shared getHeaderValueCaseInsensitive helper
from ./headers.ts for exact trim/parity) and collapse the call-site to one
line, so the frozen file stays at 6144 <= 6146.

Behavior unchanged: header lookup, trim, and empty-string rejection now go
through the same helper chatCore used directly. Tests 17/17 (8 wiring + 9
effort).
patrykkopycinski and others added 4 commits September 16, 2026 18:59
…faultReasoningEffort:auto

Gateway-side counterpart of hermes-agent#109044: resolve the thinking budget
per user turn from deterministic request-shape signals, covering every harness
(Claude Code, Cursor, Codex, opencode, Hermes) in one place, pre-translation.

- open-sse/services/adaptiveEffort.ts: pure resolver + stateless per-turn pin
  (signals scoped to the last user message; post-boundary tool traffic ignored,
  so mid-tool-loop requests resolve identically without stored state)
- chatCore.ts: wired beside applyDefaultReasoningEffort (diegosouzapw#6879) — a literal
  reasoning_effort:'auto' injected by the diegosouzapw#6879 default is treated as an opt-in
  marker and resolved, never forwarded verbatim; explicit client reasoning
  fields always win; X-OmniRoute-Effort header opt-in mirrors the diegosouzapw#6023-25
  request-controls pattern
- thresholds mirror the Hermes resolver (three coarse bands; near-miss costs a
  slightly over/under-thought answer, not a wrong route)
- docs/routing/AUTO-COMBO.md: X-OmniRoute-Effort row in the request-controls table

Gates: typecheck:core clean; 9/9 new + 46/46 related suites; cycles OK; lint
parity with base (27 pre-existing chatCore errors, 0 new); mutation-proofed
(pin/empty-band/explicit-wins all bite).
The file-size gate freezes chatCore.ts (cap 5984) and the inline wiring pushed
it to 6006. Move the call-site adapter to open-sse/handlers/chatCore/
adaptiveEffortWiring.ts, matching the in-flight chatCore/ decomposition on the
release line; chatCore.ts keeps only an import and the call.

Extraction exposed a real bug in the inline version: hasExplicitReasoningField()
returns true for the literal "auto" that applyDefaultReasoningEffort injects from
ModelSpec.defaultReasoningEffort, so the early return fired and the model-default
path was dead code -- "auto" would ship upstream verbatim instead of resolving to
a concrete level. The marker is now checked before the explicit-field guard.

Tests: 8 new wiring cases (precedence, marker resolution, header opt-in, identity
on no opt-in). 30/30 across the three effort suites. Mutation-verified: reverting
the guard reorder turns 3 of 8 red.
chatCore.ts is frozen at 6146 lines; rebasing onto the new release tip put
the branch's wiring call-site 2 lines over the cap. Move the header read into
wireAdaptiveEffort (reusing the shared getHeaderValueCaseInsensitive helper
from ./headers.ts for exact trim/parity) and collapse the call-site to one
line, so the frozen file stays at 6144 <= 6146.

Behavior unchanged: header lookup, trim, and empty-string rejection now go
through the same helper chatCore used directly. Tests 17/17 (8 wiring + 9
effort).
…aintainer rework

# Conflicts:
#	changelog.d/features/13448-adaptive-reasoning-effort.md
#	docs/routing/AUTO-COMBO.md
#	open-sse/handlers/chatCore.ts
#	open-sse/handlers/chatCore/adaptiveEffortWiring.ts
#	tests/unit/adaptive-effort-wiring.test.ts
…oping, auto in ModelSpec) onto the contributor's rebased head

# Conflicts:
#	changelog.d/features/13448-adaptive-reasoning-effort.md
#	docs/routing/AUTO-COMBO.md
#	open-sse/handlers/chatCore.ts
#	open-sse/handlers/chatCore/adaptiveEffortWiring.ts
#	tests/unit/adaptive-effort-wiring.test.ts
hartmark and others added 3 commits September 16, 2026 15:09
…ount (diegosouzapw#13079)

Merged after a maintainer rework that kept every one of @hartmark's commits intact.

**What the rework added:** the reclaimable-space gate for the auto-cleanup VACUUM sits behind a default-off feature flag so the release default is unchanged, with the flag documented in `docs/reference/FEATURE_FLAGS.md` and described in all 66 locales; the rest is your change as submitted.

Validated as a combined board first (this PR merged with the 21 siblings of the same wave on the release tip): eslint with the frozen suppressions, typecheck:core, check:open-sse-typecheck, complexity, cognitive-complexity, changelog-integrity, i18n new-key coverage, docs-counts, docs-sync, migration-numbering, provider-consistency and a duplicate-identifier audit all green, plus 176 passing / 0 failing focused node:test cases across the 25 test files the wave touches and the dashboard test under Vitest (2/0). Then re-validated alone on the fresh tip before this merge: conflicts re-resolved, file sizes rebaselined for this PR's own growth, eslint and this PR's focused tests re-run.

Thank you — gating VACUUM on reclaimable pages instead of row count is the right signal.
…rmat key (diegosouzapw#13617)

Merged after a maintainer rework that kept every one of @patrykkopycinski's commits intact, including the changelog fragment you added afterwards.

**What the rework added:** the `eslint-suppressions.json` diff was corrected (the PR had dropped live entries) and a test now proves the CLI-probe fallback path is actually taken when the HTTP API rejects a CLI-format key — before, the fallback existed but nothing exercised it.

Validated as a combined board first (this PR merged with the 21 siblings of the same wave on the release tip): eslint with the frozen suppressions, typecheck:core, check:open-sse-typecheck, complexity, cognitive-complexity, changelog-integrity, i18n new-key coverage, docs-counts, docs-sync, migration-numbering, provider-consistency and a duplicate-identifier audit all green, plus 176 passing / 0 failing focused node:test cases across the 25 test files the wave touches and the dashboard test under Vitest (2/0). Then re-validated alone on the fresh tip before this merge: conflicts re-resolved, file sizes rebaselined for this PR's own growth, eslint and this PR's focused tests re-run.

Thank you.
@diegosouzapw
diegosouzapw merged commit 23c5772 into diegosouzapw:release/v3.8.51 Sep 16, 2026
4 of 7 checks passed
@diegosouzapw diegosouzapw mentioned this pull request Sep 21, 2026
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…inned, all harnesses (diegosouzapw#13448)

Merged after a maintainer rework that kept every one of @patrykkopycinski's commits intact — including the two refactors you pushed later (extracting the adaptive-effort wiring out of `chatCore.ts` and reading `x-omniroute-effort` inside the wiring module), which were merged into the rework rather than overwritten.

**What the rework added:** the adaptive-effort wiring is scoped to OpenAI-dispatch requests only (the claim in `docs/routing` was corrected to match), and `defaultReasoningEffort` was widened to accept `auto` explicitly instead of relying on a loose string.

Validated as a combined board first (this PR merged with the 21 siblings of the same wave on the release tip): eslint with the frozen suppressions, typecheck:core, check:open-sse-typecheck, complexity, cognitive-complexity, changelog-integrity, i18n new-key coverage, docs-counts, docs-sync, migration-numbering, provider-consistency and a duplicate-identifier audit all green, plus 176 passing / 0 failing focused node:test cases across the 25 test files the wave touches and the dashboard test under Vitest (2/0). Then re-validated alone on the fresh tip before this merge: conflicts re-resolved, file sizes rebaselined for this PR's own growth, eslint and this PR's focused tests re-run.

Thank you — gateway-resolved, per-turn pinned effort is a real feature, and the header contract makes it usable from every harness.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants