Repository navigation
feat: adaptive reasoning effort (auto) — gateway-resolved, per-turn pinned, all harnesses - #13448
Conversation
2d34a51 to
89321f5
Compare
|
The adaptive-effort logic itself is solid — clean, isolated service |
…faultReasoningEffort:auto Gateway-side counterpart of hermes-agent#109044: resolve the thinking budget per user turn from deterministic request-shape signals, covering every harness (Claude Code, Cursor, Codex, opencode, Hermes) in one place, pre-translation. - open-sse/services/adaptiveEffort.ts: pure resolver + stateless per-turn pin (signals scoped to the last user message; post-boundary tool traffic ignored, so mid-tool-loop requests resolve identically without stored state) - chatCore.ts: wired beside applyDefaultReasoningEffort (diegosouzapw#6879) — a literal reasoning_effort:'auto' injected by the diegosouzapw#6879 default is treated as an opt-in marker and resolved, never forwarded verbatim; explicit client reasoning fields always win; X-OmniRoute-Effort header opt-in mirrors the diegosouzapw#6023-25 request-controls pattern - thresholds mirror the Hermes resolver (three coarse bands; near-miss costs a slightly over/under-thought answer, not a wrong route) - docs/routing/AUTO-COMBO.md: X-OmniRoute-Effort row in the request-controls table Gates: typecheck:core clean; 9/9 new + 46/46 related suites; cycles OK; lint parity with base (27 pre-existing chatCore errors, 0 new); mutation-proofed (pin/empty-band/explicit-wins all bite).
89321f5 to
e307f29
Compare
|
Rebuilt the branch surgically onto Now 1 commit, containing only the adaptive-effort change (cherry-picked verbatim from the old branch, original message kept):
Everything else has been dropped from this branch — those live in their own PRs:
New diff --stat vs release tip:
Tests, exact commands + counts: 9/9 on the adaptive-effort suite, matching what was observed before. 126 additional related chatCore/effort-wiring tests green, 0 regressions. Changelog fragment: none existed on the old head for Interplay with #13556 (
No code changes needed on this branch for the #13556 interplay — flagging it here as requested for reviewer visibility. |
|
CI attribution (cross-posted to #13355 #13448 #13617 #13627 #13359): the red Fast Quality Gates on these heads are pre-existing base debt, visible on the base branch's own runs with no PR diff applied:
|
The file-size gate freezes chatCore.ts (cap 5984) and the inline wiring pushed it to 6006. Move the call-site adapter to open-sse/handlers/chatCore/ adaptiveEffortWiring.ts, matching the in-flight chatCore/ decomposition on the release line; chatCore.ts keeps only an import and the call. Extraction exposed a real bug in the inline version: hasExplicitReasoningField() returns true for the literal "auto" that applyDefaultReasoningEffort injects from ModelSpec.defaultReasoningEffort, so the early return fired and the model-default path was dead code -- "auto" would ship upstream verbatim instead of resolving to a concrete level. The marker is now checked before the explicit-field guard. Tests: 8 new wiring cases (precedence, marker resolution, header opt-in, identity on no opt-in). 30/30 across the three effort suites. Mutation-verified: reverting the guard reorder turns 3 of 8 red.
|
Pushed file-size gate. The extraction exposed a real bug, which is the part worth your attention. Tests. 8 new cases in Remaining red on this head ( |
8867c1c to
f3fb5eb
Compare
…ultReasoningEffort to allow auto
f3fb5eb to
0e6c495
Compare
…faultReasoningEffort:auto Gateway-side counterpart of hermes-agent#109044: resolve the thinking budget per user turn from deterministic request-shape signals, covering every harness (Claude Code, Cursor, Codex, opencode, Hermes) in one place, pre-translation. - open-sse/services/adaptiveEffort.ts: pure resolver + stateless per-turn pin (signals scoped to the last user message; post-boundary tool traffic ignored, so mid-tool-loop requests resolve identically without stored state) - chatCore.ts: wired beside applyDefaultReasoningEffort (diegosouzapw#6879) — a literal reasoning_effort:'auto' injected by the diegosouzapw#6879 default is treated as an opt-in marker and resolved, never forwarded verbatim; explicit client reasoning fields always win; X-OmniRoute-Effort header opt-in mirrors the diegosouzapw#6023-25 request-controls pattern - thresholds mirror the Hermes resolver (three coarse bands; near-miss costs a slightly over/under-thought answer, not a wrong route) - docs/routing/AUTO-COMBO.md: X-OmniRoute-Effort row in the request-controls table Gates: typecheck:core clean; 9/9 new + 46/46 related suites; cycles OK; lint parity with base (27 pre-existing chatCore errors, 0 new); mutation-proofed (pin/empty-band/explicit-wins all bite).
The file-size gate freezes chatCore.ts (cap 5984) and the inline wiring pushed it to 6006. Move the call-site adapter to open-sse/handlers/chatCore/ adaptiveEffortWiring.ts, matching the in-flight chatCore/ decomposition on the release line; chatCore.ts keeps only an import and the call. Extraction exposed a real bug in the inline version: hasExplicitReasoningField() returns true for the literal "auto" that applyDefaultReasoningEffort injects from ModelSpec.defaultReasoningEffort, so the early return fired and the model-default path was dead code -- "auto" would ship upstream verbatim instead of resolving to a concrete level. The marker is now checked before the explicit-field guard. Tests: 8 new wiring cases (precedence, marker resolution, header opt-in, identity on no opt-in). 30/30 across the three effort suites. Mutation-verified: reverting the guard reorder turns 3 of 8 red.
chatCore.ts is frozen at 6146 lines; rebasing onto the new release tip put the branch's wiring call-site 2 lines over the cap. Move the header read into wireAdaptiveEffort (reusing the shared getHeaderValueCaseInsensitive helper from ./headers.ts for exact trim/parity) and collapse the call-site to one line, so the frozen file stays at 6144 <= 6146. Behavior unchanged: header lookup, trim, and empty-string rejection now go through the same helper chatCore used directly. Tests 17/17 (8 wiring + 9 effort).
0e6c495 to
447b4fc
Compare
…faultReasoningEffort:auto Gateway-side counterpart of hermes-agent#109044: resolve the thinking budget per user turn from deterministic request-shape signals, covering every harness (Claude Code, Cursor, Codex, opencode, Hermes) in one place, pre-translation. - open-sse/services/adaptiveEffort.ts: pure resolver + stateless per-turn pin (signals scoped to the last user message; post-boundary tool traffic ignored, so mid-tool-loop requests resolve identically without stored state) - chatCore.ts: wired beside applyDefaultReasoningEffort (diegosouzapw#6879) — a literal reasoning_effort:'auto' injected by the diegosouzapw#6879 default is treated as an opt-in marker and resolved, never forwarded verbatim; explicit client reasoning fields always win; X-OmniRoute-Effort header opt-in mirrors the diegosouzapw#6023-25 request-controls pattern - thresholds mirror the Hermes resolver (three coarse bands; near-miss costs a slightly over/under-thought answer, not a wrong route) - docs/routing/AUTO-COMBO.md: X-OmniRoute-Effort row in the request-controls table Gates: typecheck:core clean; 9/9 new + 46/46 related suites; cycles OK; lint parity with base (27 pre-existing chatCore errors, 0 new); mutation-proofed (pin/empty-band/explicit-wins all bite).
The file-size gate freezes chatCore.ts (cap 5984) and the inline wiring pushed it to 6006. Move the call-site adapter to open-sse/handlers/chatCore/ adaptiveEffortWiring.ts, matching the in-flight chatCore/ decomposition on the release line; chatCore.ts keeps only an import and the call. Extraction exposed a real bug in the inline version: hasExplicitReasoningField() returns true for the literal "auto" that applyDefaultReasoningEffort injects from ModelSpec.defaultReasoningEffort, so the early return fired and the model-default path was dead code -- "auto" would ship upstream verbatim instead of resolving to a concrete level. The marker is now checked before the explicit-field guard. Tests: 8 new wiring cases (precedence, marker resolution, header opt-in, identity on no opt-in). 30/30 across the three effort suites. Mutation-verified: reverting the guard reorder turns 3 of 8 red.
chatCore.ts is frozen at 6146 lines; rebasing onto the new release tip put the branch's wiring call-site 2 lines over the cap. Move the header read into wireAdaptiveEffort (reusing the shared getHeaderValueCaseInsensitive helper from ./headers.ts for exact trim/parity) and collapse the call-site to one line, so the frozen file stays at 6144 <= 6146. Behavior unchanged: header lookup, trim, and empty-string rejection now go through the same helper chatCore used directly. Tests 17/17 (8 wiring + 9 effort).
…aintainer rework # Conflicts: # changelog.d/features/13448-adaptive-reasoning-effort.md # docs/routing/AUTO-COMBO.md # open-sse/handlers/chatCore.ts # open-sse/handlers/chatCore/adaptiveEffortWiring.ts # tests/unit/adaptive-effort-wiring.test.ts
447b4fc to
2152564
Compare
…oping, auto in ModelSpec) onto the contributor's rebased head # Conflicts: # changelog.d/features/13448-adaptive-reasoning-effort.md # docs/routing/AUTO-COMBO.md # open-sse/handlers/chatCore.ts # open-sse/handlers/chatCore/adaptiveEffortWiring.ts # tests/unit/adaptive-effort-wiring.test.ts
…ount (diegosouzapw#13079) Merged after a maintainer rework that kept every one of @hartmark's commits intact. **What the rework added:** the reclaimable-space gate for the auto-cleanup VACUUM sits behind a default-off feature flag so the release default is unchanged, with the flag documented in `docs/reference/FEATURE_FLAGS.md` and described in all 66 locales; the rest is your change as submitted. Validated as a combined board first (this PR merged with the 21 siblings of the same wave on the release tip): eslint with the frozen suppressions, typecheck:core, check:open-sse-typecheck, complexity, cognitive-complexity, changelog-integrity, i18n new-key coverage, docs-counts, docs-sync, migration-numbering, provider-consistency and a duplicate-identifier audit all green, plus 176 passing / 0 failing focused node:test cases across the 25 test files the wave touches and the dashboard test under Vitest (2/0). Then re-validated alone on the fresh tip before this merge: conflicts re-resolved, file sizes rebaselined for this PR's own growth, eslint and this PR's focused tests re-run. Thank you — gating VACUUM on reclaimable pages instead of row count is the right signal.
…rmat key (diegosouzapw#13617) Merged after a maintainer rework that kept every one of @patrykkopycinski's commits intact, including the changelog fragment you added afterwards. **What the rework added:** the `eslint-suppressions.json` diff was corrected (the PR had dropped live entries) and a test now proves the CLI-probe fallback path is actually taken when the HTTP API rejects a CLI-format key — before, the fallback existed but nothing exercised it. Validated as a combined board first (this PR merged with the 21 siblings of the same wave on the release tip): eslint with the frozen suppressions, typecheck:core, check:open-sse-typecheck, complexity, cognitive-complexity, changelog-integrity, i18n new-key coverage, docs-counts, docs-sync, migration-numbering, provider-consistency and a duplicate-identifier audit all green, plus 176 passing / 0 failing focused node:test cases across the 25 test files the wave touches and the dashboard test under Vitest (2/0). Then re-validated alone on the fresh tip before this merge: conflicts re-resolved, file sizes rebaselined for this PR's own growth, eslint and this PR's focused tests re-run. Thank you.
23c5772
into
diegosouzapw:release/v3.8.51
…inned, all harnesses (diegosouzapw#13448) Merged after a maintainer rework that kept every one of @patrykkopycinski's commits intact — including the two refactors you pushed later (extracting the adaptive-effort wiring out of `chatCore.ts` and reading `x-omniroute-effort` inside the wiring module), which were merged into the rework rather than overwritten. **What the rework added:** the adaptive-effort wiring is scoped to OpenAI-dispatch requests only (the claim in `docs/routing` was corrected to match), and `defaultReasoningEffort` was widened to accept `auto` explicitly instead of relying on a loose string. Validated as a combined board first (this PR merged with the 21 siblings of the same wave on the release tip): eslint with the frozen suppressions, typecheck:core, check:open-sse-typecheck, complexity, cognitive-complexity, changelog-integrity, i18n new-key coverage, docs-counts, docs-sync, migration-numbering, provider-consistency and a duplicate-identifier audit all green, plus 176 passing / 0 failing focused node:test cases across the 25 test files the wave touches and the dashboard test under Vitest (2/0). Then re-validated alone on the fresh tip before this merge: conflicts re-resolved, file sizes rebaselined for this PR's own growth, eslint and this PR's focused tests re-run. Thank you — gateway-resolved, per-turn pinned effort is a real feature, and the header contract makes it usable from every harness.
What
Opt-in adaptive reasoning effort at the gateway:
reasoning_effort: "auto"(or anX-OmniRoute-Effort: autoheader) is resolved to a concrete level from request-shape signals before the request goes upstream. One implementation covers every harness (Claude Code, Cursor, Codex, opencode, Hermes) with zero client changes.Why
Harnesses either pin a static effort (wastes tokens on trivial turns, under-thinks heavy ones) or omit it (model-side default). The gateway sees every request body regardless of harness, so it can pick the level per turn.
How — two levers, both opt-in
ModelSpec.defaultReasoningEffort: "auto"— extends the feat(providers): per-model default reasoning_effort; make no-think/ express "none" on the OpenAI path #6879 default-injection seam. A literal"auto"left by the default injection is treated as an opt-in marker and resolved, never forwarded verbatim.X-OmniRoute-Effort: autorequest header (feat(sse): per-request Auto-Combo controls (X-OmniRoute-Mode / X-OmniRoute-Budget) — closes #6023 #6024 #6025 #6057 request-controls pattern) for per-request control without body changes.Resolution lives in
open-sse/services/adaptiveEffort.ts:estimateMessageTokens), tool-result volume, conversation depth.low, heavy turn →high, mid →medium; empty/degenerate body →medium.autonever overrides a concrete value.Tests
tests/unit/adaptive-effort.test.ts): parse/isAdaptiveEffort, bands, pin invariance across tool-loop growth, explicit-wins, both levers.default-reasoning-effort-6879,vendor-default-thinking-effort,sync-reasoning-supported-efforts-7694(55/55 total).typecheck:coreclean,check:cyclesclean, no new eslint errors (chatCore.ts carries 27 pre-existing on base — identical count verified).Docs
docs/routing/AUTO-COMBO.md: request-controls table row + usage section.Notes
check:docs-allenv-doc-sync drift is pre-existing on base; this diff touches no env/docs-sync surface.