-
Notifications
You must be signed in to change notification settings - Fork 9.1k
feat(effort): Add model-level reasoning effort routing #1780
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Merged
Merged
Changes from all commits
Commits
Show all changes
9 commits
Select commit
Hold shift + click to select a range
e8ab768
Add model-level reasoning effort metadata
jatmn 94cfa5e
Expand reasoning effort routing
jatmn 5387506
Fix Responses API reasoning effort shape
jatmn 39badf1
Document reasoning effort metadata rules
jatmn 4621145
Stabilize effort resolver tests
jatmn f4aef70
Address effort PR review findings
jatmn 2fb5ba6
Fix Z.AI metadata high effort serialization
jatmn 5c6d121
Fix provider override effort fallback
jatmn 78d98f7
Clamp provider override effort by route metadata
jatmn File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,61 @@ | ||
| # Reasoning and /effort Metadata | ||
|
|
||
| OpenClaude treats reasoning support as a per-model capability. Provider and gateway catalogs can contain a mix of reasoning and non-reasoning models, so reasoning controls must never be inferred provider-wide. | ||
|
|
||
| ## Concepts | ||
|
|
||
| `capabilities.supportsReasoning` means the model is known to support reasoning or thinking behavior. It is safe capability metadata, but by itself it does not authorize OpenClaude to mutate API requests. | ||
|
|
||
| `reasoning` describes the request control surface OpenClaude can safely use for that exact model entry or model descriptor. | ||
|
|
||
| ```ts | ||
| reasoning: { | ||
| mode: 'levels' | 'toggle' | 'always-on' | ||
| // Any supported subset for this exact model, for example ['high', 'xhigh']. | ||
| levels?: ReasoningEffortLevel[] | ||
| defaultLevel?: 'low' | 'medium' | 'high' | 'xhigh' | 'max' | ||
| wireFormat?: | ||
| | 'reasoning_effort' | ||
| | 'deepseek_compatible' | ||
| | 'zai_compatible' | ||
| | 'none' | ||
| disableFormat?: 'thinking_type_disabled' | ||
| } | ||
| ``` | ||
|
|
||
| ## Backward Compatibility | ||
|
|
||
| The `/effort` resolver is intentionally conservative: | ||
|
|
||
| 1. Explicit per-model `reasoning` metadata wins. | ||
| 2. Existing hardcoded legacy effort support remains unchanged. | ||
| 3. `supportsReasoning: true` without `reasoning` metadata is treated as reasoning-capable but not controllable. | ||
| 4. Truly unknown models do not receive new reasoning request fields. | ||
|
|
||
| This means existing OpenAI, Codex, Claude, Gemini, and configured 3P override behavior remains active, while catalogs can safely mark models with `supportsReasoning` before their exact request shape has been audited. | ||
|
|
||
| A temporary compatibility layer also preserves verified request shaping that existed before per-model `reasoning` metadata. For example, DeepSeek-compatible routes can still map `/effort xhigh` to provider `reasoning_effort: "max"`, and Z.AI GLM routes can still map supported controls through their `thinking` request shape. Those compatibility rules also cover matching uncataloged DeepSeek/Z.AI route traffic, so the unknown-model rule only applies after explicit metadata and compatibility resolution both fail. These rules are intentionally centralized in the effort resolver so they can be removed as catalogs gain explicit `reasoning` metadata. | ||
|
|
||
| ## Provider and Gateway Rules | ||
|
|
||
| Annotate reasoning per exact model on the route where it was verified. Aggregating gateways must not add reasoning controls at the provider level because different upstream models accept different parameters and levels. | ||
|
|
||
| Prefer catalog-entry metadata when a gateway route differs from the canonical model descriptor. For example, a model may support reasoning directly from its vendor but reject `reasoning_effort` through a gateway. | ||
|
|
||
| Use `mode: 'always-on'` with `wireFormat: 'none'` for models that emit reasoning but do not have a verified control parameter on that route. | ||
|
|
||
| Currently wired metadata formats are `reasoning_effort`, `deepseek_compatible`, and `zai_compatible`. The descriptor type also reserves `reasoning_object` and `thinking_type`, but those formats are not request-plumbed yet and should not be used to enable `/effort`. | ||
|
|
||
| For `deepseek_compatible` and `zai_compatible`, metadata levels must be limited to `high` and/or `xhigh`. These serializers emit provider `high` for `high` and provider `max` for `xhigh`; they cannot faithfully represent `low`, `medium`, or standard `max` as distinct UI levels. | ||
|
|
||
| ## Adding Support | ||
|
|
||
| Before adding `reasoning` metadata for a model: | ||
|
|
||
| 1. Probe the exact route and model ID OpenClaude will send. | ||
| 2. Record accepted levels and rejected levels. | ||
| 3. Check whether disabling thinking is supported and what request shape is required. | ||
| 4. Confirm whether accepted parameters actually change behavior or are silent no-ops. | ||
| 5. Add focused tests for the resolver and request serialization path. | ||
|
|
||
| Do not use `supportsReasoning: true` alone as evidence that `reasoning_effort` or any other effort field is accepted. | ||
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
Uh oh!
There was an error while loading. Please reload this page.