Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 12 additions & 0 deletions docs/integrations/how-to/add-gateway.md
Original file line number Diff line number Diff line change
Expand Up @@ -55,6 +55,18 @@ Normal gateway examples should:

The routing decision belongs to `transportConfig.kind`, not to `category`.

## Reasoning controls in mixed catalogs

Gateway catalogs often mix models with different reasoning APIs. Keep
`capabilities.supportsReasoning` as descriptive capability metadata unless
the exact gateway route and model ID have been probed.

Add `/effort`-controllable `reasoning` metadata per catalog entry, not at the
gateway level. If a gateway accepts one upstream model's `reasoning_effort`
but rejects another model's field, each entry must say so explicitly. See
`docs/integrations/reasoning-effort.md` before adding or changing reasoning
controls.

## Generated loader and preset manifest

Normal gateway onboarding is additive now:
Expand Down
11 changes: 11 additions & 0 deletions docs/integrations/how-to/add-model.md
Original file line number Diff line number Diff line change
Expand Up @@ -65,6 +65,17 @@ editing multiple shared model files. In the common path:
one route;
- let the route catalog continue to own the offered subset.

## Reasoning and `/effort` metadata

Use `classification: ['reasoning']` and `capabilities.supportsReasoning` to
describe that a model is known to reason or think. Those fields do not by
themselves enable `/effort` request mutation.

Only add `reasoning` metadata when the model's exact control surface has been
verified for the route that will call it, including accepted levels, rejected
levels, and any thinking-disable format. For the full checklist, see
`docs/integrations/reasoning-effort.md`.

## When to add a brand descriptor

Add or update a brand descriptor when:
Expand Down
12 changes: 12 additions & 0 deletions docs/integrations/how-to/add-vendor.md
Original file line number Diff line number Diff line change
Expand Up @@ -251,6 +251,18 @@ context windows, output limits, and cross-route capability metadata in
`src/integrations/models/`, then point catalog entries at those descriptors
with `modelDescriptorId`.

## Reasoning controls

For direct vendors, record reasoning controls on the exact catalog model
entry or shared model descriptor only after the vendor API has been probed.
`capabilities.supportsReasoning` means the model can reason; it does not
mean `/effort` should send `reasoning_effort` or any other control field.

If the vendor catalog contains both controllable and non-controllable
reasoning models, annotate each model separately. See
`docs/integrations/reasoning-effort.md` for the metadata shape and audit
checklist.

## OpenAI-compatible UI capability flags

For OpenAI-compatible vendors, be explicit about the provider editor surface:
Expand Down
17 changes: 14 additions & 3 deletions docs/integrations/overview.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,6 +24,7 @@ docs/
integrations/
overview.md
glossary.md
reasoning-effort.md
how-to/
add-vendor.md
add-gateway.md
Expand All @@ -43,9 +44,11 @@ If you are onboarding to the integration system:

1. Read `docs/architecture/integrations.md` for the system boundaries.
2. Read `docs/integrations/glossary.md` for the shared vocabulary.
3. Use the how-to guides for the specific descriptor type you are adding.
4. Use `docs/integrations/reference-samples.md` once the architecture and the relevant how-to guide are clear.
5. Read `docs/integrations/common-pitfalls.md` before opening a docs or implementation PR for a new integration.
3. Read `docs/integrations/reasoning-effort.md` before marking models as
reasoning-capable or `/effort`-controllable.
4. Use the how-to guides for the specific descriptor type you are adding.
5. Use `docs/integrations/reference-samples.md` once the architecture and the relevant how-to guide are clear.
6. Read `docs/integrations/common-pitfalls.md` before opening a docs or implementation PR for a new integration.

## Core Rules

Expand Down Expand Up @@ -115,6 +118,14 @@ routes. Fixed direct vendors usually set both to `false`; broad custom routes
or gateways that intentionally accept user-supplied auth/header details set the
relevant flag to `true`.

### Reasoning support is per model and per route

`capabilities.supportsReasoning` is descriptive. It says the model is known to
reason or think, but it does not by itself authorize `/effort` to add request
fields. Only add `reasoning` metadata when the exact route/model request shape,
accepted levels, and disable behavior have been verified. See
`docs/integrations/reasoning-effort.md`.

## Descriptor Authoring Pattern

Normal descriptor files should:
Expand Down
61 changes: 61 additions & 0 deletions docs/integrations/reasoning-effort.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,61 @@
# Reasoning and /effort Metadata

OpenClaude treats reasoning support as a per-model capability. Provider and gateway catalogs can contain a mix of reasoning and non-reasoning models, so reasoning controls must never be inferred provider-wide.

## Concepts

`capabilities.supportsReasoning` means the model is known to support reasoning or thinking behavior. It is safe capability metadata, but by itself it does not authorize OpenClaude to mutate API requests.

`reasoning` describes the request control surface OpenClaude can safely use for that exact model entry or model descriptor.

```ts
reasoning: {
mode: 'levels' | 'toggle' | 'always-on'
// Any supported subset for this exact model, for example ['high', 'xhigh'].
levels?: ReasoningEffortLevel[]
defaultLevel?: 'low' | 'medium' | 'high' | 'xhigh' | 'max'
wireFormat?:
| 'reasoning_effort'
| 'deepseek_compatible'
| 'zai_compatible'
| 'none'
disableFormat?: 'thinking_type_disabled'
}
Comment thread
coderabbitai[bot] marked this conversation as resolved.
```

## Backward Compatibility

The `/effort` resolver is intentionally conservative:

1. Explicit per-model `reasoning` metadata wins.
2. Existing hardcoded legacy effort support remains unchanged.
3. `supportsReasoning: true` without `reasoning` metadata is treated as reasoning-capable but not controllable.
4. Truly unknown models do not receive new reasoning request fields.

This means existing OpenAI, Codex, Claude, Gemini, and configured 3P override behavior remains active, while catalogs can safely mark models with `supportsReasoning` before their exact request shape has been audited.

A temporary compatibility layer also preserves verified request shaping that existed before per-model `reasoning` metadata. For example, DeepSeek-compatible routes can still map `/effort xhigh` to provider `reasoning_effort: "max"`, and Z.AI GLM routes can still map supported controls through their `thinking` request shape. Those compatibility rules also cover matching uncataloged DeepSeek/Z.AI route traffic, so the unknown-model rule only applies after explicit metadata and compatibility resolution both fail. These rules are intentionally centralized in the effort resolver so they can be removed as catalogs gain explicit `reasoning` metadata.

## Provider and Gateway Rules

Annotate reasoning per exact model on the route where it was verified. Aggregating gateways must not add reasoning controls at the provider level because different upstream models accept different parameters and levels.

Prefer catalog-entry metadata when a gateway route differs from the canonical model descriptor. For example, a model may support reasoning directly from its vendor but reject `reasoning_effort` through a gateway.

Use `mode: 'always-on'` with `wireFormat: 'none'` for models that emit reasoning but do not have a verified control parameter on that route.

Currently wired metadata formats are `reasoning_effort`, `deepseek_compatible`, and `zai_compatible`. The descriptor type also reserves `reasoning_object` and `thinking_type`, but those formats are not request-plumbed yet and should not be used to enable `/effort`.

For `deepseek_compatible` and `zai_compatible`, metadata levels must be limited to `high` and/or `xhigh`. These serializers emit provider `high` for `high` and provider `max` for `xhigh`; they cannot faithfully represent `low`, `medium`, or standard `max` as distinct UI levels.

## Adding Support

Before adding `reasoning` metadata for a model:

1. Probe the exact route and model ID OpenClaude will send.
2. Record accepted levels and rejected levels.
3. Check whether disabling thinking is supported and what request shape is required.
4. Confirm whether accepted parameters actually change behavior or are silent no-ops.
5. Add focused tests for the resolver and request serialization path.

Do not use `supportsReasoning: true` alone as evidence that `reasoning_effort` or any other effort field is accepted.
26 changes: 26 additions & 0 deletions src/integrations/descriptors.ts
Original file line number Diff line number Diff line change
Expand Up @@ -55,6 +55,30 @@ export interface CapabilityFlags {
supportsEmbeddings?: boolean
}

export type ReasoningControlMode = 'levels' | 'toggle' | 'always-on'
export type ReasoningEffortLevel = 'low' | 'medium' | 'high' | 'xhigh' | 'max'
/**
* reasoning_effort, deepseek_compatible, and zai_compatible are wired into
* request serialization today. Other values are reserved until their serializer
* paths are implemented.
*/
export type ReasoningWireFormat =
| 'reasoning_effort'
| 'reasoning_object'
| 'thinking_type'
| 'deepseek_compatible'
| 'zai_compatible'
| 'none'
export type ReasoningDisableFormat = 'thinking_type_disabled'

export interface ReasoningControlMetadata {
mode: ReasoningControlMode
levels?: ReasoningEffortLevel[]
defaultLevel?: ReasoningEffortLevel
wireFormat?: ReasoningWireFormat
disableFormat?: ReasoningDisableFormat
}

export interface TransportConfig {
kind: TransportKind
headers?: Record<string, string>
Expand Down Expand Up @@ -84,6 +108,7 @@ export interface ModelCatalogEntry {
hidden?: boolean
modelDescriptorId?: string
capabilities?: CapabilityFlags
reasoning?: ReasoningControlMetadata
contextWindow?: number
maxOutputTokens?: number
transportOverrides?: CatalogTransportOverrides
Expand Down Expand Up @@ -311,6 +336,7 @@ export interface ModelDescriptor {
defaultModel: string
providerModelMap?: Partial<Record<string, string>>
capabilities: CapabilityFlags
reasoning?: ReasoningControlMetadata
contextWindow?: number
maxOutputTokens?: number
cacheConfig?: CacheConfig
Expand Down
20 changes: 20 additions & 0 deletions src/integrations/runtimeMetadata.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -184,6 +184,26 @@ describe('resolveOpenAIShimRuntimeContext - Z.AI GLM-5.2', () => {
})
})

describe('resolveOpenAIShimRuntimeContext - provider override route preference', () => {
it('does not inherit ambient route config when the preferred base URL is unrecognized', () => {
const result = resolveOpenAIShimRuntimeContext({
model: 'gpt-4o',
baseUrl: 'https://custom.example.test/v1',
preferBaseUrlRoute: true,
processEnv: {
CLAUDE_CODE_USE_OPENAI: '1',
OPENAI_BASE_URL: 'https://api.groq.com/openai/v1',
},
})

expect(result.routeId).toBeNull()
expect(result.descriptor).toBeNull()
expect(result.catalogEntry).toBeNull()
expect(result.openaiShimConfig.removeBodyFields).toBeUndefined()
expect(result.openaiShimConfig.thinkingRequestFormat).toBeUndefined()
})
})

describe('resolveOpenAIShimRuntimeContext - segment-boundary heuristic', () => {
describe('DeepSeek models', () => {
it('should NOT infer preserveReasoningContent for custom aliases (false-positive case)', () => {
Expand Down
9 changes: 6 additions & 3 deletions src/integrations/runtimeMetadata.ts
Original file line number Diff line number Diff line change
Expand Up @@ -220,6 +220,7 @@ export function resolveOpenAIShimRuntimeContext(options?: {
model?: string
activeProfileProvider?: string
treatAsLocal?: boolean
preferBaseUrlRoute?: boolean
}): OpenAIShimRuntimeContext {
const processEnv = options?.processEnv ?? process.env
const runtimeEnv: NodeJS.ProcessEnv = {
Expand All @@ -239,10 +240,12 @@ export function resolveOpenAIShimRuntimeContext(options?: {
})
const baseUrlRouteId = resolveRouteIdFromBaseUrl(options?.baseUrl)
const routeId =
baseUrlRouteId &&
(!activeRouteId || activeRouteId === 'anthropic' || activeRouteId === 'openai')
options?.preferBaseUrlRoute && options.baseUrl !== undefined
? baseUrlRouteId
: activeRouteId
: baseUrlRouteId &&
(!activeRouteId || activeRouteId === 'anthropic' || activeRouteId === 'openai')
? baseUrlRouteId
: activeRouteId
const descriptor =
routeId && routeId !== 'anthropic'
? getRouteDescriptor(routeId)
Expand Down
Loading