Skip to content

feat(chat): enhance reasoning capability for Anthropic - #655

Closed
steebchen wants to merge 6 commits into
mainfrom
fix/reasoning-anthropic
Closed

steebchen wants to merge 6 commits into
mainfrom
fix/reasoning-anthropic

Conversation

@steebchen

@steebchen steebchen commented Aug 28, 2025 •

Copy link
Copy Markdown
Member

Refined handling of Anthropic responses to extract and map reasoning content. Introduced support for streaming and budgeted reasoning tokens. Updated tests and request logic for improved reasoning support.

Summary by CodeRabbit

  • New Features

    • Streaming reasoning responses supported with OpenAI-compatible chunks; reasoning_content exposed in streaming and non-streaming replies; usage now reports prompt, completion, total, and reasoning token counts.
  • Improvements

    • Consistent normalization of reasoning across providers and streaming formats.
    • Better Anthropic support: extracts thinking content, sets finish reasons, and provides a thinking budget to enable longer reasoning.
    • Enforces provider max-token limits with clear errors when exceeded.
  • Tests

    • Expanded streaming reasoning test coverage with SSE, chunk, usage, and reasoning_content validation.

Refined handling of Anthropic responses to extract and map reasoning
content. Introduced support for streaming and budgeted reasoning tokens.
Updated tests and request logic for improved reasoning support.
@coderabbitai

coderabbitai Bot commented Aug 28, 2025 •

Copy link
Copy Markdown
Contributor

Walkthrough

Adds Anthropic "thinking" budgeting to requests, normalizes reasoning content across providers and streaming/non-streaming flows (populating reasoning_content and removing legacy reasoning), enhances parsing of Anthropic streaming thinking deltas, enforces max_tokens vs model maxOutput, and expands e2e tests to cover streaming reasoning (including SSE/OpenAI chunk shape, usage, and logs).

Changes

Cohort / File(s) Summary
E2E tests: reasoning (non-streaming & streaming)
apps/gateway/src/api.e2e.ts
Lowers reasoning_effort to "low" for non-streaming tests; asserts reasoning_content present in non-streaming responses; adds streaming reasoning tests validating text/event-stream, SSE events > 0, OpenAI-format chunk shape (id, object="chat.completion.chunk", model, choices[0].delta with content), usage chunks (prompt/completion/total tokens numeric, optional reasoning_tokens), ensures presence/non-empty reasoning_content in deltas, validates streaming logs indicate streamed=true; includes duplicated streaming test block and extra debug/non-200 guards and console logs when logMode enabled.
Chat pipeline: parsing and normalization
apps/gateway/src/chat/chat.ts
parseProviderResponse now aggregates Anthropic content_blocks into content and thinking blocks into reasoningContent; extractReasoningContentFromProvider handles Anthropic streaming thinking_delta; transformStreamingChunkToOpenAIFormat maps Anthropic thinking_delta → delta.reasoning_content and normalizes/remove legacy reasoning across streaming deltas/choices; transformToOpenAIFormat normalizes reasoning into reasoning_content for non-OpenAI providers when responses have an id; added max_tokens validation against provider/model maxOutput.
Provider request shaping (Anthropic)
packages/models/src/provider-api.ts
Computes a thinkingBudget from reasoning_effort (low=1024, default=2000, high=4000; 0 if no reasoning), sets requestBody.max_tokens = max_tokens ?? Math.max(1024, thinkingBudget + 1000), and injects requestBody.thinking = { type: "enabled", budget_tokens } when model supports reasoning.

Sequence Diagram(s)

sequenceDiagram
  autonumber
  participant C as Client
  participant G as Gateway (chat.ts)
  participant P as Provider (Anthropic)
  participant N as Normalizer

  C->>G: POST /chat (stream=true, reasoning_effort=low)
  G->>P: Send provider request (max_tokens computed, thinking enabled)
  Note right of G #e6f2ff: thinking {type: "enabled", budget_tokens: ...}

  P-->>G: SSE stream (content_block_delta, thinking_delta, text deltas)
  G->>N: transformStreamingChunkToOpenAIFormat
  Note right of N #f7f7f7: Map thinking_delta -> delta.reasoning_content\nAggregate text -> delta.content\nNormalize role/ids/usage

  N-->>C: SSE chunk (object=chat.completion.chunk, choices[0].delta.{content,reasoning_content})
  loop until done
    P-->>G: More deltas
    G->>N: Normalize chunk
    N-->>C: Streamed normalized chunk
  end

  G-->>C: Final usage/event summary (prompt_tokens, completion_tokens, reasoning_tokens?)
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~60 minutes

Possibly related PRs


📜 Recent review details

Configuration used: CodeRabbit UI

Review profile: CHILL

Plan: Pro

💡 Knowledge Base configuration:

  • MCP integration is disabled by default for public repositories
  • Jira integration is disabled by default for public repositories
  • Linear integration is disabled by default for public repositories

You can enable these sources in your CodeRabbit configuration.

📥 Commits

Reviewing files that changed from the base of the PR and between 44c9bd7 and b7d1687.

📒 Files selected for processing (3)
  • apps/gateway/src/api.e2e.ts (3 hunks)
  • apps/gateway/src/chat/chat.ts (9 hunks)
  • packages/models/src/provider-api.ts (2 hunks)
🚧 Files skipped from review as they are similar to previous changes (3)
  • apps/gateway/src/api.e2e.ts
  • apps/gateway/src/chat/chat.ts
  • packages/models/src/provider-api.ts
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (2)
  • GitHub Check: build / run
  • GitHub Check: e2e / run
✨ Finishing Touches
  • 📝 Generate Docstrings
🧪 Generate unit tests
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch fix/reasoning-anthropic

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share
🪧 Tips

Chat

There are 3 ways to chat with CodeRabbit:

  • Review comments: Directly reply to a review comment made by CodeRabbit. Example:
    • I pushed a fix in commit <commit_id>, please review it.
    • Open a follow-up GitHub issue for this discussion.
  • Files and specific lines of code (under the "Files changed" tab): Tag @coderabbit in a new review comment at the desired location with your query.
  • PR comments: Tag @coderabbit in a new PR comment to ask questions about the PR branch. For the best results, please provide a very specific query, as very limited context is provided in this mode. Examples:
    • @coderabbit gather interesting stats about this repository and render them as a table. Additionally, render a pie chart showing the language distribution in the codebase.
    • @coderabbit read the files in the src/scheduler package and generate a class diagram using mermaid and a README in the markdown format.

Support

Need help? Create a ticket on our support page for assistance with any issues or questions.

CodeRabbit Commands (Invoked using PR/Issue comments)

Type @coderabbit help to get the list of available commands.

Other keywords and placeholders

  • Add @coderabbit ignore or @coderabbitai ignore anywhere in the PR description to prevent this PR from being reviewed.
  • Add @coderabbit summary or @coderabbitai summary to generate the high-level summary at a specific location in the PR description.
  • Add @coderabbit or @coderabbitai anywhere in the PR title to generate the title automatically.

CodeRabbit Configuration File (.coderabbit.yaml)

  • You can programmatically configure CodeRabbit by adding a .coderabbit.yaml file to the root of your repository.
  • Please see the configuration documentation for more information.
  • If your editor has YAML language server enabled, you can add the path at the top of this file to enable auto-completion and validation: # yaml-language-server: $schema=https://coderabbit.ai/integrations/schema.v2.json

Status, Documentation and Community

  • Visit our Status Page to check the current availability of CodeRabbit.
  • Visit our Documentation for detailed information on how to use CodeRabbit.
  • Join our Discord Community to get help, request features, and share feedback.
  • Follow us on X/Twitter for updates and announcements.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (2)
apps/gateway/src/chat/chat.ts (2)

3094-3108: Fix tool call merging in streaming: merge by id OR index to avoid duplicating calls

Anthropic emits arguments in content_block_delta without the original id. Merging only by id creates duplicates and splits arguments. Fall back to index when id is missing, and backfill ids.

Apply this diff:

-                for (const newCall of toolCallsChunk) {
-                  const existingCall = streamingToolCalls.find(
-                    (call) => call.id === newCall.id,
-                  );
+                for (const newCall of toolCallsChunk) {
+                  let existingCall =
+                    streamingToolCalls.find(
+                      (call) => call.id && newCall.id && call.id === newCall.id,
+                    ) ||
+                    streamingToolCalls.find(
+                      (call) =>
+                        call.index !== undefined &&
+                        newCall.index !== undefined &&
+                        call.index === newCall.index,
+                    );
                   if (existingCall) {
+                    // Backfill missing id so subsequent merges are stable
+                    if (!newCall.id && existingCall.id) {
+                      newCall.id = existingCall.id;
+                    }
                     // Accumulate function arguments
                     if (newCall.function?.arguments) {
                       existingCall.function.arguments =
                         (existingCall.function.arguments || "") +
                         newCall.function.arguments;
                     }
                   } else {
                     streamingToolCalls.push({ ...newCall });
                   }
                 }

239-279: Validate effective max_tokens after building requestBody (prevents provider errors)

The Anthropic path now sets a default max_tokens based on thinking budget in prepareRequestBody. Your current validation (Lines 2418-2436) only checks the user-supplied max_tokens before calling prepareRequestBody, so defaults aren’t validated and can exceed provider maxOutput.

Add this validation immediately after const requestBody = prepareRequestBody(...) (Line 2471), and remove the earlier block:

// Validate effective max_tokens (including provider defaults) against model's maxOutput
if (finalModelInfo) {
  const providerMapping = finalModelInfo.providers.find(
    (p) => p.providerId === usedProvider && p.modelName === usedModel,
  ) as any;
  const effectiveMax = requestBody?.max_tokens as number | undefined;
  if (
    effectiveMax !== undefined &&
    providerMapping?.maxOutput !== undefined &&
    effectiveMax > providerMapping.maxOutput
  ) {
    throw new HTTPException(400, {
      message: `The requested/effective max_tokens (${effectiveMax}) exceeds the maximum output tokens allowed for model ${usedModel} (${providerMapping.maxOutput})`,
    });
  }
}
🧹 Nitpick comments (2)
packages/models/src/provider-api.ts (1)

180-187: Map reasoning_effort to budget_tokens and confirm header requirements
Hard-coding budget_tokens: 2000 prevents callers from controlling cost—use a lookup (e.g. { low: 500, medium: 2000, high: 4000 }) keyed by reasoning_effort with a 2000-token default. No special anthropic-beta header is required for basic extended thinking on the 2023-06-01 API; only the interleaved-thinking beta (interleaved-thinking-2025-05-14) requires adding its header.

apps/gateway/src/chat/chat.ts (1)

239-279: Anthropic parsing: solid normalization; ensure total_tokens includes reasoning where needed

Collecting text and thinking blocks and mapping tool_use is correct. Consider whether your cost/usage model expects total_tokens to include reasoning tokens for Anthropic; currently it sums input+output only. Tests allow either, but logs/costs might need parity with Google where you add reasoning tokens to totals.

📜 Review details

Configuration used: CodeRabbit UI

Review profile: CHILL

Plan: Pro

💡 Knowledge Base configuration:

  • MCP integration is disabled by default for public repositories
  • Jira integration is disabled by default for public repositories
  • Linear integration is disabled by default for public repositories

You can enable these sources in your CodeRabbit configuration.

📥 Commits

Reviewing files that changed from the base of the PR and between a9a7e02 and da7ebc5.

📒 Files selected for processing (3)
  • apps/gateway/src/api.e2e.ts (2 hunks)
  • apps/gateway/src/chat/chat.ts (8 hunks)
  • packages/models/src/provider-api.ts (2 hunks)
🧰 Additional context used
📓 Path-based instructions (6)
**/*.{js,jsx,ts,tsx}

📄 CodeRabbit inference engine (.github/copilot-instructions.md)

Use localStorage instead of cookies for client-side data persistence

Files:

  • apps/gateway/src/api.e2e.ts
  • packages/models/src/provider-api.ts
  • apps/gateway/src/chat/chat.ts
**/*.{js,ts}

📄 CodeRabbit inference engine (.github/copilot-instructions.md)

**/*.{js,ts}: Use drizzle with the latest object syntax for database operations
For read queries, always use db().query.<table>.findMany() or db().query.<table>.findFirst()

Files:

  • apps/gateway/src/api.e2e.ts
  • packages/models/src/provider-api.ts
  • apps/gateway/src/chat/chat.ts
**/*.{ts,tsx}

📄 CodeRabbit inference engine (.cursor/rules/general.mdc)

Never use as any or : any in TypeScript files.

Files:

  • apps/gateway/src/api.e2e.ts
  • packages/models/src/provider-api.ts
  • apps/gateway/src/chat/chat.ts
apps/{api,gateway}/**/*.{ts,tsx}

📄 CodeRabbit inference engine (CLAUDE.md)

apps/{api,gateway}/**/*.{ts,tsx}: Use Drizzle ORM with the latest object syntax in backend services
For reads, use db().query.

.findMany() or db().query.
.findFirst()

Files:

  • apps/gateway/src/api.e2e.ts
  • apps/gateway/src/chat/chat.ts
**/*.e2e.ts

📄 CodeRabbit inference engine (CLAUDE.md)

Name end-to-end tests with the .e2e.ts suffix

Files:

  • apps/gateway/src/api.e2e.ts
{**/*.spec.ts,**/*.e2e.ts}

📄 CodeRabbit inference engine (CLAUDE.md)

Write tests using Vitest

Files:

  • apps/gateway/src/api.e2e.ts
🧬 Code graph analysis (1)
apps/gateway/src/api.e2e.ts (1)
packages/db/src/schema.ts (1)
  • log (243-305)
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (2)
  • GitHub Check: build / run
  • GitHub Check: e2e / run
🔇 Additional comments (6)
apps/gateway/src/api.e2e.ts (2)

457-458: Good: assert reasoning with low effort in non-streaming path

The change lowers effort to "low" and adds coverage for reasoning responses. Looks correct.


489-492: Good: enforce presence of reasoning_content in non-streaming

Asserting message.reasoning_content ensures normalization works.

apps/gateway/src/chat/chat.ts (4)

514-526: Streaming: handle Anthropic thinking deltas

Good coverage for thinking_delta. Double-check if any providers ever emit initial content_block_start for thinking with an immediate payload; if so, you may want a small handler to catch a first "thinking" value without a delta.


844-852: Normalize reasoning → reasoning_content when upstream already provides an id

Good defensive normalization to keep one field name in responses.


916-938: Streaming: emit reasoning_content for Anthropic thinking deltas

Looks correct and aligns with the e2e assertions for delta.reasoning_content.


1227-1283: Default streaming transformation: ensure role and normalize reasoning

The delta construction and reasoning normalization are sound and match OpenAI chunk shape.

Comment thread apps/gateway/src/api.e2e.ts Outdated
Comment thread packages/models/src/provider-api.ts Outdated
Added validation to ensure the max_tokens value does not exceed the
maximum allowed by the provider mapping for the selected model.
Updated token calculation logic to dynamically adjust thinking budgets
based on reasoning effort levels.
Updated the token calculation for "low" effort levels from 1000 to 1024
to align with Anthropic's minimum token requirements. Added a clarifying
comment.
Updated test cases to use only models with streaming-enabled providers. Added logic to determine streaming availability at both the model and provider levels. Improved filtering to ensure accurate test coverage for streaming use cases.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

♻️ Duplicate comments (3)
packages/models/src/provider-api.ts (1)

121-137: Good switch to budget-aware defaults and preserving explicit 0 via ??

The dynamic budget + cushion and max_tokens ?? minMaxTokens resolve the earlier precedence/0 issue and align with downstream validation.

apps/gateway/src/api.e2e.ts (1)

153-163: Nice — filter reasoning models to only those with streaming support

This removes flakiness by intersecting reasoning with streaming capability.

apps/gateway/src/chat/chat.ts (1)

2472-2489: Validate effective max_tokens after prepareRequestBody — aligned with earlier guidance

Prevents silent overshoot when defaults inflate max_tokens; behavior is explicit (400). Consider clamping if you prefer permissive UX.

-      if (requestBody.max_tokens > providerMapping.maxOutput) {
-        throw new HTTPException(400, {
-          message: `The effective max_tokens (${requestBody.max_tokens}) exceeds the maximum output tokens allowed for model ${usedModel} (${providerMapping.maxOutput})`,
-        });
-      }
+      if (requestBody.max_tokens > providerMapping.maxOutput) {
+        // Alternatively clamp to provider limit for permissive behavior
+        requestBody.max_tokens = providerMapping.maxOutput;
+      }
🧹 Nitpick comments (3)
apps/gateway/src/api.e2e.ts (3)

468-468: Reasoning effort set to "low" for baseline coverage

Reasonable default to reduce cost/latency; consider a follow-up matrix test for “high” on at least one Anthropic/Gemini model.


500-504: Assert type of reasoning_content as well

Add a quick type check to ensure normalized shape.

   // check for reasoning response
   expect(json.choices[0].message).toHaveProperty("reasoning_content");
+  expect(typeof json.choices[0].message.reasoning_content).toBe("string");

505-613: Streaming reasoning test looks solid; add reasoning_content type assertions per-chunk

This tightens guarantees on the normalized streaming delta.

   const reasoningChunks = streamResult.chunks.filter(
     (chunk: any) =>
       chunk.choices?.[0]?.delta?.reasoning_content &&
       chunk.choices[0].delta.reasoning_content.length > 0,
   );
   expect(reasoningChunks.length).toBeGreaterThan(0);
+  for (const rc of reasoningChunks) {
+    expect(typeof rc.choices[0].delta.reasoning_content).toBe("string");
+  }
📜 Review details

Configuration used: CodeRabbit UI

Review profile: CHILL

Plan: Pro

💡 Knowledge Base configuration:

  • MCP integration is disabled by default for public repositories
  • Jira integration is disabled by default for public repositories
  • Linear integration is disabled by default for public repositories

You can enable these sources in your CodeRabbit configuration.

📥 Commits

Reviewing files that changed from the base of the PR and between da7ebc5 and 44c9bd7.

📒 Files selected for processing (3)
  • apps/gateway/src/api.e2e.ts (3 hunks)
  • apps/gateway/src/chat/chat.ts (9 hunks)
  • packages/models/src/provider-api.ts (2 hunks)
🧰 Additional context used
📓 Path-based instructions (6)
**/*.{js,jsx,ts,tsx}

📄 CodeRabbit inference engine (.github/copilot-instructions.md)

Use localStorage instead of cookies for client-side data persistence

Files:

  • packages/models/src/provider-api.ts
  • apps/gateway/src/api.e2e.ts
  • apps/gateway/src/chat/chat.ts
**/*.{js,ts}

📄 CodeRabbit inference engine (.github/copilot-instructions.md)

**/*.{js,ts}: Use drizzle with the latest object syntax for database operations
For read queries, always use db().query.<table>.findMany() or db().query.<table>.findFirst()

Files:

  • packages/models/src/provider-api.ts
  • apps/gateway/src/api.e2e.ts
  • apps/gateway/src/chat/chat.ts
**/*.{ts,tsx}

📄 CodeRabbit inference engine (.cursor/rules/general.mdc)

Never use as any or : any in TypeScript files.

Files:

  • packages/models/src/provider-api.ts
  • apps/gateway/src/api.e2e.ts
  • apps/gateway/src/chat/chat.ts
apps/{api,gateway}/**/*.{ts,tsx}

📄 CodeRabbit inference engine (CLAUDE.md)

apps/{api,gateway}/**/*.{ts,tsx}: Use Drizzle ORM with the latest object syntax in backend services
For reads, use db().query.

.findMany() or db().query.
.findFirst()

Files:

  • apps/gateway/src/api.e2e.ts
  • apps/gateway/src/chat/chat.ts
**/*.e2e.ts

📄 CodeRabbit inference engine (CLAUDE.md)

Name end-to-end tests with the .e2e.ts suffix

Files:

  • apps/gateway/src/api.e2e.ts
{**/*.spec.ts,**/*.e2e.ts}

📄 CodeRabbit inference engine (CLAUDE.md)

Write tests using Vitest

Files:

  • apps/gateway/src/api.e2e.ts
🧬 Code graph analysis (1)
apps/gateway/src/api.e2e.ts (3)
packages/models/src/models.ts (1)
  • ProviderModelMapping (22-77)
packages/models/src/providers.ts (1)
  • providers (19-236)
packages/db/src/schema.ts (1)
  • log (243-305)
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (2)
  • GitHub Check: build / run
  • GitHub Check: e2e / run
🔇 Additional comments (6)
apps/gateway/src/chat/chat.ts (6)

239-278: Anthropic non-streaming parsing correctly aggregates text vs thinking blocks

Joining text blocks into content and “thinking” into reasoningContent matches the new contract and keeps tool_use separate.


766-806: Anthropic normalization to OpenAI format (reasoning_content) — LGTM

Correctly maps finish reasons and usage, and exposes reasoning_content consistently.


843-853: Normalize legacy reasoning to reasoning_content when id already present

Prevents mixed fields in downstream consumers.


916-938: Streaming: map Anthropic thinking_delta to OpenAI delta.reasoning_content

Matches the tests’ expectations; preserves role and usage.


1227-1283: Default streaming transform: role hardening + reasoning→reasoning_content normalization

Good normalization pass; ensures OpenAI-compatible deltas even from heterogeneous providers.


514-526: Anthropic thinking_delta schema confirmed
Code matches Anthropic streaming docs (August 2025): content_block_delta events include delta.type === "thinking_delta" and delta.thinking: string. No changes required.

Comment on lines +193 to +199
// Enable thinking for reasoning-capable Anthropic models
if (supportsReasoning) {
requestBody.thinking = {
type: "enabled",
budget_tokens: thinkingBudget,
};
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Verification agent

🧩 Analysis chain

Anthropic “thinking” payload enabled — verify required beta header/version

Enabling thinking looks right. Please confirm the correct Anthropic beta/version flags that expose thinking/thinking_delta in Messages; otherwise responses may omit reasoning blocks even with thinking set.

Would you like me to wire this to a feature flag while we validate?


🌐 Web query:

What Anthropic headers and versions are required to enable "thinking" (reasoning) blocks and "thinking_delta" SSE events in the Messages API as of August 2025? Provide the exact header names/values (anthropic-version, anthropic-beta) and any docs links.

💡 Result:

Required headers (exact names/values):

  • anthropic-version: 2023-06-01. [3]
  • anthropic-beta: interleaved-thinking-2025-05-14. [2][1]

Notes / docs:

  • Extended/interleaved thinking beta feature (name to pass in anthropic-beta). [1]
  • Beta headers usage (anthropic-beta header). [2]
  • Messages streaming / thinking_delta SSE events sample and example request showing anthropic-version and thinking streaming. [3]

Sources: [1] https://docs.anthropic.com/en/docs/build-with-claude/extended-thinking (docs.anthropic.com); [2] https://docs.anthropic.com/en/api/beta-headers (docs.anthropic.com); [3] https://docs.anthropic.com/en/api/messages-streaming (docs.anthropic.com).

Citations:


Include Anthropic extended-thinking headers on supportsReasoning

  • Add the headers
    • anthropic-version: 2023-06-01
    • anthropic-beta: interleaved-thinking-2025-05-14
    to your Messages API requests when requestBody.thinking is enabled to ensure reasoning blocks and thinking_delta SSE events are exposed.
  • (Optional) Gate this header injection behind a feature flag to control rollout.
🤖 Prompt for AI Agents
In packages/models/src/provider-api.ts around lines 193 to 199, when
supportsReasoning is true and requestBody.thinking is set you must add the
Anthropic extended-thinking headers to the Messages API request: set
"anthropic-version" to "2023-06-01" and "anthropic-beta" to
"interleaved-thinking-2025-05-14" on the outgoing request headers; implement
this where the request headers are constructed for the Messages API (only when
requestBody.thinking is enabled) and optionally wrap the header injection behind
an existing or new feature flag to control rollout.

@steebchen

Copy link
Copy Markdown
Member Author

superseded by #659

@steebchen steebchen closed this Aug 28, 2025
@steebchen
steebchen deleted the fix/reasoning-anthropic branch August 28, 2025 23:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant