Conversation
Refined handling of Anthropic responses to extract and map reasoning content. Introduced support for streaming and budgeted reasoning tokens. Updated tests and request logic for improved reasoning support.
WalkthroughAdds Anthropic "thinking" budgeting to requests, normalizes reasoning content across providers and streaming/non-streaming flows (populating Changes
Sequence Diagram(s)sequenceDiagram
autonumber
participant C as Client
participant G as Gateway (chat.ts)
participant P as Provider (Anthropic)
participant N as Normalizer
C->>G: POST /chat (stream=true, reasoning_effort=low)
G->>P: Send provider request (max_tokens computed, thinking enabled)
Note right of G #e6f2ff: thinking {type: "enabled", budget_tokens: ...}
P-->>G: SSE stream (content_block_delta, thinking_delta, text deltas)
G->>N: transformStreamingChunkToOpenAIFormat
Note right of N #f7f7f7: Map thinking_delta -> delta.reasoning_content\nAggregate text -> delta.content\nNormalize role/ids/usage
N-->>C: SSE chunk (object=chat.completion.chunk, choices[0].delta.{content,reasoning_content})
loop until done
P-->>G: More deltas
G->>N: Normalize chunk
N-->>C: Streamed normalized chunk
end
G-->>C: Final usage/event summary (prompt_tokens, completion_tokens, reasoning_tokens?)
Estimated code review effort🎯 4 (Complex) | ⏱️ ~60 minutes Possibly related PRs
📜 Recent review detailsConfiguration used: CodeRabbit UI Review profile: CHILL Plan: Pro 💡 Knowledge Base configuration:
You can enable these sources in your CodeRabbit configuration. 📒 Files selected for processing (3)
🚧 Files skipped from review as they are similar to previous changes (3)
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (2)
✨ Finishing Touches
🧪 Generate unit tests
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. 🪧 TipsChatThere are 3 ways to chat with CodeRabbit:
SupportNeed help? Create a ticket on our support page for assistance with any issues or questions. CodeRabbit Commands (Invoked using PR/Issue comments)Type Other keywords and placeholders
CodeRabbit Configuration File (
|
There was a problem hiding this comment.
Actionable comments posted: 2
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (2)
apps/gateway/src/chat/chat.ts (2)
3094-3108: Fix tool call merging in streaming: merge by id OR index to avoid duplicating callsAnthropic emits arguments in
content_block_deltawithout the original id. Merging only byidcreates duplicates and splits arguments. Fall back toindexwhenidis missing, and backfill ids.Apply this diff:
- for (const newCall of toolCallsChunk) { - const existingCall = streamingToolCalls.find( - (call) => call.id === newCall.id, - ); + for (const newCall of toolCallsChunk) { + let existingCall = + streamingToolCalls.find( + (call) => call.id && newCall.id && call.id === newCall.id, + ) || + streamingToolCalls.find( + (call) => + call.index !== undefined && + newCall.index !== undefined && + call.index === newCall.index, + ); if (existingCall) { + // Backfill missing id so subsequent merges are stable + if (!newCall.id && existingCall.id) { + newCall.id = existingCall.id; + } // Accumulate function arguments if (newCall.function?.arguments) { existingCall.function.arguments = (existingCall.function.arguments || "") + newCall.function.arguments; } } else { streamingToolCalls.push({ ...newCall }); } }
239-279: Validate effective max_tokens after building requestBody (prevents provider errors)The Anthropic path now sets a default
max_tokensbased on thinking budget inprepareRequestBody. Your current validation (Lines 2418-2436) only checks the user-suppliedmax_tokensbefore callingprepareRequestBody, so defaults aren’t validated and can exceed providermaxOutput.Add this validation immediately after
const requestBody = prepareRequestBody(...)(Line 2471), and remove the earlier block:// Validate effective max_tokens (including provider defaults) against model's maxOutput if (finalModelInfo) { const providerMapping = finalModelInfo.providers.find( (p) => p.providerId === usedProvider && p.modelName === usedModel, ) as any; const effectiveMax = requestBody?.max_tokens as number | undefined; if ( effectiveMax !== undefined && providerMapping?.maxOutput !== undefined && effectiveMax > providerMapping.maxOutput ) { throw new HTTPException(400, { message: `The requested/effective max_tokens (${effectiveMax}) exceeds the maximum output tokens allowed for model ${usedModel} (${providerMapping.maxOutput})`, }); } }
🧹 Nitpick comments (2)
packages/models/src/provider-api.ts (1)
180-187: Mapreasoning_efforttobudget_tokensand confirm header requirements
Hard-codingbudget_tokens: 2000prevents callers from controlling cost—use a lookup (e.g.{ low: 500, medium: 2000, high: 4000 }) keyed byreasoning_effortwith a 2000-token default. No specialanthropic-betaheader is required for basic extended thinking on the 2023-06-01 API; only the interleaved-thinking beta (interleaved-thinking-2025-05-14) requires adding its header.apps/gateway/src/chat/chat.ts (1)
239-279: Anthropic parsing: solid normalization; ensure total_tokens includes reasoning where neededCollecting
textandthinkingblocks and mapping tool_use is correct. Consider whether your cost/usage model expectstotal_tokensto include reasoning tokens for Anthropic; currently it sums input+output only. Tests allow either, but logs/costs might need parity with Google where you add reasoning tokens to totals.
📜 Review details
Configuration used: CodeRabbit UI
Review profile: CHILL
Plan: Pro
💡 Knowledge Base configuration:
- MCP integration is disabled by default for public repositories
- Jira integration is disabled by default for public repositories
- Linear integration is disabled by default for public repositories
You can enable these sources in your CodeRabbit configuration.
📒 Files selected for processing (3)
apps/gateway/src/api.e2e.ts(2 hunks)apps/gateway/src/chat/chat.ts(8 hunks)packages/models/src/provider-api.ts(2 hunks)
🧰 Additional context used
📓 Path-based instructions (6)
**/*.{js,jsx,ts,tsx}
📄 CodeRabbit inference engine (.github/copilot-instructions.md)
Use localStorage instead of cookies for client-side data persistence
Files:
apps/gateway/src/api.e2e.tspackages/models/src/provider-api.tsapps/gateway/src/chat/chat.ts
**/*.{js,ts}
📄 CodeRabbit inference engine (.github/copilot-instructions.md)
**/*.{js,ts}: Use drizzle with the latest object syntax for database operations
For read queries, always usedb().query.<table>.findMany()ordb().query.<table>.findFirst()
Files:
apps/gateway/src/api.e2e.tspackages/models/src/provider-api.tsapps/gateway/src/chat/chat.ts
**/*.{ts,tsx}
📄 CodeRabbit inference engine (.cursor/rules/general.mdc)
Never use
as anyor: anyin TypeScript files.
Files:
apps/gateway/src/api.e2e.tspackages/models/src/provider-api.tsapps/gateway/src/chat/chat.ts
apps/{api,gateway}/**/*.{ts,tsx}
📄 CodeRabbit inference engine (CLAUDE.md)
.findMany() or db().query.
apps/{api,gateway}/**/*.{ts,tsx}: Use Drizzle ORM with the latest object syntax in backend services
For reads, use db().query..findFirst() Files:
apps/gateway/src/api.e2e.tsapps/gateway/src/chat/chat.ts**/*.e2e.ts
📄 CodeRabbit inference engine (CLAUDE.md)
Name end-to-end tests with the .e2e.ts suffix
Files:
apps/gateway/src/api.e2e.ts{**/*.spec.ts,**/*.e2e.ts}
📄 CodeRabbit inference engine (CLAUDE.md)
Write tests using Vitest
Files:
apps/gateway/src/api.e2e.ts🧬 Code graph analysis (1)
apps/gateway/src/api.e2e.ts (1)
packages/db/src/schema.ts (1)
log(243-305)⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (2)
- GitHub Check: build / run
- GitHub Check: e2e / run
🔇 Additional comments (6)
apps/gateway/src/api.e2e.ts (2)
457-458: Good: assert reasoning with low effort in non-streaming pathThe change lowers effort to "low" and adds coverage for reasoning responses. Looks correct.
489-492: Good: enforce presence of reasoning_content in non-streamingAsserting
message.reasoning_contentensures normalization works.apps/gateway/src/chat/chat.ts (4)
514-526: Streaming: handle Anthropic thinking deltasGood coverage for
thinking_delta. Double-check if any providers ever emit initialcontent_block_startfor thinking with an immediate payload; if so, you may want a small handler to catch a first "thinking" value without a delta.
844-852: Normalize reasoning → reasoning_content when upstream already provides an idGood defensive normalization to keep one field name in responses.
916-938: Streaming: emit reasoning_content for Anthropic thinking deltasLooks correct and aligns with the e2e assertions for
delta.reasoning_content.
1227-1283: Default streaming transformation: ensure role and normalize reasoningThe delta construction and reasoning normalization are sound and match OpenAI chunk shape.
Added validation to ensure the max_tokens value does not exceed the maximum allowed by the provider mapping for the selected model. Updated token calculation logic to dynamically adjust thinking budgets based on reasoning effort levels.
Updated the token calculation for "low" effort levels from 1000 to 1024 to align with Anthropic's minimum token requirements. Added a clarifying comment.
Updated test cases to use only models with streaming-enabled providers. Added logic to determine streaming availability at both the model and provider levels. Improved filtering to ensure accurate test coverage for streaming use cases.
There was a problem hiding this comment.
Actionable comments posted: 1
♻️ Duplicate comments (3)
packages/models/src/provider-api.ts (1)
121-137: Good switch to budget-aware defaults and preserving explicit 0 via ??The dynamic budget + cushion and
max_tokens ?? minMaxTokensresolve the earlier precedence/0 issue and align with downstream validation.apps/gateway/src/api.e2e.ts (1)
153-163: Nice — filter reasoning models to only those with streaming supportThis removes flakiness by intersecting reasoning with streaming capability.
apps/gateway/src/chat/chat.ts (1)
2472-2489: Validate effective max_tokens after prepareRequestBody — aligned with earlier guidancePrevents silent overshoot when defaults inflate max_tokens; behavior is explicit (400). Consider clamping if you prefer permissive UX.
- if (requestBody.max_tokens > providerMapping.maxOutput) { - throw new HTTPException(400, { - message: `The effective max_tokens (${requestBody.max_tokens}) exceeds the maximum output tokens allowed for model ${usedModel} (${providerMapping.maxOutput})`, - }); - } + if (requestBody.max_tokens > providerMapping.maxOutput) { + // Alternatively clamp to provider limit for permissive behavior + requestBody.max_tokens = providerMapping.maxOutput; + }
🧹 Nitpick comments (3)
apps/gateway/src/api.e2e.ts (3)
468-468: Reasoning effort set to "low" for baseline coverageReasonable default to reduce cost/latency; consider a follow-up matrix test for “high” on at least one Anthropic/Gemini model.
500-504: Assert type of reasoning_content as wellAdd a quick type check to ensure normalized shape.
// check for reasoning response expect(json.choices[0].message).toHaveProperty("reasoning_content"); + expect(typeof json.choices[0].message.reasoning_content).toBe("string");
505-613: Streaming reasoning test looks solid; add reasoning_content type assertions per-chunkThis tightens guarantees on the normalized streaming delta.
const reasoningChunks = streamResult.chunks.filter( (chunk: any) => chunk.choices?.[0]?.delta?.reasoning_content && chunk.choices[0].delta.reasoning_content.length > 0, ); expect(reasoningChunks.length).toBeGreaterThan(0); + for (const rc of reasoningChunks) { + expect(typeof rc.choices[0].delta.reasoning_content).toBe("string"); + }
📜 Review details
Configuration used: CodeRabbit UI
Review profile: CHILL
Plan: Pro
💡 Knowledge Base configuration:
- MCP integration is disabled by default for public repositories
- Jira integration is disabled by default for public repositories
- Linear integration is disabled by default for public repositories
You can enable these sources in your CodeRabbit configuration.
📒 Files selected for processing (3)
apps/gateway/src/api.e2e.ts(3 hunks)apps/gateway/src/chat/chat.ts(9 hunks)packages/models/src/provider-api.ts(2 hunks)
🧰 Additional context used
📓 Path-based instructions (6)
**/*.{js,jsx,ts,tsx}
📄 CodeRabbit inference engine (.github/copilot-instructions.md)
Use localStorage instead of cookies for client-side data persistence
Files:
packages/models/src/provider-api.tsapps/gateway/src/api.e2e.tsapps/gateway/src/chat/chat.ts
**/*.{js,ts}
📄 CodeRabbit inference engine (.github/copilot-instructions.md)
**/*.{js,ts}: Use drizzle with the latest object syntax for database operations
For read queries, always usedb().query.<table>.findMany()ordb().query.<table>.findFirst()
Files:
packages/models/src/provider-api.tsapps/gateway/src/api.e2e.tsapps/gateway/src/chat/chat.ts
**/*.{ts,tsx}
📄 CodeRabbit inference engine (.cursor/rules/general.mdc)
Never use
as anyor: anyin TypeScript files.
Files:
packages/models/src/provider-api.tsapps/gateway/src/api.e2e.tsapps/gateway/src/chat/chat.ts
apps/{api,gateway}/**/*.{ts,tsx}
📄 CodeRabbit inference engine (CLAUDE.md)
.findMany() or db().query.
apps/{api,gateway}/**/*.{ts,tsx}: Use Drizzle ORM with the latest object syntax in backend services
For reads, use db().query..findFirst() Files:
apps/gateway/src/api.e2e.tsapps/gateway/src/chat/chat.ts**/*.e2e.ts
📄 CodeRabbit inference engine (CLAUDE.md)
Name end-to-end tests with the .e2e.ts suffix
Files:
apps/gateway/src/api.e2e.ts{**/*.spec.ts,**/*.e2e.ts}
📄 CodeRabbit inference engine (CLAUDE.md)
Write tests using Vitest
Files:
apps/gateway/src/api.e2e.ts🧬 Code graph analysis (1)
apps/gateway/src/api.e2e.ts (3)
packages/models/src/models.ts (1)
ProviderModelMapping(22-77)packages/models/src/providers.ts (1)
providers(19-236)packages/db/src/schema.ts (1)
log(243-305)⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (2)
- GitHub Check: build / run
- GitHub Check: e2e / run
🔇 Additional comments (6)
apps/gateway/src/chat/chat.ts (6)
239-278: Anthropic non-streaming parsing correctly aggregates text vs thinking blocksJoining text blocks into content and “thinking” into reasoningContent matches the new contract and keeps tool_use separate.
766-806: Anthropic normalization to OpenAI format (reasoning_content) — LGTMCorrectly maps finish reasons and usage, and exposes reasoning_content consistently.
843-853: Normalize legacyreasoningtoreasoning_contentwhen id already presentPrevents mixed fields in downstream consumers.
916-938: Streaming: map Anthropic thinking_delta to OpenAI delta.reasoning_contentMatches the tests’ expectations; preserves role and usage.
1227-1283: Default streaming transform: role hardening + reasoning→reasoning_content normalizationGood normalization pass; ensures OpenAI-compatible deltas even from heterogeneous providers.
514-526: Anthropic thinking_delta schema confirmed
Code matches Anthropic streaming docs (August 2025):content_block_deltaevents includedelta.type === "thinking_delta"anddelta.thinking: string. No changes required.
| // Enable thinking for reasoning-capable Anthropic models | ||
| if (supportsReasoning) { | ||
| requestBody.thinking = { | ||
| type: "enabled", | ||
| budget_tokens: thinkingBudget, | ||
| }; | ||
| } |
There was a problem hiding this comment.
💡 Verification agent
🧩 Analysis chain
Anthropic “thinking” payload enabled — verify required beta header/version
Enabling thinking looks right. Please confirm the correct Anthropic beta/version flags that expose thinking/thinking_delta in Messages; otherwise responses may omit reasoning blocks even with thinking set.
Would you like me to wire this to a feature flag while we validate?
🌐 Web query:
What Anthropic headers and versions are required to enable "thinking" (reasoning) blocks and "thinking_delta" SSE events in the Messages API as of August 2025? Provide the exact header names/values (anthropic-version, anthropic-beta) and any docs links.
💡 Result:
Required headers (exact names/values):
- anthropic-version: 2023-06-01. [3]
- anthropic-beta: interleaved-thinking-2025-05-14. [2][1]
Notes / docs:
- Extended/interleaved thinking beta feature (name to pass in anthropic-beta). [1]
- Beta headers usage (anthropic-beta header). [2]
- Messages streaming / thinking_delta SSE events sample and example request showing anthropic-version and thinking streaming. [3]
Sources: [1] https://docs.anthropic.com/en/docs/build-with-claude/extended-thinking (docs.anthropic.com); [2] https://docs.anthropic.com/en/api/beta-headers (docs.anthropic.com); [3] https://docs.anthropic.com/en/api/messages-streaming (docs.anthropic.com).
Citations:
- 1: https://docs.anthropic.com/en/docs/build-with-claude/extended-thinking?utm_source=openai
- 2: https://docs.anthropic.com/en/api/beta-headers?utm_source=openai
- 3: https://docs.anthropic.com/en/api/messages-streaming?utm_source=openai
Include Anthropic extended-thinking headers on supportsReasoning
- Add the headers
•anthropic-version: 2023-06-01
•anthropic-beta: interleaved-thinking-2025-05-14
to your Messages API requests whenrequestBody.thinkingis enabled to ensure reasoning blocks andthinking_deltaSSE events are exposed. - (Optional) Gate this header injection behind a feature flag to control rollout.
🤖 Prompt for AI Agents
In packages/models/src/provider-api.ts around lines 193 to 199, when
supportsReasoning is true and requestBody.thinking is set you must add the
Anthropic extended-thinking headers to the Messages API request: set
"anthropic-version" to "2023-06-01" and "anthropic-beta" to
"interleaved-thinking-2025-05-14" on the outgoing request headers; implement
this where the request headers are constructed for the Messages API (only when
requestBody.thinking is enabled) and optionally wrap the header injection behind
an existing or new feature flag to control rollout.
|
superseded by #659 |
Refined handling of Anthropic responses to extract and map reasoning content. Introduced support for streaming and budgeted reasoning tokens. Updated tests and request logic for improved reasoning support.
Summary by CodeRabbit
New Features
Improvements
Tests