fix(gateway): web_search tool in responses api - #2992
Conversation
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
|
Warning Review limit reached
Next review available in: 27 minutes Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available. How can I continue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews. How do review limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: Repository UI Review profile: CHILL Plan: Pro Run ID: 📒 Files selected for processing (5)
WalkthroughThe Responses API now forwards ChangesResponses web search and annotations
Estimated code review effort: 3 (Moderate) | ~25 minutes Sequence Diagram(s)sequenceDiagram
participant Client
participant ResponsesEndpoint
participant ChatCompletions
participant RequestLogs
Client->>ResponsesEndpoint: POST /v1/responses with web_search
ResponsesEndpoint->>ChatCompletions: Forward web_search tool
ChatCompletions-->>ResponsesEndpoint: Return search response
ResponsesEndpoint->>RequestLogs: Record webSearchCost
ResponsesEndpoint-->>Client: Return Responses output
sequenceDiagram
participant ChatCompletionsStream
participant StreamConverter
participant StreamingState
participant ResponsesClient
ChatCompletionsStream->>StreamConverter: Send delta annotations
StreamConverter->>StreamingState: Normalize and accumulate annotations
StreamConverter-->>ResponsesClient: Emit annotation.added event
StreamingState-->>ResponsesClient: Include annotations in completed output
Possibly related PRs
Suggested reviewers: 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
The /v1/responses handler dropped all non-function tools when converting to the internal chat completions request, so web_search never reached the provider even though chat completions supports it natively. Pass web_search through, forward url citations in both the non-streaming and streaming converters (flattened to the Responses API annotation shape), and accept allowed_domains/blocked_domains on the responses web_search tool schema for parity with chat completions. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
41096ce to
cd7ddae
Compare
## Problem A user reported (Discord) that prompt caching for `meta/muse-spark-1.1` stopped working: they were billed full input price on every turn of a pi-agent session, while it had worked the day before. ## Root cause — provider-side routing, reproduced live Meta's Model API load-balances unkeyed requests randomly across cache shards, so implicit prefix caching almost never hits without `prompt_cache_key`. Reproduced directly against `api.meta.ai` (no gateway involved) with ~4k and ~9k token prefixes: **1 hit in 25 unkeyed repeats (~4%)** regardless of prefix size or wait time, vs ~80% hits with a stable key after the first write. It presumably "worked yesterday" (launch day) because a small preview fleet made accidental shard affinity common. The gateway only forwarded `prompt_cache_key` for `usedProvider === "openai"` — and coding agents like pi don't send the field at all. No gateway regression was involved: nothing on the meta request/billing path changed since theopenco#2985 (audited theopenco#2991, theopenco#2992, and all July 11 merges). ## Fix Send an upstream `prompt_cache_key` wherever the upstream supports the field, in this priority order: 1. **Caller-supplied `prompt_cache_key`** — forwarded verbatim (unchanged). 2. **Salted hash of the resolved session id** (`x-session-id` → `x-session-affinity` → Claude Code's `metadata.user_id` session, i.e. the same resolution sticky routing already uses). The id is hashed with HMAC-SHA256 keyed by the existing `GATEWAY_API_KEY_HASH_SECRET` (required in production — no insecure fallback; a `prompt-cache-key:` prefix domain-separates these digests from API-key fingerprints) so **raw session ids are never exposed to providers**. 3. **Meta only:** a key derived from the conversation's first two processed messages, so pi-style agents that send no session signal still get cache hits. Provider coverage, per research: - **OpenAI** — already supported; now also gets the session-derived fallback (both chat completions and Responses API). - **Azure** — enabled on the Responses-API path ([Microsoft docs confirm](https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/prompt-caching) `prompt_cache_key` on the v1 surface, combined with the prefix hash for routing). The chat-completions path is intentionally excluded: it can hit legacy deployment-based api-versions that reject unknown body fields, and the deployment type isn't visible at body-preparation time. - **Meta** — required for cache hits at all (measurements above). - **Sakana** — NOT enabled: their docs document prompt caching pricing but not the `prompt_cache_key` field, and the API blocked direct verification. Excluded per the only-if-supported rule. Also documents the behavior in the Sessions docs page. ## Verification (live, local gateway on :4101 → real provider APIs) - Meta two-turn conversation with `x-session-id`: turn 2 reported **1777/1963 cached tokens**. - The queued log row's `upstreamRequest` carried exactly `prompt_cache_key: sha256(salt:session)[:32]`; the raw session id appears nowhere in the upstream body. - OpenAI `gpt-5-mini` (Responses) and `gpt-4o-mini` (chat completions) both accepted the hashed key and returned normally. - Azure could not be live-tested (the dev resource has zero deployments); covered by docs research + unit tests. - Meta no-session fallback verified earlier: turn 2 cached 2097/2296 (streaming 2161/2320), with billing-queue costs exact to the mapping's prices ($0.15/M cached, $1.25/M uncached). Note: Meta's cache writes take a few seconds to propagate; the first repeat after a write can still miss. That part is provider-side. ## Tests - 11 unit specs covering: stable per-conversation derivation (meta), caller key precedence, session-hash precedence over conversation derivation, openai/azure/meta session paths, sakana exclusion, no key when no signals, and a sweep asserting the raw session id never appears in any upstream body. - `pnpm build` and the full `prepare-request-body` suite (136 tests) pass. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Enhanced upstream prompt-cache routing: when a session id is present (and no override is provided), the gateway forwards a deterministic, salted HMAC-SHA256–based cache key to supported providers; when a key is supplied by the caller, it is forwarded as-is. * Meta now uses a conversation-derived cache key to keep caching consistent across turns. * **Bug Fixes** * Avoids sending `prompt_cache_key` to providers/surfaces that don’t support it, and prevents the raw session id from appearing in upstream request bodies. * **Documentation** * Documented “Upstream prompt-cache routing” behavior and provider-specific support. * **Tests** * Added coverage for stability/differences across conversations, caller overrides, session-hash derivation, and provider-specific request assertions. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Summary
A user reported that
web_searchtools work via/v1/chat/completionsbut are silently ignored via/v1/responses(the model answers "I don't have real-time access" andweb_search_coststays 0).Root cause: the Responses API handler converts the request to an internal chat completions request but dropped every non-
functiontool, includingweb_search— even though the chat completions layer supports it natively (provider grounding / native web search).Changes
responses.ts: passweb_searchtools through to the internal chat completions request instead of dropping them. Other OpenAI-native built-ins (computer_use, code_interpreter, …) are still dropped.schemas.ts: acceptallowed_domains/blocked_domainson the responsesweb_searchtool for parity with chat completions (zod was silently stripping them).convert-chat-to-responses.ts: flatten chat-styleurl_citationannotations (nested underurl_citation) to the Responses API annotation shape.convert-streaming-to-responses.ts: forward citation annotations in the streaming path (previously hardcoded to[]) — accumulatesdelta.annotations, emitsresponse.output_text.annotation.addedevents, and includes them incontent_part.done/output_item.done/final output.chat-websearch.e2e.ts: new e2e test exercisingweb_searchvia/v1/responses, asserting output text andwebSearchCost > 0in logs.Testing
TEST_WEB_SEARCH=1 TEST_MODELS="google-ai-studio/gemini-3.5-flash" pnpm test:e2e— all three web search tests pass (non-streaming, responses api, streaming) against the real provider withwebSearchCost > 0. (Remaining failures in that run are pre-existing gemini JSON-mode flakes on the untouched chat completions path.)pnpm test:unit— only pre-existing failure (api.spec.tshybrid-escape test, fails on a clean tree too).pnpm buildpasses.🤖 Generated with Claude Code
Summary by CodeRabbit