Repository navigation
fix(gateway): support service_tier in responses API - #2991
Conversation
The /v1/responses request schema stripped service_tier before the internal chat-completions hop, so flex/priority requests were silently served (and logged) as the default tier. Accept the parameter, forward it to the chat pipeline, and echo the tier the provider actually served on both non-streaming responses and streaming completion events. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
WalkthroughThe Responses API now accepts and forwards ChangesResponses service tier support
Estimated code review effort: 3 (Moderate) | ~20 minutes Sequence Diagram(s)sequenceDiagram
participant Client
participant ResponsesAPI
participant ChatCompletions
participant Provider
participant ResponseConverter
Client->>ResponsesAPI: Send service_tier
ResponsesAPI->>ChatCompletions: Forward service_tier
ChatCompletions->>Provider: Request completion
Provider-->>ChatCompletions: Return served tier
ChatCompletions-->>ResponseConverter: Return response or stream metadata
ResponseConverter-->>Client: Emit service_tier
Possibly related PRs
Suggested reviewers: 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
🧹 Nitpick comments (1)
apps/gateway/src/responses/tools/convert-chat-to-responses.ts (1)
192-213: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick winDuplicated served-tier resolution logic.
resolveServedServiceTierhere is logically identical (same comment, same branching) to the inline block inapps/gateway/src/responses/tools/convert-streaming-to-responses.ts(processStreamChunk, lines 195-210). Extracting a shared helper (e.g. into a small shared utils module) would prevent the two implementations from silently drifting apart as the tier-resolution rules evolve.♻️ Proposed extraction
+// shared-tier-resolution.ts +export function resolveServedServiceTierFromFields( + serviceTier: unknown, + metadata: Record<string, unknown> | undefined, +): string | undefined { + if (typeof serviceTier === "string") { + return serviceTier; + } + if (metadata && typeof metadata.requested_service_tier === "string") { + return typeof metadata.used_service_tier === "string" + ? metadata.used_service_tier + : "default"; + } + return undefined; +}Both
resolveServedServiceTier(non-streaming) and the inline block inprocessStreamChunk(streaming) can then call this shared helper with their respectiveservice_tier/metadatafields.Also applies to: 341-344
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@apps/gateway/src/responses/tools/convert-chat-to-responses.ts` around lines 192 - 213, Extract the duplicated served-tier resolution logic from resolveServedServiceTier and processStreamChunk into a shared utility, preserving the existing handling of top-level service_tier, metadata.requested_service_tier, metadata.used_service_tier, and the "default" fallback. Update both non-streaming and streaming callers to use the helper with their respective response fields, and remove the duplicated inline logic and local implementation.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Nitpick comments:
In `@apps/gateway/src/responses/tools/convert-chat-to-responses.ts`:
- Around line 192-213: Extract the duplicated served-tier resolution logic from
resolveServedServiceTier and processStreamChunk into a shared utility,
preserving the existing handling of top-level service_tier,
metadata.requested_service_tier, metadata.used_service_tier, and the "default"
fallback. Update both non-streaming and streaming callers to use the helper with
their respective response fields, and remove the duplicated inline logic and
local implementation.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository UI
Review profile: CHILL
Plan: Pro
Run ID: f074c825-f610-4685-a452-eedffd61ab14
📒 Files selected for processing (8)
apps/docs/content/features/service-tiers.mdxapps/gateway/src/api.spec.tsapps/gateway/src/chat/tools/transform-streaming-to-openai.tsapps/gateway/src/responses/responses.spec.tsapps/gateway/src/responses/responses.tsapps/gateway/src/responses/schemas.tsapps/gateway/src/responses/tools/convert-chat-to-responses.tsapps/gateway/src/responses/tools/convert-streaming-to-responses.ts
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
## Problem A user reported (Discord) that prompt caching for `meta/muse-spark-1.1` stopped working: they were billed full input price on every turn of a pi-agent session, while it had worked the day before. ## Root cause — provider-side routing, reproduced live Meta's Model API load-balances unkeyed requests randomly across cache shards, so implicit prefix caching almost never hits without `prompt_cache_key`. Reproduced directly against `api.meta.ai` (no gateway involved) with ~4k and ~9k token prefixes: **1 hit in 25 unkeyed repeats (~4%)** regardless of prefix size or wait time, vs ~80% hits with a stable key after the first write. It presumably "worked yesterday" (launch day) because a small preview fleet made accidental shard affinity common. The gateway only forwarded `prompt_cache_key` for `usedProvider === "openai"` — and coding agents like pi don't send the field at all. No gateway regression was involved: nothing on the meta request/billing path changed since theopenco#2985 (audited theopenco#2991, theopenco#2992, and all July 11 merges). ## Fix Send an upstream `prompt_cache_key` wherever the upstream supports the field, in this priority order: 1. **Caller-supplied `prompt_cache_key`** — forwarded verbatim (unchanged). 2. **Salted hash of the resolved session id** (`x-session-id` → `x-session-affinity` → Claude Code's `metadata.user_id` session, i.e. the same resolution sticky routing already uses). The id is hashed with HMAC-SHA256 keyed by the existing `GATEWAY_API_KEY_HASH_SECRET` (required in production — no insecure fallback; a `prompt-cache-key:` prefix domain-separates these digests from API-key fingerprints) so **raw session ids are never exposed to providers**. 3. **Meta only:** a key derived from the conversation's first two processed messages, so pi-style agents that send no session signal still get cache hits. Provider coverage, per research: - **OpenAI** — already supported; now also gets the session-derived fallback (both chat completions and Responses API). - **Azure** — enabled on the Responses-API path ([Microsoft docs confirm](https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/prompt-caching) `prompt_cache_key` on the v1 surface, combined with the prefix hash for routing). The chat-completions path is intentionally excluded: it can hit legacy deployment-based api-versions that reject unknown body fields, and the deployment type isn't visible at body-preparation time. - **Meta** — required for cache hits at all (measurements above). - **Sakana** — NOT enabled: their docs document prompt caching pricing but not the `prompt_cache_key` field, and the API blocked direct verification. Excluded per the only-if-supported rule. Also documents the behavior in the Sessions docs page. ## Verification (live, local gateway on :4101 → real provider APIs) - Meta two-turn conversation with `x-session-id`: turn 2 reported **1777/1963 cached tokens**. - The queued log row's `upstreamRequest` carried exactly `prompt_cache_key: sha256(salt:session)[:32]`; the raw session id appears nowhere in the upstream body. - OpenAI `gpt-5-mini` (Responses) and `gpt-4o-mini` (chat completions) both accepted the hashed key and returned normally. - Azure could not be live-tested (the dev resource has zero deployments); covered by docs research + unit tests. - Meta no-session fallback verified earlier: turn 2 cached 2097/2296 (streaming 2161/2320), with billing-queue costs exact to the mapping's prices ($0.15/M cached, $1.25/M uncached). Note: Meta's cache writes take a few seconds to propagate; the first repeat after a write can still miss. That part is provider-side. ## Tests - 11 unit specs covering: stable per-conversation derivation (meta), caller key precedence, session-hash precedence over conversation derivation, openai/azure/meta session paths, sakana exclusion, no key when no signals, and a sweep asserting the raw session id never appears in any upstream body. - `pnpm build` and the full `prepare-request-body` suite (136 tests) pass. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Enhanced upstream prompt-cache routing: when a session id is present (and no override is provided), the gateway forwards a deterministic, salted HMAC-SHA256–based cache key to supported providers; when a key is supplied by the caller, it is forwarded as-is. * Meta now uses a conversation-derived cache key to keep caching consistent across turns. * **Bug Fixes** * Avoids sending `prompt_cache_key` to providers/surfaces that don’t support it, and prevents the raw session id from appearing in upstream request bodies. * **Documentation** * Documented “Upstream prompt-cache routing” behavior and provider-specific support. * **Tests** * Added coverage for stability/differences across conversations, caller overrides, session-hash derivation, and provider-specific request assertions. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Problem
service_tierwas ignored on/v1/responsesfor all models (reported with gpt-5.x): even withservice_tier: "flex", the request log showed the tier as default. The parameter only worked on/v1/chat/completions.Root cause
responsesRequestSchemahad noservice_tierfield, so Zod stripped it from the parsed body./v1/chat/completionshop, so the chat pipeline neither requested flex/priority upstream nor recordedrequestedServiceTier/usedServiceTieron the log.request.service_tier ?? "default", which was always"default".Fix
service_tier(auto/default/flex/priority, null-tolerant) in the Responses API request schema and forward it to the internal chat request. Tier validation (400unsupported_service_tier) now applies to /v1/responses exactly like chat completions.service_tierfrom the chat response (OpenAI), falling back to the gateway'smetadata.used_service_tier(Google flex/priority, incl. downgrades), then the requested tier.response.completed/response.incompletechunks built from the upstream OpenAI Responses API now carryservice_tier, and the streaming translator captures it (or the final usage chunk metadata) soresponse.completedevents echo the served tier.Tests
requestedServiceTier/usedServiceTier), rejects unsupported tiers with 400, and streams the served tier inresponse.completed.🤖 Generated with Claude Code
Summary by CodeRabbit
New Features
service_tiersupport to the Responses API.auto,default,flex, andpriority.Bug Fixes
default.Documentation