Skip to content

fix(gateway): support service_tier in responses API - #2991

Merged
steebchen merged 1 commit into
mainfrom
fix-service-tier-responses-api
Jul 10, 2026
Merged

steebchen merged 1 commit into
mainfrom
fix-service-tier-responses-api

Conversation

@steebchen

@steebchen steebchen commented Jul 10, 2026 •

Copy link
Copy Markdown
Member

Problem

service_tier was ignored on /v1/responses for all models (reported with gpt-5.x): even with service_tier: "flex", the request log showed the tier as default. The parameter only worked on /v1/chat/completions.

Root cause

  • responsesRequestSchema had no service_tier field, so Zod stripped it from the parsed body.
  • The /v1/responses handler never forwarded the tier on the internal /v1/chat/completions hop, so the chat pipeline neither requested flex/priority upstream nor recorded requestedServiceTier/usedServiceTier on the log.
  • Both response converters hardcoded the echo to request.service_tier ?? "default", which was always "default".

Fix

  • Accept service_tier (auto/default/flex/priority, null-tolerant) in the Responses API request schema and forward it to the internal chat request. Tier validation (400 unsupported_service_tier) now applies to /v1/responses exactly like chat completions.
  • Echo the tier the provider actually served on the response: top-level service_tier from the chat response (OpenAI), falling back to the gateway's metadata.used_service_tier (Google flex/priority, incl. downgrades), then the requested tier.
  • Streaming: response.completed/response.incomplete chunks built from the upstream OpenAI Responses API now carry service_tier, and the streaming translator captures it (or the final usage chunk metadata) so response.completed events echo the served tier.

Tests

  • Unit specs for the schema, non-streaming echo (served / downgraded / requested-only), and streaming capture paths.
  • Integration tests against the mock OpenAI server: /v1/responses forwards the tier (log records requestedServiceTier/usedServiceTier), rejects unsupported tiers with 400, and streams the served tier in response.completed.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Added service_tier support to the Responses API.
    • Supported tiers include auto, default, flex, and priority.
    • Responses now report the tier actually used by the provider, including in streaming completion events.
    • Tier information is preserved when requests are converted between API formats.
  • Bug Fixes

    • Unsupported service-tier values now return a clear validation error.
    • Downgraded or unavailable tiers correctly report default.
  • Documentation

    • Clarified how service tiers work with the Responses API.

The /v1/responses request schema stripped service_tier before the
internal chat-completions hop, so flex/priority requests were silently
served (and logged) as the default tier. Accept the parameter, forward
it to the chat pipeline, and echo the tier the provider actually served
on both non-streaming responses and streaming completion events.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings July 10, 2026 17:42

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@coderabbitai

coderabbitai Bot commented Jul 10, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

Walkthrough

The Responses API now accepts and forwards service_tier, resolves the provider-served tier for standard and streaming responses, emits it in response payloads and OpenAI-compatible chunks, and adds validation, integration tests, and documentation.

Changes

Responses service tier support

Layer / File(s) Summary
Request contract and forwarding
apps/gateway/src/responses/schemas.ts, apps/gateway/src/responses/responses.ts, apps/gateway/src/responses/responses.spec.ts
Validates optional nullable service tiers, normalizes null, and forwards provided values to the internal chat-completions request.
Non-streaming served-tier resolution
apps/gateway/src/responses/tools/convert-chat-to-responses.ts, apps/gateway/src/responses/responses.spec.ts
Resolves served tiers from response fields or metadata, with requested-tier and default fallbacks.
Streaming tier propagation and output
apps/gateway/src/responses/tools/convert-streaming-to-responses.ts, apps/gateway/src/chat/tools/transform-streaming-to-openai.ts, apps/gateway/src/responses/responses.spec.ts
Tracks served tiers through streaming state and emits them in completed Responses payloads and OpenAI-compatible chunks.
Integration coverage and documentation
apps/gateway/src/api.spec.ts, apps/docs/content/features/service-tiers.mdx
Covers valid, invalid, and streaming /v1/responses requests and documents Responses API service-tier behavior.

Estimated code review effort: 3 (Moderate) | ~20 minutes

Sequence Diagram(s)

sequenceDiagram
  participant Client
  participant ResponsesAPI
  participant ChatCompletions
  participant Provider
  participant ResponseConverter
  Client->>ResponsesAPI: Send service_tier
  ResponsesAPI->>ChatCompletions: Forward service_tier
  ChatCompletions->>Provider: Request completion
  Provider-->>ChatCompletions: Return served tier
  ChatCompletions-->>ResponseConverter: Return response or stream metadata
  ResponseConverter-->>Client: Emit service_tier
Loading

Possibly related PRs

Suggested reviewers: smakosh, RATCHAW

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely summarizes the main change: adding service_tier support to the Responses API.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix-service-tier-responses-api

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
apps/gateway/src/responses/tools/convert-chat-to-responses.ts (1)

192-213: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Duplicated served-tier resolution logic.

resolveServedServiceTier here is logically identical (same comment, same branching) to the inline block in apps/gateway/src/responses/tools/convert-streaming-to-responses.ts (processStreamChunk, lines 195-210). Extracting a shared helper (e.g. into a small shared utils module) would prevent the two implementations from silently drifting apart as the tier-resolution rules evolve.

♻️ Proposed extraction
+// shared-tier-resolution.ts
+export function resolveServedServiceTierFromFields(
+	serviceTier: unknown,
+	metadata: Record<string, unknown> | undefined,
+): string | undefined {
+	if (typeof serviceTier === "string") {
+		return serviceTier;
+	}
+	if (metadata && typeof metadata.requested_service_tier === "string") {
+		return typeof metadata.used_service_tier === "string"
+			? metadata.used_service_tier
+			: "default";
+	}
+	return undefined;
+}

Both resolveServedServiceTier (non-streaming) and the inline block in processStreamChunk (streaming) can then call this shared helper with their respective service_tier/metadata fields.

Also applies to: 341-344

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@apps/gateway/src/responses/tools/convert-chat-to-responses.ts` around lines
192 - 213, Extract the duplicated served-tier resolution logic from
resolveServedServiceTier and processStreamChunk into a shared utility,
preserving the existing handling of top-level service_tier,
metadata.requested_service_tier, metadata.used_service_tier, and the "default"
fallback. Update both non-streaming and streaming callers to use the helper with
their respective response fields, and remove the duplicated inline logic and
local implementation.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Nitpick comments:
In `@apps/gateway/src/responses/tools/convert-chat-to-responses.ts`:
- Around line 192-213: Extract the duplicated served-tier resolution logic from
resolveServedServiceTier and processStreamChunk into a shared utility,
preserving the existing handling of top-level service_tier,
metadata.requested_service_tier, metadata.used_service_tier, and the "default"
fallback. Update both non-streaming and streaming callers to use the helper with
their respective response fields, and remove the duplicated inline logic and
local implementation.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro

Run ID: f074c825-f610-4685-a452-eedffd61ab14

📥 Commits

Reviewing files that changed from the base of the PR and between 973b27d and 49ca4b1.

📒 Files selected for processing (8)
  • apps/docs/content/features/service-tiers.mdx
  • apps/gateway/src/api.spec.ts
  • apps/gateway/src/chat/tools/transform-streaming-to-openai.ts
  • apps/gateway/src/responses/responses.spec.ts
  • apps/gateway/src/responses/responses.ts
  • apps/gateway/src/responses/schemas.ts
  • apps/gateway/src/responses/tools/convert-chat-to-responses.ts
  • apps/gateway/src/responses/tools/convert-streaming-to-responses.ts

@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@steebchen
steebchen merged commit 22d1af7 into main Jul 10, 2026
17 checks passed
@steebchen
steebchen deleted the fix-service-tier-responses-api branch July 10, 2026 17:54
pull Bot pushed a commit to soitun/llmgateway that referenced this pull request Jul 12, 2026
## Problem

A user reported (Discord) that prompt caching for `meta/muse-spark-1.1`
stopped working: they were billed full input price on every turn of a
pi-agent session, while it had worked the day before.

## Root cause — provider-side routing, reproduced live

Meta's Model API load-balances unkeyed requests randomly across cache
shards, so implicit prefix caching almost never hits without
`prompt_cache_key`. Reproduced directly against `api.meta.ai` (no
gateway involved) with ~4k and ~9k token prefixes: **1 hit in 25 unkeyed
repeats (~4%)** regardless of prefix size or wait time, vs ~80% hits
with a stable key after the first write. It presumably "worked
yesterday" (launch day) because a small preview fleet made accidental
shard affinity common.

The gateway only forwarded `prompt_cache_key` for `usedProvider ===
"openai"` — and coding agents like pi don't send the field at all. No
gateway regression was involved: nothing on the meta request/billing
path changed since theopenco#2985 (audited theopenco#2991, theopenco#2992, and all July 11 merges).

## Fix

Send an upstream `prompt_cache_key` wherever the upstream supports the
field, in this priority order:

1. **Caller-supplied `prompt_cache_key`** — forwarded verbatim
(unchanged).
2. **Salted hash of the resolved session id** (`x-session-id` →
`x-session-affinity` → Claude Code's `metadata.user_id` session, i.e.
the same resolution sticky routing already uses). The id is hashed with
HMAC-SHA256 keyed by the existing `GATEWAY_API_KEY_HASH_SECRET`
(required in production — no insecure fallback; a `prompt-cache-key:`
prefix domain-separates these digests from API-key fingerprints) so
**raw session ids are never exposed to providers**.
3. **Meta only:** a key derived from the conversation's first two
processed messages, so pi-style agents that send no session signal still
get cache hits.

Provider coverage, per research:

- **OpenAI** — already supported; now also gets the session-derived
fallback (both chat completions and Responses API).
- **Azure** — enabled on the Responses-API path ([Microsoft docs
confirm](https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/prompt-caching)
`prompt_cache_key` on the v1 surface, combined with the prefix hash for
routing). The chat-completions path is intentionally excluded: it can
hit legacy deployment-based api-versions that reject unknown body
fields, and the deployment type isn't visible at body-preparation time.
- **Meta** — required for cache hits at all (measurements above).
- **Sakana** — NOT enabled: their docs document prompt caching pricing
but not the `prompt_cache_key` field, and the API blocked direct
verification. Excluded per the only-if-supported rule.

Also documents the behavior in the Sessions docs page.

## Verification (live, local gateway on :4101 → real provider APIs)

- Meta two-turn conversation with `x-session-id`: turn 2 reported
**1777/1963 cached tokens**.
- The queued log row's `upstreamRequest` carried exactly
`prompt_cache_key: sha256(salt:session)[:32]`; the raw session id
appears nowhere in the upstream body.
- OpenAI `gpt-5-mini` (Responses) and `gpt-4o-mini` (chat completions)
both accepted the hashed key and returned normally.
- Azure could not be live-tested (the dev resource has zero
deployments); covered by docs research + unit tests.
- Meta no-session fallback verified earlier: turn 2 cached 2097/2296
(streaming 2161/2320), with billing-queue costs exact to the mapping's
prices ($0.15/M cached, $1.25/M uncached).

Note: Meta's cache writes take a few seconds to propagate; the first
repeat after a write can still miss. That part is provider-side.

## Tests

- 11 unit specs covering: stable per-conversation derivation (meta),
caller key precedence, session-hash precedence over conversation
derivation, openai/azure/meta session paths, sakana exclusion, no key
when no signals, and a sweep asserting the raw session id never appears
in any upstream body.
- `pnpm build` and the full `prepare-request-body` suite (136 tests)
pass.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Enhanced upstream prompt-cache routing: when a session id is present
(and no override is provided), the gateway forwards a deterministic,
salted HMAC-SHA256–based cache key to supported providers; when a key is
supplied by the caller, it is forwarded as-is.
* Meta now uses a conversation-derived cache key to keep caching
consistent across turns.
* **Bug Fixes**
* Avoids sending `prompt_cache_key` to providers/surfaces that don’t
support it, and prevents the raw session id from appearing in upstream
request bodies.
* **Documentation**
* Documented “Upstream prompt-cache routing” behavior and
provider-specific support.
* **Tests**
* Added coverage for stability/differences across conversations, caller
overrides, session-hash derivation, and provider-specific request
assertions.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants