feat(models): add Meta provider with Muse Spark 1.1 - #2985
Conversation
Add Meta's Model API (dev.meta.ai) as a provider using the LLM_META_API_KEY env var, with the muse-spark-1.1 multimodal reasoning model. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository UI Review profile: CHILL Plan: Pro Run ID: 📒 Files selected for processing (9)
🚧 Files skipped from review as they are similar to previous changes (3)
WalkthroughChangesMeta provider integration
Estimated code review effort: 3 (Moderate) | ~25 minutes Possibly related PRs
Suggested reviewers: Sequence Diagram(s)sequenceDiagram
participant Gateway
participant getProviderEndpoint
participant getProviderHeaders
participant MetaAPI
participant transformResponseToOpenai
participant transformStreamingToOpenai
Gateway->>getProviderEndpoint: resolve Meta endpoint
Gateway->>getProviderHeaders: create Bearer authorization header
Gateway->>MetaAPI: send prepared request
MetaAPI-->>transformResponseToOpenai: return response
MetaAPI-->>transformStreamingToOpenai: return streaming chunks
transformResponseToOpenai-->>Gateway: return OpenAI-compatible response
transformStreamingToOpenai-->>Gateway: return normalized stream
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@packages/models/src/providers.ts`:
- Around line 1281-1309: Add valid Meta terms-of-service and privacy-policy URLs
to the `termsUrl` and `privacyPolicyUrl` fields in the `meta` provider
definition, replacing both null values; prefer Meta Model API-specific legal
pages, or use Meta’s public terms and privacy pages as a fallback.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository UI
Review profile: CHILL
Plan: Pro
Run ID: d3a45e89-d780-4f33-a515-c308dda617d5
📒 Files selected for processing (10)
.github/workflows/e2e.ymlapps/gateway/src/chat/tools/transform-streaming-to-openai.tsapps/ui/src/components/provider-keys/provider-logo.tsapps/ui/src/lib/model-category-filters.tspackages/actions/src/get-provider-endpoint.tspackages/actions/src/get-provider-headers.tspackages/models/src/models/meta.tspackages/models/src/providers.tspackages/scripts/src/export-models-dev.tspackages/shared/src/components/provider-icons.tsx
Route meta through /v1/responses so reasoning summaries stream as delta.reasoning (Chat Completions redacts reasoning entirely). Omit reasoning.effort when unset so Muse Spark keeps its adaptive default. Also stop billing reasoning_tokens on top of output for providers whose output token counts already include reasoning (OpenAI, Azure, Sakana, Meta) — verified live that OpenAI/Meta count reasoning inside output_tokens, so the previous math double-billed reasoning output. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
## Problem A user reported (Discord) that prompt caching for `meta/muse-spark-1.1` stopped working: they were billed full input price on every turn of a pi-agent session, while it had worked the day before. ## Root cause — provider-side routing, reproduced live Meta's Model API load-balances unkeyed requests randomly across cache shards, so implicit prefix caching almost never hits without `prompt_cache_key`. Reproduced directly against `api.meta.ai` (no gateway involved) with ~4k and ~9k token prefixes: **1 hit in 25 unkeyed repeats (~4%)** regardless of prefix size or wait time, vs ~80% hits with a stable key after the first write. It presumably "worked yesterday" (launch day) because a small preview fleet made accidental shard affinity common. The gateway only forwarded `prompt_cache_key` for `usedProvider === "openai"` — and coding agents like pi don't send the field at all. No gateway regression was involved: nothing on the meta request/billing path changed since theopenco#2985 (audited theopenco#2991, theopenco#2992, and all July 11 merges). ## Fix Send an upstream `prompt_cache_key` wherever the upstream supports the field, in this priority order: 1. **Caller-supplied `prompt_cache_key`** — forwarded verbatim (unchanged). 2. **Salted hash of the resolved session id** (`x-session-id` → `x-session-affinity` → Claude Code's `metadata.user_id` session, i.e. the same resolution sticky routing already uses). The id is hashed with HMAC-SHA256 keyed by the existing `GATEWAY_API_KEY_HASH_SECRET` (required in production — no insecure fallback; a `prompt-cache-key:` prefix domain-separates these digests from API-key fingerprints) so **raw session ids are never exposed to providers**. 3. **Meta only:** a key derived from the conversation's first two processed messages, so pi-style agents that send no session signal still get cache hits. Provider coverage, per research: - **OpenAI** — already supported; now also gets the session-derived fallback (both chat completions and Responses API). - **Azure** — enabled on the Responses-API path ([Microsoft docs confirm](https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/prompt-caching) `prompt_cache_key` on the v1 surface, combined with the prefix hash for routing). The chat-completions path is intentionally excluded: it can hit legacy deployment-based api-versions that reject unknown body fields, and the deployment type isn't visible at body-preparation time. - **Meta** — required for cache hits at all (measurements above). - **Sakana** — NOT enabled: their docs document prompt caching pricing but not the `prompt_cache_key` field, and the API blocked direct verification. Excluded per the only-if-supported rule. Also documents the behavior in the Sessions docs page. ## Verification (live, local gateway on :4101 → real provider APIs) - Meta two-turn conversation with `x-session-id`: turn 2 reported **1777/1963 cached tokens**. - The queued log row's `upstreamRequest` carried exactly `prompt_cache_key: sha256(salt:session)[:32]`; the raw session id appears nowhere in the upstream body. - OpenAI `gpt-5-mini` (Responses) and `gpt-4o-mini` (chat completions) both accepted the hashed key and returned normally. - Azure could not be live-tested (the dev resource has zero deployments); covered by docs research + unit tests. - Meta no-session fallback verified earlier: turn 2 cached 2097/2296 (streaming 2161/2320), with billing-queue costs exact to the mapping's prices ($0.15/M cached, $1.25/M uncached). Note: Meta's cache writes take a few seconds to propagate; the first repeat after a write can still miss. That part is provider-side. ## Tests - 11 unit specs covering: stable per-conversation derivation (meta), caller key precedence, session-hash precedence over conversation derivation, openai/azure/meta session paths, sakana exclusion, no key when no signals, and a sweep asserting the raw session id never appears in any upstream body. - `pnpm build` and the full `prepare-request-body` suite (136 tests) pass. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Enhanced upstream prompt-cache routing: when a session id is present (and no override is provided), the gateway forwards a deterministic, salted HMAC-SHA256–based cache key to supported providers; when a key is supplied by the caller, it is forwarded as-is. * Meta now uses a conversation-derived cache key to keep caching consistent across turns. * **Bug Fixes** * Avoids sending `prompt_cache_key` to providers/surfaces that don’t support it, and prevents the raw session id from appearing in upstream request bodies. * **Documentation** * Documented “Upstream prompt-cache routing” behavior and provider-specific support. * **Tests** * Added coverage for stability/differences across conversations, caller overrides, session-hash derivation, and provider-specific request assertions. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Summary
LLM_META_API_KEYenv varstability: "beta") to the existingmetafamily:reasoning_contententirely, so the gateway routesmetathrough the Responses API (supportsResponsesApi: true) and requestsreasoning.summary. Reasoning summaries stream asdelta.reasoning(verified live). Non-streaming requests returnsummary: []upstream — a Meta-side limitation.reasoning.effortis only forwarded when the caller setsreasoning_effort, preserving Muse Spark's adaptive default (it also rejectseffort: "none")output_tokensinclusive of reasoning;reasoning_tokensis informational.calculateCostspreviously addedreasoningTokenson top of completion tokens for all non-Google providers, double-billing reasoning output. Verified live: gpt-5-mini billed 721 output tokens for a 401-token response before, exactly 401 after; muse-spark billed 1580 before, 809 after. Token reporting is unchanged (completion tokens keep including reasoning)muse-spark-1.1from the open-source filter (proprietary model in the otherwise-open Llama family)Test plan
pnpm buildpassesapi.spec.ts, reproduced on clean tree)x-no-fallback: meta streaming (reasoning deltas + correct usage/costs), meta non-streaming, openai gpt-5-mini (double-billing gone, reasoning still exposed)LLM_META_API_KEYrepo secret so e2e covers the new mapping🤖 Generated with Claude Code
Summary by CodeRabbit
New Features
Bug Fixes