Skip to content

feat(models): add Meta provider with Muse Spark 1.1 - #2985

Merged
steebchen merged 4 commits into
mainfrom
add-meta-ai-provider
Jul 10, 2026
Merged

steebchen merged 4 commits into
mainfrom
add-meta-ai-provider

Conversation

@steebchen

@steebchen steebchen commented Jul 10, 2026 •

Copy link
Copy Markdown
Member

Summary

  • Add Meta (Model API, https://dev.meta.ai) as a new OpenAI-compatible provider, authenticated via the LLM_META_API_KEY env var
  • Add the muse-spark-1.1 model (released 2026-07-09, public preview → stability: "beta") to the existing meta family:
    • $1.25/M input, $4.25/M output, $0.15/M cached input, 1,048,576-token context, 131,072 max output
    • streaming, tools, vision, JSON output + schema, adaptive reasoning
  • Reasoning exposure: Meta's Chat Completions endpoint redacts reasoning_content entirely, so the gateway routes meta through the Responses API (supportsResponsesApi: true) and requests reasoning.summary. Reasoning summaries stream as delta.reasoning (verified live). Non-streaming requests return summary: [] upstream — a Meta-side limitation. reasoning.effort is only forwarded when the caller sets reasoning_effort, preserving Muse Spark's adaptive default (it also rejects effort: "none")
  • Multi-part Responses reasoning summaries are now joined instead of only taking the first part
  • Billing fix (also affects OpenAI/Azure/Sakana): providers using the Responses API report output_tokens inclusive of reasoning; reasoning_tokens is informational. calculateCosts previously added reasoningTokens on top of completion tokens for all non-Google providers, double-billing reasoning output. Verified live: gpt-5-mini billed 721 output tokens for a 401-token response before, exactly 401 after; muse-spark billed 1580 before, 809 after. Token reporting is unchanged (completion tokens keep including reasoning)
  • Wire endpoint routing, auth headers, streaming transform, provider icons, e2e workflow secret; exclude muse-spark-1.1 from the open-source filter (proprietary model in the otherwise-open Llama family)

Test plan

  • pnpm build passes
  • Gateway/actions/costs unit suites pass (one pre-existing env-dependent routing failure in api.spec.ts, reproduced on clean tree)
  • Verified live via local gateway with x-no-fallback: meta streaming (reasoning deltas + correct usage/costs), meta non-streaming, openai gpt-5-mini (double-billing gone, reasoning still exposed)
  • Add the LLM_META_API_KEY repo secret so e2e covers the new mapping

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Added Meta as a supported AI provider, including API configuration, provider branding, streaming, cancellation, and model selection.
    • Added support for the Muse Spark 1.1 model, including vision, tools, JSON output, and reasoning capabilities.
    • Added Meta provider icons and UI branding.
  • Bug Fixes

    • Improved compatibility with Meta’s responses and streaming formats.
    • Preserved complete multi-part reasoning summaries.
    • Improved token usage and cost calculations across supported providers.
    • Refined open-source model filtering to exclude closed-source models with matching families.

Add Meta's Model API (dev.meta.ai) as a provider using the
LLM_META_API_KEY env var, with the muse-spark-1.1 multimodal
reasoning model.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings July 10, 2026 14:11
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@coderabbitai

coderabbitai Bot commented Jul 10, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro

Run ID: 44d5aae7-a805-4fc8-ba4f-030616e298b4

📥 Commits

Reviewing files that changed from the base of the PR and between 5b11dc0 and d5f0406.

📒 Files selected for processing (9)
  • apps/gateway/src/chat/tools/parse-provider-response.ts
  • apps/gateway/src/chat/tools/transform-response-to-openai.ts
  • apps/gateway/src/chat/tools/transform-streaming-to-openai.ts
  • apps/gateway/src/lib/costs.ts
  • packages/actions/src/get-provider-endpoint.ts
  • packages/actions/src/prepare-request-body.ts
  • packages/models/src/models/meta.ts
  • packages/models/src/providers.ts
  • packages/models/src/types.ts
🚧 Files skipped from review as they are similar to previous changes (3)
  • packages/models/src/models/meta.ts
  • apps/gateway/src/chat/tools/transform-streaming-to-openai.ts
  • packages/models/src/providers.ts

Walkthrough

Changes

Meta provider integration

Layer / File(s) Summary
Provider and model registration
packages/models/src/providers.ts, packages/models/src/models/meta.ts, packages/scripts/src/export-models-dev.ts, packages/shared/src/components/provider-icons.tsx
Adds the Meta provider, Muse Spark 1.1 model, family detection, and Meta icon mappings.
Meta request routing and transformation
packages/actions/..., apps/gateway/src/chat/tools/..., packages/models/src/types.ts, .github/workflows/e2e.yml
Routes Meta requests to its API with Bearer authentication, supports Responses and Chat Completions endpoints, prepares optional reasoning fields, transforms responses, and supplies the e2e API key.
Reasoning parsing and token accounting
apps/gateway/src/chat/tools/parse-provider-response.ts, apps/gateway/src/lib/costs.ts
Combines multi-part reasoning summaries and updates completion-token handling for providers whose totals include reasoning tokens.
UI provider display and model filtering
apps/ui/src/components/provider-keys/provider-logo.ts, apps/ui/src/lib/model-category-filters.ts
Adds the Meta logo and excludes configured closed-source model IDs from the open-source filter.

Estimated code review effort: 3 (Moderate) | ~25 minutes

Possibly related PRs

Suggested reviewers: smakosh

Sequence Diagram(s)

sequenceDiagram
  participant Gateway
  participant getProviderEndpoint
  participant getProviderHeaders
  participant MetaAPI
  participant transformResponseToOpenai
  participant transformStreamingToOpenai
  Gateway->>getProviderEndpoint: resolve Meta endpoint
  Gateway->>getProviderHeaders: create Bearer authorization header
  Gateway->>MetaAPI: send prepared request
  MetaAPI-->>transformResponseToOpenai: return response
  MetaAPI-->>transformStreamingToOpenai: return streaming chunks
  transformResponseToOpenai-->>Gateway: return OpenAI-compatible response
  transformStreamingToOpenai-->>Gateway: return normalized stream
Loading
🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 42.86% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly summarizes the main change: adding the Meta provider and the Muse Spark 1.1 model.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch add-meta-ai-provider

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@packages/models/src/providers.ts`:
- Around line 1281-1309: Add valid Meta terms-of-service and privacy-policy URLs
to the `termsUrl` and `privacyPolicyUrl` fields in the `meta` provider
definition, replacing both null values; prefer Meta Model API-specific legal
pages, or use Meta’s public terms and privacy pages as a fallback.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro

Run ID: d3a45e89-d780-4f33-a515-c308dda617d5

📥 Commits

Reviewing files that changed from the base of the PR and between ab43bad and 5b11dc0.

📒 Files selected for processing (10)
  • .github/workflows/e2e.yml
  • apps/gateway/src/chat/tools/transform-streaming-to-openai.ts
  • apps/ui/src/components/provider-keys/provider-logo.ts
  • apps/ui/src/lib/model-category-filters.ts
  • packages/actions/src/get-provider-endpoint.ts
  • packages/actions/src/get-provider-headers.ts
  • packages/models/src/models/meta.ts
  • packages/models/src/providers.ts
  • packages/scripts/src/export-models-dev.ts
  • packages/shared/src/components/provider-icons.tsx

Comment thread packages/models/src/providers.ts
steebchen and others added 3 commits July 10, 2026 16:16
Route meta through /v1/responses so reasoning summaries stream as
delta.reasoning (Chat Completions redacts reasoning entirely). Omit
reasoning.effort when unset so Muse Spark keeps its adaptive default.

Also stop billing reasoning_tokens on top of output for providers
whose output token counts already include reasoning (OpenAI, Azure,
Sakana, Meta) — verified live that OpenAI/Meta count reasoning inside
output_tokens, so the previous math double-billed reasoning output.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@steebchen
steebchen enabled auto-merge July 10, 2026 16:30
@steebchen
steebchen added this pull request to the merge queue Jul 10, 2026
Merged via the queue into main with commit 973b27d Jul 10, 2026
16 of 17 checks passed
@steebchen
steebchen deleted the add-meta-ai-provider branch July 10, 2026 16:52
pull Bot pushed a commit to soitun/llmgateway that referenced this pull request Jul 12, 2026
## Problem

A user reported (Discord) that prompt caching for `meta/muse-spark-1.1`
stopped working: they were billed full input price on every turn of a
pi-agent session, while it had worked the day before.

## Root cause — provider-side routing, reproduced live

Meta's Model API load-balances unkeyed requests randomly across cache
shards, so implicit prefix caching almost never hits without
`prompt_cache_key`. Reproduced directly against `api.meta.ai` (no
gateway involved) with ~4k and ~9k token prefixes: **1 hit in 25 unkeyed
repeats (~4%)** regardless of prefix size or wait time, vs ~80% hits
with a stable key after the first write. It presumably "worked
yesterday" (launch day) because a small preview fleet made accidental
shard affinity common.

The gateway only forwarded `prompt_cache_key` for `usedProvider ===
"openai"` — and coding agents like pi don't send the field at all. No
gateway regression was involved: nothing on the meta request/billing
path changed since theopenco#2985 (audited theopenco#2991, theopenco#2992, and all July 11 merges).

## Fix

Send an upstream `prompt_cache_key` wherever the upstream supports the
field, in this priority order:

1. **Caller-supplied `prompt_cache_key`** — forwarded verbatim
(unchanged).
2. **Salted hash of the resolved session id** (`x-session-id` →
`x-session-affinity` → Claude Code's `metadata.user_id` session, i.e.
the same resolution sticky routing already uses). The id is hashed with
HMAC-SHA256 keyed by the existing `GATEWAY_API_KEY_HASH_SECRET`
(required in production — no insecure fallback; a `prompt-cache-key:`
prefix domain-separates these digests from API-key fingerprints) so
**raw session ids are never exposed to providers**.
3. **Meta only:** a key derived from the conversation's first two
processed messages, so pi-style agents that send no session signal still
get cache hits.

Provider coverage, per research:

- **OpenAI** — already supported; now also gets the session-derived
fallback (both chat completions and Responses API).
- **Azure** — enabled on the Responses-API path ([Microsoft docs
confirm](https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/prompt-caching)
`prompt_cache_key` on the v1 surface, combined with the prefix hash for
routing). The chat-completions path is intentionally excluded: it can
hit legacy deployment-based api-versions that reject unknown body
fields, and the deployment type isn't visible at body-preparation time.
- **Meta** — required for cache hits at all (measurements above).
- **Sakana** — NOT enabled: their docs document prompt caching pricing
but not the `prompt_cache_key` field, and the API blocked direct
verification. Excluded per the only-if-supported rule.

Also documents the behavior in the Sessions docs page.

## Verification (live, local gateway on :4101 → real provider APIs)

- Meta two-turn conversation with `x-session-id`: turn 2 reported
**1777/1963 cached tokens**.
- The queued log row's `upstreamRequest` carried exactly
`prompt_cache_key: sha256(salt:session)[:32]`; the raw session id
appears nowhere in the upstream body.
- OpenAI `gpt-5-mini` (Responses) and `gpt-4o-mini` (chat completions)
both accepted the hashed key and returned normally.
- Azure could not be live-tested (the dev resource has zero
deployments); covered by docs research + unit tests.
- Meta no-session fallback verified earlier: turn 2 cached 2097/2296
(streaming 2161/2320), with billing-queue costs exact to the mapping's
prices ($0.15/M cached, $1.25/M uncached).

Note: Meta's cache writes take a few seconds to propagate; the first
repeat after a write can still miss. That part is provider-side.

## Tests

- 11 unit specs covering: stable per-conversation derivation (meta),
caller key precedence, session-hash precedence over conversation
derivation, openai/azure/meta session paths, sakana exclusion, no key
when no signals, and a sweep asserting the raw session id never appears
in any upstream body.
- `pnpm build` and the full `prepare-request-body` suite (136 tests)
pass.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Enhanced upstream prompt-cache routing: when a session id is present
(and no override is provided), the gateway forwards a deterministic,
salted HMAC-SHA256–based cache key to supported providers; when a key is
supplied by the caller, it is forwarded as-is.
* Meta now uses a conversation-derived cache key to keep caching
consistent across turns.
* **Bug Fixes**
* Avoids sending `prompt_cache_key` to providers/surfaces that don’t
support it, and prevents the raw session id from appearing in upstream
request bodies.
* **Documentation**
* Documented “Upstream prompt-cache routing” behavior and
provider-specific support.
* **Tests**
* Added coverage for stability/differences across conversations, caller
overrides, session-hash derivation, and provider-specific request
assertions.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants