Skip to content

feat(models): add RanoAI provider - #3265

Merged
steebchen merged 12 commits into
mainfrom
add-ranoai-provider
Aug 7, 2026
Merged

steebchen merged 12 commits into
mainfrom
add-ranoai-provider

Conversation

@steebchen

@steebchen steebchen commented Jul 27, 2026 •

Copy link
Copy Markdown
Member

Adds RanoAI (ranoai) — an OpenAI-compatible inference provider serving Gemma 4 31B on Furiosa RNGD NPU hardware, via https://api.ranoai.com/v1/chat/completions.

Changes

  • Provider registration (packages/models/src/providers.ts): ranoai, API key via LLM_RANOAI_API_KEY, streaming + cancellation, full metadata (website, terms, privacy, HQ US, data policy).
  • Endpoint resolution (packages/actions/src/get-provider-endpoint.ts): default base URL https://api.ranoai.com, routed to /v1/chat/completions.
  • Streaming transform (apps/gateway/src/chat/tools/transform-streaming-to-openai.ts): added to the OpenAI-compatible passthrough branch.
  • Model mapping (packages/models/src/models/google.ts): new ranoai mapping on the existing gemma-4-31b-it model (externalId: gemma-4-31b).
  • Brand logo (packages/shared/src/components/provider-icons.tsx, apps/ui/.../provider-logo.ts): RanoAIIcon registered in ProviderIcons + both providerLogoUrls maps.
  • Env: .env.example, .env.unified.example, Helm values.yaml, and the e2e CI workflow.

Pricing

Provider list price, passed through as-is: $0.10/M input, $0.30/M output — currently the cheapest Gemma 4 31B mapping in the catalogue.

Capability flags — measured against the live deployment

Every flag was verified by probing the deployment directly rather than reading the docs. RanoAI shipped several fixes during review (image input, tool-call formatting including streaming deltas and tool-result round-trips, tool_choice: "none", and a window expansion from 8K to 20480), so this reflects the state as of 2026-07-28.

Capability Measured Flag
Context window 20480 tokens (prompt + max_tokens), rejected with a proper OpenAI-style error contextSize: 20480, maxOutput: 20480
Vision works — data URIs and remote HTTPS URLs both decode and are described accurately vision: true
Tools correct tool_calls non-streaming and streaming; parallel multi-tool works; tool-result round-trips clean tools: true
JSON schema strict json_schema conforms to nested enum schemas (3/3) jsonOutputSchema: true
Multiple choices n=3 returns 3 distinct indexed choices, streaming included supportsN: true
tool_choice auto and none honoured; required accepted but not enforced; named choices ignored (see below) supportedToolChoices: ["auto", "none"]
Streaming works, incl. stream_options.include_usage streaming: true
JSON output json_object and json_schema both return valid JSON jsonOutput: true
Reasoning reasoning_content across all 7 effort tiers, streamed as deltas; only emitted when reasoning_effort is passed reasoning: true + full reasoningEfforts

Vision was confirmed with a synthetic image the model could not have guessed (a 64×64 half-red/half-blue PNG, correctly described as "red and blue", 285 prompt tokens confirming real tokenization), and separately with a remote URL using the exact message shape our e2e sends.

tool_choice coverage

auto and none are passed through — none correctly returns finish_reason: "stop" with no tool calls and a clean refusal (3/3). Two modes are imperfect upstream and therefore downgrade to auto:

  • required is accepted but not enforced: with a tool-inviting prompt it returns a correct tool call, but with a neutral prompt the model simply answers instead of being forced to call one (3/3).
  • A named function choice is ignored — requesting get_time returned a get_weather call (3/3).

Since the upstream disregards both constraints, downgrading them to auto is behaviourally identical and keeps the mapping from implying a guarantee it cannot make. Widening the list is a one-line change once either is enforced upstream.

An earlier /v1/models mismatch (stale context_length: 131072 / modality: "text->text") and a malformed non-streaming finish_reason under required were both reported upstream during review and have since been fixed.

Testing

  • TEST_MODELS="ranoai/gemma-4-31b-it" FULL_MODE=true pnpm test:e2e — 27 files passed, 0 failed; 89 passed. All 11 RanoAI cases green: basic (exercises vision), streaming, tool calls, streaming tool calls, tool-call results, JSON output, JSON output streaming, Responses API single-turn/multi-turn/tool-calls, empty-response DONE handling.
  • CONTEXT_SIZE_TEST=true context-size e2e — passed, driving 14,336 real input tokens through the gateway to confirm the 20480 window.
  • pnpm test:unit — 3088 passed / 2 failed; both pre-existing and unrelated (fallback.spec.ts fails identically on a clean checkout of main; stealth-error-redaction.spec.ts passes in isolation and is order-flaky).
  • pnpm build 17/17, pnpm lint 17/17, pnpm format clean.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features
    • Added RanoAI as a supported LLM provider, including default endpoint routing and streaming transformation support.
    • Updated the UI provider selector to display the RanoAI icon/logo.
    • Added Gemma 4 (RanoAI) model support with configured token limits and enabled capabilities (streaming, vision, tools, and JSON output).
    • Expanded RanoAI API key configuration across environment examples, Helm chart values, and e2e workflow runs.

Copilot AI review requested due to automatic review settings July 27, 2026 20:51
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@coderabbitai

coderabbitai Bot commented Jul 27, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

Walkthrough

RanoAI is registered as an LLM provider with Gemma model support, OpenAI-compatible endpoint and streaming routing, API key configuration, e2e secret wiring, and provider icon mappings.

Changes

RanoAI provider integration

Layer / File(s) Summary
Provider and model registration
packages/models/src/providers.ts, packages/models/src/models/google.ts
Adds RanoAI metadata and configures gemma-4-31b-it with RanoAI capabilities, pricing, limits, and streaming support.
Endpoint and streaming routing
packages/actions/src/get-provider-endpoint.ts, apps/gateway/src/chat/tools/transform-streaming-to-openai.ts
Maps RanoAI to https://api.ranoai.com, the shared chat-completions endpoint, and OpenAI-style stream transformation.
Provider icon registration
packages/shared/src/components/provider-icons.tsx, apps/ui/src/components/provider-keys/provider-logo.ts
Adds the RanoAI SVG icon and registers it in shared and UI logo lookups.
Configuration and e2e wiring
.env.example, .env.unified.example, infra/helm/llmgateway/values.yaml, .github/workflows/e2e.yml
Adds the RanoAI API key configuration and passes its repository secret to e2e shard execution.

Estimated code review effort: 2 (Simple) | ~10 minutes

Possibly related PRs

Sequence Diagram(s)

sequenceDiagram
  participant Client
  participant getProviderEndpoint
  participant RanoAI
  participant transformStreamingToOpenai
  Client->>getProviderEndpoint: Resolve the RanoAI endpoint
  getProviderEndpoint->>RanoAI: Send chat completion request
  RanoAI-->>transformStreamingToOpenai: Return streaming chunks
  transformStreamingToOpenai-->>Client: Return OpenAI-style stream
Loading
🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly matches the main change: adding the RanoAI provider, and is concise and specific.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch add-ranoai-provider

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

steebchen and others added 11 commits August 5, 2026 18:03
Add the RanoAI OpenAI-compatible provider (https://api.ranoai.com/v1)
serving Gemma 4 31B on Furiosa RNGD NPU hardware.

Co-Authored-By: Claude <noreply@anthropic.com>
RanoAI fixed image input and non-streaming tool calls, and expanded the
NPU deployment window from 8K to 20480 tokens.

Verified against the live deployment:
- vision works (data URIs and remote URLs) -> vision: true
- window is 20480 (prompt + max_tokens), enforced with a proper
  OpenAI-style error -> contextSize/maxOutput 20480
- Gemma 4 has no reasoning feature -> reasoning stays false

Tools stay disabled: streaming tool calls emit finish_reason
"tool_calls" with no tool_calls delta, and tool-result follow-ups leak
raw <|channel>thought<channel|> markers into content or drop it to null.

Co-Authored-By: Claude <noreply@anthropic.com>
RanoAI fixed streaming tool calls (a proper tool_calls delta is now
emitted) and tool-result follow-ups (clean content, no <|channel>
template leak, with or without tools re-declared).

Verified against the live deployment: auto tool_choice, parallel
multi-tool calls, streaming deltas and tool-result round-trips all
behave correctly.

supportedToolChoices is pinned to ["auto"] because the other modes are
still wrong upstream: "none" emits raw call:name{args} text into
content whenever the model wants a tool, "required" returns
finish_reason "tool_calls" with an empty tool_calls array, and a named
function choice is ignored. Unsupported modes downgrade to auto.

Co-Authored-By: Claude <noreply@anthropic.com>
RanoAI fixed tool_choice "none": it now returns finish_reason "stop"
with no tool_calls and a clean refusal, instead of leaking the raw
call:name{args} template text into content.

Pass "none" through instead of downgrading it to "auto", which was
turning an explicit "do not call tools" into a tool call.

"required" (empty tool_calls array with finish_reason "tool_calls") and
named function choices (get_time request yields get_weather) are still
wrong upstream, so they continue to downgrade to "auto".

Co-Authored-By: Claude <noreply@anthropic.com>
"required" is accepted rather than rejected: it behaves like "auto" and
only misreports finish_reason as "tool_calls" (with an empty tool_calls
array) when the model declines to call a tool. With a tool-inviting
prompt it returns a correct tool call.

The previous comment implied the empty array was unconditional, which
overstated the defect. No behaviour change — downgrading "required" and
named choices to "auto" remains correct, since it leaves tool selection
identical while restoring an accurate finish_reason.

Co-Authored-By: Claude <noreply@anthropic.com>
Narrow the description of the tool_choice "required" defect: it is
specific to the non-streaming response path, and the message carries no
tool_calls key at all rather than an empty array. Streaming returns the
correct "stop" for an identical request.

Co-Authored-By: Claude <noreply@anthropic.com>
RanoAI fixed the non-streaming finish_reason defect: "required" with a
prompt that needs no tool now returns "stop" instead of "tool_calls" on
a message with no tool_calls key. /v1/models also now reports
context_length 20480 and modality text+image->text.

"required" and named function choices are still not enforced, so they
continue to downgrade to "auto" — behaviourally identical, since the
upstream disregards both constraints either way.

Co-Authored-By: Claude <noreply@anthropic.com>
Verified against the live deployment: strict json_schema conforms to a
nested schema with enums and additionalProperties:false (3/3, identical
output), and n=3 returns three distinct correctly-indexed choices with
the prompt billed once and output summed across choices. Streaming n
also emits index 0/1/2.

Also correct the reasoning note: the deepinfra, together-ai, cerebras
and runware deployments of this same model do emit reasoning, so the
absence here is a deployment gap rather than a Gemma 4 limitation.
RanoAI accepts reasoning_effort but returns no reasoning field and no
reasoning_tokens.

Co-Authored-By: Claude <noreply@anthropic.com>
RanoAI now emits reasoning as `reasoning_content`, streamed as deltas,
across all seven effort tiers (none/minimal/low/medium/high/xhigh/max);
an invalid tier is rejected with a validation error.

Reasoning only appears when reasoning_effort is passed explicitly — a
request without it returns none — and "none" suppresses it, so the
mapping declares the full reasoningEfforts list. Usage still does not
break out reasoning_tokens.

This also fixes a real defect: with reasoning:false the streaming
transform folds reasoning_content into content, so the reasoning text
was leaking into the assistant message.

Co-Authored-By: Claude <noreply@anthropic.com>
RanoAI now reports reasoning_tokens inside completion_tokens_details.
The streaming transform hoists that to a top-level reasoning_tokens, and
costs.ts adds reasoning on top of completion tokens for any provider not
in the completionIncludesReasoning list — but RanoAI already counts
reasoning inside completion_tokens.

Measured on the live deployment, reasoning is 81-86% of completion
tokens, so a streaming reasoning request was billed roughly 1.8x its
real output (330 output tokens billed as 597).

Adds a regression test, verified to fail without the fix:
expected 0.0001791 to be close to 0.000099.

Co-Authored-By: Claude <noreply@anthropic.com>
RanoAI now enforces every tool_choice mode, so the ["auto","none"]
restriction is dropped and requests pass through unchanged.

Verified on the live deployment: "required" forces a tool call on five
varied prompts that need no tool (5/5, previously 0/5), a named function
choice returns exactly the requested function on five weather-biased
prompts (5/5, previously always get_weather), and "none"/"auto" are
unchanged with no raw call: template text leaking into content.

logprobs are now returned as well. Prompt caching is still not offered.

Co-Authored-By: Claude <noreply@anthropic.com>
@steebchen
steebchen force-pushed the add-ranoai-provider branch from 4e610a6 to 3181231 Compare August 5, 2026 17:18
RanoAI now does automatic prefix caching and advertises the rate as
input_cache_read in /v1/models: 5e-8, exactly half the input price.

Verified live: the first request misses, subsequent identical requests
report cached_tokens 2496/2521, and an 8k prompt reuses the shorter
prefix before caching the full 8000.

Without cachedInputPrice the cost engine falls back to inputPrice, so
cached tokens were billed at 2x what RanoAI charges. Adds a regression
test, verified to fail without the price (0.0002496 vs 0.0001248).

The e2e prompt-caching suite now covers this mapping under
TEST_CACHE_MODE.

Co-Authored-By: Claude <noreply@anthropic.com>
@steebchen
steebchen merged commit 0c749d1 into main Aug 7, 2026
27 checks passed
@steebchen
steebchen deleted the add-ranoai-provider branch August 7, 2026 13:21
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants