Repository navigation
feat(models): add RanoAI provider - #3265
Conversation
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
|
Note Reviews pausedIt looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the Use the following commands to manage reviews:
Use the checkboxes below for quick actions:
WalkthroughRanoAI is registered as an LLM provider with Gemma model support, OpenAI-compatible endpoint and streaming routing, API key configuration, e2e secret wiring, and provider icon mappings. ChangesRanoAI provider integration
Estimated code review effort: 2 (Simple) | ~10 minutes Possibly related PRs
Sequence Diagram(s)sequenceDiagram
participant Client
participant getProviderEndpoint
participant RanoAI
participant transformStreamingToOpenai
Client->>getProviderEndpoint: Resolve the RanoAI endpoint
getProviderEndpoint->>RanoAI: Send chat completion request
RanoAI-->>transformStreamingToOpenai: Return streaming chunks
transformStreamingToOpenai-->>Client: Return OpenAI-style stream
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
48f22ac to
21e3877
Compare
2b3db83 to
355bcef
Compare
881f048 to
7e7c1b6
Compare
Add the RanoAI OpenAI-compatible provider (https://api.ranoai.com/v1) serving Gemma 4 31B on Furiosa RNGD NPU hardware. Co-Authored-By: Claude <noreply@anthropic.com>
RanoAI fixed image input and non-streaming tool calls, and expanded the NPU deployment window from 8K to 20480 tokens. Verified against the live deployment: - vision works (data URIs and remote URLs) -> vision: true - window is 20480 (prompt + max_tokens), enforced with a proper OpenAI-style error -> contextSize/maxOutput 20480 - Gemma 4 has no reasoning feature -> reasoning stays false Tools stay disabled: streaming tool calls emit finish_reason "tool_calls" with no tool_calls delta, and tool-result follow-ups leak raw <|channel>thought<channel|> markers into content or drop it to null. Co-Authored-By: Claude <noreply@anthropic.com>
RanoAI fixed streaming tool calls (a proper tool_calls delta is now
emitted) and tool-result follow-ups (clean content, no <|channel>
template leak, with or without tools re-declared).
Verified against the live deployment: auto tool_choice, parallel
multi-tool calls, streaming deltas and tool-result round-trips all
behave correctly.
supportedToolChoices is pinned to ["auto"] because the other modes are
still wrong upstream: "none" emits raw call:name{args} text into
content whenever the model wants a tool, "required" returns
finish_reason "tool_calls" with an empty tool_calls array, and a named
function choice is ignored. Unsupported modes downgrade to auto.
Co-Authored-By: Claude <noreply@anthropic.com>
RanoAI fixed tool_choice "none": it now returns finish_reason "stop"
with no tool_calls and a clean refusal, instead of leaking the raw
call:name{args} template text into content.
Pass "none" through instead of downgrading it to "auto", which was
turning an explicit "do not call tools" into a tool call.
"required" (empty tool_calls array with finish_reason "tool_calls") and
named function choices (get_time request yields get_weather) are still
wrong upstream, so they continue to downgrade to "auto".
Co-Authored-By: Claude <noreply@anthropic.com>
"required" is accepted rather than rejected: it behaves like "auto" and only misreports finish_reason as "tool_calls" (with an empty tool_calls array) when the model declines to call a tool. With a tool-inviting prompt it returns a correct tool call. The previous comment implied the empty array was unconditional, which overstated the defect. No behaviour change — downgrading "required" and named choices to "auto" remains correct, since it leaves tool selection identical while restoring an accurate finish_reason. Co-Authored-By: Claude <noreply@anthropic.com>
Narrow the description of the tool_choice "required" defect: it is specific to the non-streaming response path, and the message carries no tool_calls key at all rather than an empty array. Streaming returns the correct "stop" for an identical request. Co-Authored-By: Claude <noreply@anthropic.com>
RanoAI fixed the non-streaming finish_reason defect: "required" with a prompt that needs no tool now returns "stop" instead of "tool_calls" on a message with no tool_calls key. /v1/models also now reports context_length 20480 and modality text+image->text. "required" and named function choices are still not enforced, so they continue to downgrade to "auto" — behaviourally identical, since the upstream disregards both constraints either way. Co-Authored-By: Claude <noreply@anthropic.com>
Verified against the live deployment: strict json_schema conforms to a nested schema with enums and additionalProperties:false (3/3, identical output), and n=3 returns three distinct correctly-indexed choices with the prompt billed once and output summed across choices. Streaming n also emits index 0/1/2. Also correct the reasoning note: the deepinfra, together-ai, cerebras and runware deployments of this same model do emit reasoning, so the absence here is a deployment gap rather than a Gemma 4 limitation. RanoAI accepts reasoning_effort but returns no reasoning field and no reasoning_tokens. Co-Authored-By: Claude <noreply@anthropic.com>
RanoAI now emits reasoning as `reasoning_content`, streamed as deltas, across all seven effort tiers (none/minimal/low/medium/high/xhigh/max); an invalid tier is rejected with a validation error. Reasoning only appears when reasoning_effort is passed explicitly — a request without it returns none — and "none" suppresses it, so the mapping declares the full reasoningEfforts list. Usage still does not break out reasoning_tokens. This also fixes a real defect: with reasoning:false the streaming transform folds reasoning_content into content, so the reasoning text was leaking into the assistant message. Co-Authored-By: Claude <noreply@anthropic.com>
RanoAI now reports reasoning_tokens inside completion_tokens_details. The streaming transform hoists that to a top-level reasoning_tokens, and costs.ts adds reasoning on top of completion tokens for any provider not in the completionIncludesReasoning list — but RanoAI already counts reasoning inside completion_tokens. Measured on the live deployment, reasoning is 81-86% of completion tokens, so a streaming reasoning request was billed roughly 1.8x its real output (330 output tokens billed as 597). Adds a regression test, verified to fail without the fix: expected 0.0001791 to be close to 0.000099. Co-Authored-By: Claude <noreply@anthropic.com>
RanoAI now enforces every tool_choice mode, so the ["auto","none"] restriction is dropped and requests pass through unchanged. Verified on the live deployment: "required" forces a tool call on five varied prompts that need no tool (5/5, previously 0/5), a named function choice returns exactly the requested function on five weather-biased prompts (5/5, previously always get_weather), and "none"/"auto" are unchanged with no raw call: template text leaking into content. logprobs are now returned as well. Prompt caching is still not offered. Co-Authored-By: Claude <noreply@anthropic.com>
4e610a6 to
3181231
Compare
RanoAI now does automatic prefix caching and advertises the rate as input_cache_read in /v1/models: 5e-8, exactly half the input price. Verified live: the first request misses, subsequent identical requests report cached_tokens 2496/2521, and an 8k prompt reuses the shorter prefix before caching the full 8000. Without cachedInputPrice the cost engine falls back to inputPrice, so cached tokens were billed at 2x what RanoAI charges. Adds a regression test, verified to fail without the price (0.0002496 vs 0.0001248). The e2e prompt-caching suite now covers this mapping under TEST_CACHE_MODE. Co-Authored-By: Claude <noreply@anthropic.com>
Adds RanoAI (
ranoai) — an OpenAI-compatible inference provider serving Gemma 4 31B on Furiosa RNGD NPU hardware, viahttps://api.ranoai.com/v1/chat/completions.Changes
packages/models/src/providers.ts):ranoai, API key viaLLM_RANOAI_API_KEY, streaming + cancellation, full metadata (website, terms, privacy, HQUS, data policy).packages/actions/src/get-provider-endpoint.ts): default base URLhttps://api.ranoai.com, routed to/v1/chat/completions.apps/gateway/src/chat/tools/transform-streaming-to-openai.ts): added to the OpenAI-compatible passthrough branch.packages/models/src/models/google.ts): newranoaimapping on the existinggemma-4-31b-itmodel (externalId: gemma-4-31b).packages/shared/src/components/provider-icons.tsx,apps/ui/.../provider-logo.ts):RanoAIIconregistered inProviderIcons+ bothproviderLogoUrlsmaps..env.example,.env.unified.example, Helmvalues.yaml, and the e2e CI workflow.Pricing
Provider list price, passed through as-is: $0.10/M input, $0.30/M output — currently the cheapest Gemma 4 31B mapping in the catalogue.
Capability flags — measured against the live deployment
Every flag was verified by probing the deployment directly rather than reading the docs. RanoAI shipped several fixes during review (image input, tool-call formatting including streaming deltas and tool-result round-trips,
tool_choice: "none", and a window expansion from 8K to 20480), so this reflects the state as of 2026-07-28.prompt + max_tokens), rejected with a proper OpenAI-style errorcontextSize: 20480,maxOutput: 20480vision: truetool_callsnon-streaming and streaming; parallel multi-tool works; tool-result round-trips cleantools: truejson_schemaconforms to nested enum schemas (3/3)jsonOutputSchema: truen=3returns 3 distinct indexed choices, streaming includedsupportsN: truetool_choiceautoandnonehonoured;requiredaccepted but not enforced; named choices ignored (see below)supportedToolChoices: ["auto", "none"]stream_options.include_usagestreaming: truejson_objectandjson_schemaboth return valid JSONjsonOutput: truereasoning_contentacross all 7 effort tiers, streamed as deltas; only emitted whenreasoning_effortis passedreasoning: true+ fullreasoningEffortsVision was confirmed with a synthetic image the model could not have guessed (a 64×64 half-red/half-blue PNG, correctly described as "red and blue", 285 prompt tokens confirming real tokenization), and separately with a remote URL using the exact message shape our e2e sends.
tool_choicecoverageautoandnoneare passed through —nonecorrectly returnsfinish_reason: "stop"with no tool calls and a clean refusal (3/3). Two modes are imperfect upstream and therefore downgrade toauto:requiredis accepted but not enforced: with a tool-inviting prompt it returns a correct tool call, but with a neutral prompt the model simply answers instead of being forced to call one (3/3).get_timereturned aget_weathercall (3/3).Since the upstream disregards both constraints, downgrading them to
autois behaviourally identical and keeps the mapping from implying a guarantee it cannot make. Widening the list is a one-line change once either is enforced upstream.An earlier
/v1/modelsmismatch (stalecontext_length: 131072/modality: "text->text") and a malformed non-streamingfinish_reasonunderrequiredwere both reported upstream during review and have since been fixed.Testing
TEST_MODELS="ranoai/gemma-4-31b-it" FULL_MODE=true pnpm test:e2e— 27 files passed, 0 failed; 89 passed. All 11 RanoAI cases green: basic (exercises vision), streaming, tool calls, streaming tool calls, tool-call results, JSON output, JSON output streaming, Responses API single-turn/multi-turn/tool-calls, empty-response DONE handling.CONTEXT_SIZE_TEST=truecontext-size e2e — passed, driving 14,336 real input tokens through the gateway to confirm the 20480 window.pnpm test:unit— 3088 passed / 2 failed; both pre-existing and unrelated (fallback.spec.tsfails identically on a clean checkout ofmain;stealth-error-redaction.spec.tspasses in isolation and is order-flaky).pnpm build17/17,pnpm lint17/17,pnpm formatclean.🤖 Generated with Claude Code
Summary by CodeRabbit