add Sarvam AI provider (chat, text-to-speech, speech-to-text) - #5068
Conversation
|
Caution Review failedThe pull request is closed. ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (1)
📝 WalkthroughSummary by CodeRabbit
WalkthroughAdds the Sarvam AI provider with OpenAI-compatible chat and Responses support, custom speech and transcription APIs, provider registration, CI and configuration wiring, UI metadata, tests, and documentation. ChangesSarvam AI Provider Integration
Estimated code review effort: 4 (Complex) | ~60 minutes Sequence Diagram(s)sequenceDiagram
participant Client
participant SarvamProvider
participant SarvamAPI
Client->>SarvamProvider: ChatCompletion or Responses
SarvamProvider->>SarvamAPI: POST /v1/chat/completions
SarvamAPI-->>SarvamProvider: Chat JSON
SarvamProvider-->>Client: Bifrost response
Client->>SarvamProvider: Speech or SpeechStream
SarvamProvider->>SarvamAPI: POST /text-to-speech or /text-to-speech/stream
SarvamAPI-->>SarvamProvider: Audio payload
SarvamProvider-->>Client: Bifrost audio response
Client->>SarvamProvider: Transcription
SarvamProvider->>SarvamAPI: Multipart POST /speech-to-text
SarvamAPI-->>SarvamProvider: Transcript and timing data
SarvamProvider-->>Client: Normalized transcription
Possibly related PRs
Suggested reviewers: 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Confidence Score: 5/5This looks safe to merge.
Important Files Changed
Reviews (10): Last reviewed commit: "Merge branch 'dev' into feat/sarvam-prov..." | Re-trigger Greptile |
There was a problem hiding this comment.
Actionable comments posted: 1
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
.github/workflows/release-pipeline.yml (1)
155-185: 🎯 Functional Correctness | 🟠 Major | ⚡ Quick winAllow
api.sarvam.ai:443in the Sarvam test jobsAdd
api.sarvam.ai:443to theallowed-endpointsblocks in.github/workflows/release-pipeline.yml:155-185, 947-992, 1079-1115.test-coreand bothtest-docker-image-*jobs setSARVAM_API_KEY, and the Sarvam client targetshttps://api.sarvam.ai, so harden-runner will block those calls otherwise.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In @.github/workflows/release-pipeline.yml around lines 155 - 185, Add api.sarvam.ai:443 to the allowed-endpoints lists used by the Sarvam test jobs in release-pipeline.yml so hardened runner permits the Sarvam API calls. Update the allowed-endpoints blocks for test-core and both test-docker-image jobs, keeping the new endpoint alongside the other API host entries referenced by those job definitions and the SARVAM_API_KEY usage.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@core/providers/sarvam/sarvam.go`:
- Around line 150-172: ChatCompletionStream is hardcoding the provider
identifier instead of using the provider’s configured key. Update
SarvamProvider.ChatCompletionStream to pass provider.GetProviderKey() into
openai.HandleOpenAIChatCompletionStreaming, matching ChatCompletion and
preserving correct behavior for custom base_provider_type configurations. Keep
the change localized to the ChatCompletionStream call site and avoid using the
literal schemas.Sarvam there.
---
Outside diff comments:
In @.github/workflows/release-pipeline.yml:
- Around line 155-185: Add api.sarvam.ai:443 to the allowed-endpoints lists used
by the Sarvam test jobs in release-pipeline.yml so hardened runner permits the
Sarvam API calls. Update the allowed-endpoints blocks for test-core and both
test-docker-image jobs, keeping the new endpoint alongside the other API host
entries referenced by those job definitions and the SARVAM_API_KEY usage.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: CHILL
Plan: Pro Plus
Run ID: e18b615c-568f-4b30-b608-083628b5c30b
📒 Files selected for processing (21)
.github/workflows/pr-tests.yml.github/workflows/release-pipeline.ymlcore/bifrost.gocore/internal/llmtests/account.gocore/providers/sarvam/cachedcontents.gocore/providers/sarvam/errors.gocore/providers/sarvam/sarvam.gocore/providers/sarvam/sarvam_test.gocore/providers/sarvam/speech.gocore/providers/sarvam/transcription.gocore/providers/sarvam/types.gocore/schemas/bifrost.gocore/utils.godocs/docs.jsondocs/openapi/openapi.jsondocs/providers/supported-providers/overview.mdxdocs/providers/supported-providers/sarvam.mdxtransports/config.schema.jsonui/lib/constants/config.tsui/lib/constants/icons.tsxui/lib/constants/logs.ts
There was a problem hiding this comment.
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
core/providers/sarvam/sarvam.go (1)
106-128: 🎯 Functional Correctness | 🟠 Major | ⚡ Quick winUse
GetPathFromContextfor streaming chat requests.
ChatCompletionStreamhardcodes/v1/chat/completions, soBifrostContextKeyURLPathis ignored for streaming while the unary path honors it.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@core/providers/sarvam/sarvam.go` around lines 106 - 128, ChatCompletionStream is hardcoding the chat completions path, so the context-derived URL path override is ignored for streaming. Update SarvamProvider.ChatCompletionStream to use the same path resolution as the unary chat flow by calling GetPathFromContext with the BifrostContext and falling back to the default completions path only when no override is present. Keep the change localized in ChatCompletionStream and preserve the existing HandleOpenAIChatCompletionStreaming call structure.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Outside diff comments:
In `@core/providers/sarvam/sarvam.go`:
- Around line 106-128: ChatCompletionStream is hardcoding the chat completions
path, so the context-derived URL path override is ignored for streaming. Update
SarvamProvider.ChatCompletionStream to use the same path resolution as the unary
chat flow by calling GetPathFromContext with the BifrostContext and falling back
to the default completions path only when no override is present. Keep the
change localized in ChatCompletionStream and preserve the existing
HandleOpenAIChatCompletionStreaming call structure.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: CHILL
Plan: Pro Plus
Run ID: b9fda956-d6d1-404e-abc1-f1a86eb779d9
📒 Files selected for processing (4)
core/providers/sarvam/sarvam.gocore/providers/sarvam/speech.gocore/providers/sarvam/transcription.gocore/providers/sarvam/types.go
🚧 Files skipped from review as they are similar to previous changes (3)
- core/providers/sarvam/transcription.go
- core/providers/sarvam/types.go
- core/providers/sarvam/speech.go
|
@Purvi09 can you share the test report output here please? |
|
@akshaydeo here is the test report Test ReportDashboard
Live gateway tests (localhost:8080)curl -s http://localhost:8080/v1/chat/completions -H "Content-Type: application/json" \
-d '{"model":"sarvam/sarvam-30b","messages":[{"role":"user","content":"Namaste, reply in one short sentence"}]}'curl -s http://localhost:8080/v1/chat/completions -H "Content-Type: application/json" \
-d '{"model":"sarvam/sarvam-105b","messages":[{"role":"user","content":"Hi in one word"}]}'curl -s http://localhost:8080/v1/responses -H "Content-Type: application/json" \
-d '{"model":"sarvam/sarvam-30b","input":"Say hello in one word"}'C4 — Text-to-Speech (Bulbul) · curl -s http://localhost:8080/v1/audio/speech -H "Content-Type: application/json" \
-d '{"model":"sarvam/bulbul:v2","input":"Namaste, aap kaise hain?","voice":"anushka","target_language_code":"hi-IN"}' \
--output tts.wav && file tts.wavC5 — Speech-to-Text (Saaras) · curl -s http://localhost:8080/v1/audio/transcriptions -F "model=sarvam/saaras:v3" -F "file=@tts.wav"Summary
|
|
|
| } | ||
|
|
||
| // Speech performs a text-to-speech request to Sarvam's API. | ||
| func (provider *SarvamProvider) Speech(ctx *schemas.BifrostContext, key schemas.Key, request *schemas.BifrostSpeechRequest) (*schemas.BifrostSpeechResponse, *schemas.BifrostError) { |
There was a problem hiding this comment.
you can shift this method to sarvam.go - to maintain parity with the existing conventions
| } | ||
|
|
||
| // Transcription performs a speech-to-text request to Sarvam's API using multipart/form-data. | ||
| func (provider *SarvamProvider) Transcription(ctx *schemas.BifrostContext, key schemas.Key, request *schemas.BifrostTranscriptionRequest) (*schemas.BifrostTranscriptionResponse, *schemas.BifrostError) { |
There was a problem hiding this comment.
same, can shift this method to sarvam.go - to maintain parity with the existing conventions
| // SarvamError models Sarvam's error responses. | ||
| type SarvamError struct { | ||
| Error *sarvamErrorBody `json:"error"` | ||
| Detail json.RawMessage `json:"detail"` |
There was a problem hiding this comment.
we can keep this as a combined struct string and []string and unmarshal whichever is present, instead of unmarshaling at runtime in .Message() (we already do something similar for some other provider ig)
|
Hey @Purvi09 I have added some comments |
There was a problem hiding this comment.
🧹 Nitpick comments (1)
core/providers/sarvam/sarvam.go (1)
227-229: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low valueMinor: error string starts with capital per Go convention.
"Sarvam text-to-speech response contained no audio"starts with a capital letter. While "Sarvam" is a proper noun, Go convention prefers lowercase error strings. Consider rephrasing to keep the convention.✏️ Suggested rewording
- return nil, providerUtils.EnrichError(ctx, providerUtils.NewBifrostOperationError("Sarvam text-to-speech response contained no audio", nil), jsonData, body, provider.sendBackRawRequest, provider.sendBackRawResponse, latency) + return nil, providerUtils.EnrichError(ctx, providerUtils.NewBifrostOperationError("no audio in Sarvam text-to-speech response", nil), jsonData, body, provider.sendBackRawRequest, provider.sendBackRawResponse, latency)🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@core/providers/sarvam/sarvam.go` around lines 227 - 229, Update the error message passed to NewBifrostOperationError in the Sarvam response validation to begin with a lowercase word while preserving the existing provider context and no-audio meaning.Source: Coding guidelines
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Nitpick comments:
In `@core/providers/sarvam/sarvam.go`:
- Around line 227-229: Update the error message passed to
NewBifrostOperationError in the Sarvam response validation to begin with a
lowercase word while preserving the existing provider context and no-audio
meaning.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: CHILL
Plan: Pro Plus
Run ID: 0029d524-2845-4c10-bd05-a2d31d646db9
📒 Files selected for processing (1)
core/providers/sarvam/sarvam.go
|
Hi @Pratham-Mishra04, thanks for the review! I've addressed all the comments: 1)Moved Speech() and Transcription() into sarvam.go
|
|
|
||
| // ListModels is not supported by the Sarvam provider. | ||
| func (provider *SarvamProvider) ListModels(ctx *schemas.BifrostContext, keys []schemas.Key, request *schemas.BifrostListModelsRequest) (*schemas.BifrostListModelsResponse, *schemas.BifrostError) { | ||
| return nil, providerUtils.NewUnsupportedOperationError(schemas.ListModelsRequest, provider.GetProviderKey()) |
There was a problem hiding this comment.
This prs implements list models, can you pls cross check?
|
Hi @Pratham-Mishra04 - added ListModels. /v1/models isn't in Sarvam's docs but is live and OpenAI-shaped, so it delegates to the shared OpenAI handler. Verified live (returns sarvam-30b, sarvam-105b), updated the support matrix and enabled the ListModels test. |
|
awesome! reviewing |
* add Sarvam AI provider (chat, text-to-speech, speech-to-text) * fix: address review feedback on Sarvam provider * fix: enrich Sarvam STT decode/parse error paths with raw diagnostics * refactor+feat: address review feedback on Sarvam provider * fix: enrich Sarvam TTS malformed-response error paths * feat: implement Sarvam ListModels --------- Co-authored-by: Akshay Deo <akshay@akshaydeo.com> Co-authored-by: Pratham Mishra <99235987+Pratham-Mishra04@users.noreply.github.com>
…q#5068) * add Sarvam AI provider (chat, text-to-speech, speech-to-text) * fix: address review feedback on Sarvam provider * fix: enrich Sarvam STT decode/parse error paths with raw diagnostics * refactor+feat: address review feedback on Sarvam provider * fix: enrich Sarvam TTS malformed-response error paths * feat: implement Sarvam ListModels --------- Co-authored-by: Akshay Deo <akshay@akshaydeo.com> Co-authored-by: Pratham Mishra <99235987+Pratham-Mishra04@users.noreply.github.com>
…q#5068) * add Sarvam AI provider (chat, text-to-speech, speech-to-text) * fix: address review feedback on Sarvam provider * fix: enrich Sarvam STT decode/parse error paths with raw diagnostics * refactor+feat: address review feedback on Sarvam provider * fix: enrich Sarvam TTS malformed-response error paths * feat: implement Sarvam ListModels --------- Co-authored-by: Akshay Deo <akshay@akshaydeo.com> Co-authored-by: Pratham Mishra <99235987+Pratham-Mishra04@users.noreply.github.com>









Summary
Adds Sarvam AI as a built-in provider, closing #5051. Sarvam is a voice/LLM
provider focused on Indian languages (10 Indic languages + English), so this gives
Bifrost users a single integration for both text (chat) and voice (TTS/STT)
workloads in that language segment.
Changes
core/providers/sarvam/covering three capabilities:/v1/chat/completionsis OpenAI-compatible, sothis delegates to the shared
openai.*handlers (base URLhttps://api.sarvam.ai,Authorization: Bearer). Mirrors the Cerebras pattern.audio in an
audios[]array (not raw binary), which is decoded to bytes. Indic fields(
target_language_code,speaker,pace,dict_id, …) are mapped from the request /ExtraParams.transcript/language_code/timestamps/diarized_transcriptmapped onto Bifrost'stranscription shape.
StandardProviders(core/schemas/bifrost.go),registration in
createBaseProvider(core/bifrost.go),dynamicallyConfigurableProviders(
core/utils.go).transports/config.schema.json(provider +base_provider_typeenum),docs/openapi/openapi.json(ModelProviderenum).ui/lib/constants/— provider label, model placeholder, key-requiredflag, and brand icon.
docs/providers/supported-providers/sarvam.mdx, nav entry indocs/docs.json,and a row in the provider support matrix (
overview.mdx).core/providers/sarvam/sarvam_test.go, test-harness keywiring (
core/internal/llmtests/account.go), andSARVAM_API_KEYenv references inpr-tests.yml/release-pipeline.yml.Notable design decisions / trade-offs
voice required hand-written mapping because Sarvam's TTS/STT are not OpenAI-compatible
and use a different auth header (
api-subscription-key).SimpleChat,MultiTurnConversation),which pass in the comprehensive harness. Streaming and voice are verified manually (see
below) but left out of the automated suite for concrete reasons documented in the test:
harness's 500-chunk safety cap.
target_language_codefor TTS.endpoint used here); the mapping handles them defensively and this is noted in the docs.
Type of change
Affected areas
How to test
All three capabilities were verified end-to-end against the live Sarvam API.
Automated test (chat)
UI
cd ui npm i npm run typecheck npm run buildManual end-to-end (gateway)
Configure a Sarvam key (dashboard, or config.json with
"value": "env.SARVAM_API_KEY"), then:Expected: a chat completion, a valid WAV file, and a JSON transcript respectively.
New env var
SARVAM_API_KEY— Sarvam API key. Used by config examples (env.SARVAM_API_KEY) and CI.The same key works for both chat (Bearer) and voice (
api-subscription-key); Bifrost appliesthe correct header per operation. Maintainers: please add the
SARVAM_API_KEYsecret inrepo settings so the provider's tests run in CI (until then they skip, not fail).
Screenshots/Recordings
Breaking changes
Related issues
Closes #5051
Security considerations
SecretVar/env.mechanismand is never logged.
Authorization: Bearerfor chat,api-subscription-keyfor voice — appliedserver-side per request.
Checklist
docs/contributing/README.mdand followed the guidelines