Skip to content

chunking_strategy as extra params for openai models - #4741

Merged
akshaydeo merged 2 commits into
devfrom
06-27-chunking_strategy_as_extra_params_for_openai_models
Jun 28, 2026
Merged

chunking_strategy as extra params for openai models#4741
akshaydeo merged 2 commits into
devfrom
06-27-chunking_strategy_as_extra_params_for_openai_models

Conversation

@akshaydeo

@akshaydeo akshaydeo commented Jun 27, 2026

Copy link
Copy Markdown
Contributor

Summary

Adds support for forwarding ExtraParams (provider-specific passthrough fields) through the OpenAI transcription multipart form body. This unblocks use of OpenAI diarization models that require chunking_strategy to be present in the multipart request. Closes #4720

Changes

  • ExtraParams from TranscriptionParameters are now copied onto OpenAITranscriptionRequest and written as multipart form fields in ParseTranscriptionFormDataBodyFromRequest. String values are written verbatim; object/map values are JSON-encoded since multipart fields are strings. Keys are sorted for deterministic output.
  • The HTTP transport's parseTranscriptionMultipartRequest now reads chunking_strategy from the incoming multipart form and stores it in ExtraParams, decoding it as a JSON object when possible or passing it through as a plain string (e.g. "auto").
  • Tests cover both the string ("auto") and object (server_vad with threshold) variants of chunking_strategy, verifying correct field values and that the file part remains last.

Type of change

  • Bug fix
  • Feature
  • Refactor
  • Documentation
  • Chore/CI

Affected areas

  • Core (Go)
  • Transports (HTTP)
  • Providers/Integrations
  • Plugins
  • UI (React)
  • Docs

How to test

go test ./core/providers/openai/... ./transports/bifrost-http/...

To validate end-to-end, send a transcription request targeting an OpenAI diarization model (e.g. gpt-4o-transcribe-diarize) with chunking_strategy set to either "auto" or a server_vad object and confirm the field is forwarded correctly in the outgoing multipart request.

Screenshots/Recordings

N/A

Breaking changes

  • No

Related issues

N/A

Security considerations

ExtraParams values are caller-supplied and written directly into the multipart body. No secrets or PII are introduced by this change; callers should ensure they do not pass sensitive data through ExtraParams.

Checklist

  • I read docs/contributing/README.md and followed the guidelines
  • I added/updated tests where appropriate
  • I updated documentation where needed
  • I verified builds succeed (Go and UI)
  • I verified the CI pipeline passes locally if applicable

@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

@coderabbitai

coderabbitai Bot commented Jun 27, 2026

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: b9ebf984-38b1-4071-b74c-1884cb1ce9f9

📥 Commits

Reviewing files that changed from the base of the PR and between 0c13539 and 8dd6dfc.

📒 Files selected for processing (3)
  • core/providers/openai/transcription.go
  • core/providers/openai/transcription_test.go
  • transports/bifrost-http/integrations/openai.go

📝 Walkthrough

Summary by CodeRabbit

  • New Features

    • Transcription requests now support additional provider-specific parameters being passed through end-to-end.
    • chunking_strategy is now accepted in multipart transcription requests and preserved as either plain text or structured JSON.
  • Bug Fixes

    • Improved multipart transcription handling so extra fields are included consistently and in a stable order.
    • Added clearer error handling when transcription form data can’t be written.

Walkthrough

OpenAI transcription requests now preserve provider-specific ExtraParams through multipart parsing, request construction, and outbound multipart encoding. chunking_strategy is accepted as either a raw string or a JSON object, and tests cover both forms.

Changes

OpenAI transcription extra params

Layer / File(s) Summary
Ingest and propagate ExtraParams
transports/bifrost-http/integrations/openai.go, core/providers/openai/transcription.go
parseTranscriptionMultipartRequest stores chunking_strategy in transcriptionReq.ExtraParams, and ToOpenAITranscriptionRequest copies params.ExtraParams onto the OpenAI transcription request.
Encode extra multipart fields
core/providers/openai/transcription.go, core/providers/openai/transcription_test.go
ParseTranscriptionFormDataBodyFromRequest writes sorted ExtraParams fields before the file part, keeps string values unchanged, JSON-encodes other values, and the tests cover string and object chunking_strategy values.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Poem

A bunny hopped through multipart hay,
with chunking_strategy tucked away.
String or JSON, both shine bright,
in little fields that feel just right.
🐇✨

🚥 Pre-merge checks | ✅ 2 | ❌ 3

❌ Failed checks (3 warnings)

Check name Status Explanation Resolution
Linked Issues check ⚠️ Warning The PR implements transcription multipart extra params, but issue #123 requires Files API support such as POST /v1/files. Implement the file upload API workflows required by #123, or relink this PR to the transcription chunking_strategy work instead.
Out of Scope Changes check ⚠️ Warning The code changes are unrelated to the linked Files API objective and focus instead on transcription chunking_strategy handling. Remove the transcription chunking_strategy changes or update the linked issue so the PR scope matches the requested Files API support.
Docstring Coverage ⚠️ Warning Docstring coverage is 57.14% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly describes the main change: forwarding chunking_strategy as OpenAI transcription extra params.
Description check ✅ Passed The PR description covers summary, changes, testing, affected areas, and security, matching the template well.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch 06-27-chunking_strategy_as_extra_params_for_openai_models

Warning

There were issues while running some tools. Please review the errors and either fix the tool's configuration or disable the tool if it's a critical failure.

🔧 golangci-lint (2.12.2)

level=error msg="[linters_context] typechecking error: pattern ./...: directory prefix . does not contain main module or its selected dependencies"


Comment @coderabbitai help to get the list of available commands.

@akshaydeo
akshaydeo marked this pull request as ready for review June 27, 2026 05:53

Copy link
Copy Markdown
Contributor Author

This stack of pull requests is managed by Graphite. Learn more about stacking.

@greptile-apps

greptile-apps Bot commented Jun 27, 2026

Copy link
Copy Markdown
Contributor

Confidence Score: 5/5

This looks safe to merge.

  • No blocking issues found in the changed code.

Important Files Changed

Filename Overview
core/providers/openai/transcription.go Adds transcription ExtraParams forwarding into the OpenAI multipart body.
transports/bifrost-http/integrations/openai.go Reads incoming chunking_strategy multipart values into transcription ExtraParams.
core/providers/openai/transcription_test.go Covers string and object chunking_strategy multipart serialization.

Reviews (2): Last reviewed commit: "Merge branch 'dev' into 06-27-chunking_s..." | Re-trigger Greptile

@akshaydeo
akshaydeo merged commit c4dc970 into dev Jun 28, 2026
12 of 14 checks passed
@akshaydeo
akshaydeo deleted the 06-27-chunking_strategy_as_extra_params_for_openai_models branch June 28, 2026 05:46
R-droid101 pushed a commit to R-droid101/bifrost that referenced this pull request Jul 1, 2026
akshaydeo added a commit that referenced this pull request Jul 6, 2026
* fix: update audit logs page layout classes and add newline at EOF (#4719)

## Summary

Fixes the layout of the Audit Logs page to correctly fill the viewport and apply the appropriate background and border styles.

## Changes

- Replaced `h-[calc(100dvh-1rem)]` with `h-[calc(100vh-16px)]` for consistent viewport height calculation
- Swapped `mx-auto flex flex-col p-4` utility classes for `no-border-parent bg-background flex` to align with the layout conventions used elsewhere in the app
- Added missing newline at end of file

## Type of change

- [x] Bug fix
- [ ] Feature
- [ ] Refactor
- [ ] Documentation
- [ ] Chore/CI

## Affected areas

- [ ] Core (Go)
- [ ] Transports (HTTP)
- [ ] Providers/Integrations
- [ ] Plugins
- [x] UI (React)
- [ ] Docs

## How to test

Navigate to the Audit Logs page and verify:
- The page fills the full viewport height without overflow or clipping
- The background color and border styling match the rest of the workspace layout

```sh
cd ui
pnpm i || npm i
pnpm build || npm run build
```

## Screenshots/Recordings

Add before/after screenshots showing the corrected Audit Logs page layout.

## Breaking changes

- [ ] Yes
- [x] No

## Related issues

## Security considerations

None.

## Checklist

- [ ] I read `docs/contributing/README.md` and followed the guidelines
- [ ] I added/updated tests where appropriate
- [ ] I updated documentation where needed
- [ ] I verified builds succeed (Go and UI)
- [ ] I verified the CI pipeline passes locally if applicable

* feat: add collapsible tag limit to TagInput (#4730)

## Summary

Adds collapsible tag support to the `TagInput` component and applies it to the keyword lists on the Complexity Router page. When a keyword list exceeds a configurable limit, tags beyond that limit are hidden behind a gradient overlay with a "Show more" toggle, keeping the UI compact while still allowing full access to all tags.

## Changes

- Added `collapsedTagLimit` and `expandButtonTestId` props to `TagInput`. When `collapsedTagLimit` is provided, the component renders in a collapsible layout: tags beyond the limit are hidden with a fade gradient, and "Show more" / "Show less" buttons toggle the expanded state. The collapsed state auto-resets when the tag count drops back to or below the limit.
- Set `KEYWORD_COLLAPSED_LIMIT = 8` on the Complexity Router page and passed it along with a `expandButtonTestId` to each keyword `TagInput`.
- Standardized border radius tokens from `rounded-lg`/`rounded-md`/`rounded-full` to `rounded-sm` across the Complexity Router page for visual consistency.
- Reformatted `index.html` inline shell skeleton from a single minified line to readable, indented HTML and CSS.

## Type of change

- [ ] Bug fix
- [x] Feature
- [ ] Refactor
- [ ] Documentation
- [ ] Chore/CI

## Affected areas

- [ ] Core (Go)
- [ ] Transports (HTTP)
- [ ] Providers/Integrations
- [ ] Plugins
- [x] UI (React)
- [ ] Docs

## How to test

```sh
cd ui
pnpm i || npm i
pnpm test || npm test
pnpm build || npm run build
```

1. Navigate to the Complexity Router page.
2. Add more than 8 keywords to any keyword list.
3. Verify that tags beyond 8 are hidden with a gradient overlay and a "Show more" button appears.
4. Click "Show more" and confirm all tags are visible with a "Show less" button.
5. Click "Show less" and confirm the list collapses again.
6. Remove tags until 8 or fewer remain and confirm the list stays expanded without the toggle controls.

## Screenshots/Recordings

Before/after screenshots of the keyword lists with collapse behavior recommended.

## Breaking changes

- [ ] Yes
- [x] No

## Related issues

## Security considerations

None.

## Checklist

- [ ] I read `docs/contributing/README.md` and followed the guidelines
- [ ] I added/updated tests where appropriate
- [ ] I updated documentation where needed
- [ ] I verified builds succeed (Go and UI)
- [ ] I verified the CI pipeline passes locally if applicable

* fix: rebuild token_usage from denormalized columns in hybrid log list (#4722)

* fix: rebuild token_usage from denormalized columns in hybrid log list

* refactor: inline hybrid token usage reconstruction

* fix: preserve malformed serialized token usage state

---------

Co-authored-by: gexiangdong <xiangdong.ge@pandasofcaribbean.com>

* refactor: centralize Anthropic request building into `BuildAnthropicChatRequestBody` and `AnthropicProviderRequestDefaultsMap` (#3309)

## Summary

This PR consolidates Anthropic-family request building across all providers (Anthropic native, Azure, Vertex, Bedrock) into two shared builder functions — `BuildAnthropicChatRequestBody` and `BuildAnthropicResponsesRequestBody` — eliminating duplicated inline logic and provider-specific wrapper helpers that previously scattered the same field-stripping, beta-header injection, and model-field manipulation across multiple files.

## Changes

- Introduced `AnthropicProviderRequestDefaults` and `AnthropicProviderRequestDefaultsMap` to encode static, per-provider request-shaping flags (e.g. `DeleteModelField`, `DeleteStreamField`, `AddAnthropicVersion`, `InjectBetaHeadersIntoBody`) in one place. Callers no longer pass these flags directly; the builder looks them up by `cfg.Provider`.
- Renamed `Deployment` to `Model` in `AnthropicRequestBuildConfig` for clarity, since all providers now use the same field for model/deployment overrides.
- Added `BuildAnthropicChatRequestBody` as the chat-completion analogue of `BuildAnthropicResponsesRequestBody`, covering both raw-body and typed paths, including field stripping, beta-header injection, streaming flag handling, and `fallbacks` deletion.
- Bedrock now routes Anthropic models through the Anthropic Messages API format (`invoke` / `invoke-with-response-stream` endpoints) for both chat and responses, rather than the Bedrock Converse API. This includes proper response parsing via `AcquireAnthropicMessageResponse` and streaming via `AnthropicStreamState` / `AnthropicResponsesStreamState`.
- Removed private wrapper functions `getRequestBodyForResponses` (Anthropic), `getRequestBodyForAnthropicResponses` (Azure, Vertex), and the inline `CheckContextAndGetRequestBody` closures for Anthropic models in Vertex and Azure, replacing all call sites with direct `BuildAnthropicChatRequestBody` / `BuildAnthropicResponsesRequestBody` calls.
- Exported `AcquireAnthropicResponsesStreamState`, `ReleaseAnthropicResponsesStreamState`, `AcquireAnthropicMessageResponse`, and `ReleaseAnthropicMessageResponse` so Bedrock can reuse the Anthropic stream state pool.
- Removed `DefaultVertexAnthropicVersion` constant from the Vertex package; the canonical version string now lives in `AnthropicProviderRequestDefaultsMap`.
- Bedrock's `releaseBedrockChatResponse` now zeroes the struct before returning it to the pool.
- `stripUnsupportedAnthropicFields` is now called inside `BuildAnthropicResponsesRequestBody` on the typed path, making field stripping symmetric across raw and typed paths and across both APIs.

## Type of change

- [ ] Bug fix
- [ ] Feature
- [x] Refactor
- [ ] Documentation
- [ ] Chore/CI

## Affected areas

- [x] Core (Go)
- [ ] Transports (HTTP)
- [x] Providers/Integrations
- [ ] Plugins
- [ ] UI (React)
- [ ] Docs

## How to test

```sh
go test ./core/providers/anthropic/...
go test ./core/providers/azure/...
go test ./core/providers/bedrock/...
go test ./core/providers/vertex/...
go test ./...
```

Run integration tests against Anthropic, Azure (Anthropic models), Vertex (Claude models), and Bedrock (Claude models) for chat completion, streaming, responses, responses streaming, and count-tokens endpoints. Verify that raw-body passthrough requests produce the same field stripping and beta-header injection as typed requests.

## Breaking changes

- [x] Yes
- [ ] No

`AnthropicRequestBuildConfig` has a breaking field rename: `Deployment` → `Model`. Any external code constructing this struct directly must update the field name. The static shaping flags (`DeleteModelField`, `DeleteRegionField`, `AddAnthropicVersion`, `AnthropicVersion`, `StripCacheControlScope`, `RemapToolVersions`, `InjectBetaHeadersIntoBody`) have been removed from `AnthropicRequestBuildConfig` and are now looked up internally via `AnthropicProviderRequestDefaultsMap`; callers that set these fields must remove them.

## Related issues

## Security considerations

No new auth flows, secrets handling, or PII exposure introduced. Field stripping ensures provider-unsupported fields are not forwarded to external APIs.

## Checklist

- [ ] I read `docs/contributing/README.md` and followed the guidelines
- [x] I added/updated tests where appropriate
- [ ] I updated documentation where needed
- [x] I verified builds succeed (Go and UI)
- [ ] I verified the CI pipeline passes locally if applicable

* refactor: extract `HandleAnthropicChatCompletionRequest` / `HandleAnthropicResponsesRequest` and make `completeRequest` a package-level func shared by Anthropic, Azure, and Bedrock providers (#4394)

## Summary

The Anthropic provider's unary request logic was duplicated across the Anthropic, Azure, Bedrock, and Vertex providers. This PR extracts the core non-streaming request execution into a package-level `completeRequest` function and introduces two exported handler functions — `HandleAnthropicChatCompletionRequest` and `HandleAnthropicResponsesRequest` — that encapsulate the full build → send → parse pipeline for chat completions and the Responses API respectively. Azure and Bedrock now delegate directly to these shared handlers for Anthropic-family models instead of reimplementing request dispatch, response parsing, and raw request/response handling inline.

A secondary bug fix is included: the large-response streaming client was being activated for count-tokens requests (which should always be buffered) and skipped for all other requests — the condition was inverted.

## Changes

- Extracted `completeRequest` as a package-level function accepting explicit `client`, `headers`, `extraHeaders`, `betaHeaderOverrides`, `providerName`, and `logger` arguments, removing the method receiver dependency so it can be called by other providers.
- Added `anthropicRequestHeaders` as a provider method to build the `x-api-key` / `anthropic-version` header map, shared across `TextCompletion`, `ChatCompletion`, `Responses`, and `CountTokens`.
- Introduced `HandleAnthropicChatCompletionRequest` and `HandleAnthropicResponsesRequest` as exported functions that perform the full unary request lifecycle (body build, HTTP send, large-response detection, response parse, raw request/response attachment). These are now called by the Anthropic, Azure, and Bedrock providers.
- Removed `completeMantleRequest` from Bedrock — its logic is now covered by `completeRequest` inside the shared handlers.
- Azure's `ChatCompletion` and `Responses` methods now branch early for Anthropic-family models, calling the shared handlers with Azure-specific auth headers, and fall through to the OpenAI-compatible path otherwise, eliminating the post-response model-family branch.
- Fixed the inverted condition in `completeRequest` that caused the large-response streaming client to be used for count-tokens requests instead of being skipped for them.
- `AnthropicRequestBuildConfig` now carries `BetaHeaderOverrides` so callers do not need to pass it separately.

## Type of change

- [ ] Bug fix
- [ ] Feature
- [x] Refactor
- [ ] Documentation
- [ ] Chore/CI

## Affected areas

- [x] Core (Go)
- [ ] Transports (HTTP)
- [x] Providers/Integrations
- [ ] Plugins
- [ ] UI (React)
- [ ] Docs

## How to test

```sh
go test ./...
```

Validate that chat completions and Responses API requests succeed for Anthropic-family models routed through the Azure and Bedrock providers, and that count-tokens requests return buffered responses without triggering large-response mode.

## Breaking changes

- [ ] Yes
- [x] No

## Security considerations

Auth headers (`x-api-key`, Bearer tokens, SigV4-signed headers) are applied last in `completeRequest`, after network-config extra headers, ensuring they cannot be overridden by user-supplied configuration. No new secrets or PII handling paths are introduced.

## Checklist

- [ ] I read `docs/contributing/README.md` and followed the guidelines
- [ ] I added/updated tests where appropriate
- [ ] I updated documentation where needed
- [ ] I verified builds succeed (Go and UI)
- [ ] I verified the CI pipeline passes locally if applicable

* refactor: introduce `BearerAuthHeader` helper and migrate OpenAI-compatible providers from `schemas.Key` to `map[string]string` auth header param (#4425)

## Summary

This PR standardizes how Bearer token authentication headers are constructed across all OpenAI-compatible providers. Previously, each call site independently built the `Authorization: Bearer <token>` header map with duplicated inline logic. A new `BearerAuthHeader(key)` helper is introduced in the OpenAI package and used uniformly everywhere.

Additionally, the Azure provider's private `completeRequest` method is removed. Its non-Anthropic request paths (text completion, chat completion, responses, embedding, compaction) are now delegated directly to the shared `Handle*` functions in the OpenAI package, consistent with how other providers already work. The `Handle*` functions themselves are updated to accept a pre-built `authHeader map[string]string` instead of a raw `schemas.Key`, making them provider-agnostic and compatible with non-Bearer auth schemes (e.g., Azure API key headers, SigV4).

## Changes

- Added `BearerAuthHeader(key schemas.Key) map[string]string` to the OpenAI provider package, which returns an `Authorization: Bearer <token>` header map, or an empty map when the key carries no value.
- Updated all `Handle*Request` and `handleOpenAILargePayloadPassthrough` function signatures to accept `authHeader map[string]string` instead of `schemas.Key`, applying the map directly to request headers.
- Replaced all inline `var authHeader map[string]string` + conditional assignment blocks across Cerebras, Fireworks, Groq, HuggingFace, Mistral, Nebius, Ollama, Opencode, OpenRouter, Parasail, Perplexity, SGL, VLLM, xAI, and Bedrock with calls to `openai.BearerAuthHeader(key)`.
- Removed the Azure provider's `completeRequest` method and replaced its usage in `TextCompletion`, `ChatCompletion`, `Responses`, `Embedding`, and `Compaction` with direct calls to the corresponding shared OpenAI `Handle*` functions, passing Azure-specific auth headers and pre-resolved endpoint URLs.

## Type of change

- [ ] Bug fix
- [ ] Feature
- [x] Refactor
- [ ] Documentation
- [ ] Chore/CI

## Affected areas

- [x] Core (Go)
- [ ] Transports (HTTP)
- [x] Providers/Integrations
- [ ] Plugins
- [ ] UI (React)
- [ ] Docs

## How to test

```sh
go test ./...
```

Verify that all OpenAI-compatible providers (OpenAI, Azure, Cerebras, Fireworks, Groq, HuggingFace, Mistral, Nebius, Ollama, Opencode, OpenRouter, Parasail, Perplexity, SGL, VLLM, xAI, Bedrock Mantle) continue to authenticate correctly and that requests succeed for text completion, chat completion, responses, embeddings, and compaction endpoints.

## Breaking changes

- [ ] Yes
- [x] No

## Security considerations

The `BearerAuthHeader` helper preserves the existing behavior of omitting the `Authorization` header when the key value is empty, which is intentional for providers that supply auth via other mechanisms (e.g., extra headers or SigV4 signing). No secrets are logged or exposed.

## Checklist

- [ ] I read `docs/contributing/README.md` and followed the guidelines
- [ ] I added/updated tests where appropriate
- [ ] I updated documentation where needed
- [ ] I verified builds succeed (Go and UI)
- [ ] I verified the CI pipeline passes locally if applicable

* refactor: replace pre-built SigV4 body signing with lazy `BodySigner` closure passed through request handlers (#4735)

## Summary

Replaces the pre-build-and-sign approach for Bedrock Mantle SigV4 authentication with a `BodySigner` callback that is invoked after the request handler has marshaled the body. This ensures the signature always covers the exact bytes sent on the wire, eliminating the previous double-marshal pattern where the body was built once for signing and again inside the handler.

## Changes

- Introduces a new `BodySigner` type (`func(jsonData []byte) (map[string]string, *schemas.BifrostError)`) in `core/providers/utils/bodysigner.go`. Handlers call it after building the request body and apply the returned headers to the outgoing request.
- Adds the `signer` parameter to `HandleOpenAIChatCompletionRequest`, `HandleOpenAIChatCompletionStreaming`, `HandleOpenAIResponsesRequest`, `HandleOpenAIResponsesStreaming`, `HandleAnthropicChatCompletionRequest`, `HandleAnthropicChatCompletionStreaming`, `HandleAnthropicResponsesRequest`, and `HandleAnthropicResponsesStream`. All existing callers pass `nil`.
- Rewrites Bedrock Mantle's SigV4 paths (`mantleChatCompletions`, `mantleChatCompletionsStream`, `mantleResponses`, `mantleResponsesStream`) to construct a `BodySigner` closure when no API key is present, instead of pre-building the body, signing it, and merging the signature headers into `extraHeaders`. The Bearer path no longer needs a separate early-return branch.
- Removes the now-unnecessary `maps` import and the intermediate `extraHeaders` map copies in the Mantle code paths.

## Type of change

- [ ] Bug fix
- [x] Refactor
- [ ] Feature
- [ ] Documentation
- [ ] Chore/CI

## Affected areas

- [x] Core (Go)
- [ ] Transports (HTTP)
- [x] Providers/Integrations
- [ ] Plugins
- [ ] UI (React)
- [ ] Docs

## How to test

```sh
go test ./...
```

For Bedrock Mantle with SigV4 (empty key value), verify that requests to chat completions, streaming chat completions, responses, and streaming responses are signed correctly and accepted by the Bedrock endpoint. For Bearer key paths, confirm that no signing is attempted and the `Authorization` header is set as expected.

## Breaking changes

- [ ] Yes
- [x] No

## Security considerations

The `BodySigner` callback signs the exact serialized bytes that are placed on the wire. Previously, the body was serialized twice (once for signing, once inside the handler), which could in theory produce a signature mismatch if marshaling were non-deterministic. This change closes that gap by signing after the final body is set.

## Checklist

- [ ] I read `docs/contributing/README.md` and followed the guidelines
- [ ] I added/updated tests where appropriate
- [ ] I updated documentation where needed
- [ ] I verified builds succeed (Go and UI)
- [ ] I verified the CI pipeline passes locally if applicable

* feat: add `bedrock_mantle` as a first-class provider with native-Anthropic and OpenAI-compatible routing (#4736)

## Summary

Introduces `bedrock_mantle` as a first-class, standalone provider that owns the Bedrock Mantle surface (`bedrock-mantle.{region}.api.aws`). Previously, Mantle routing was handled as an internal routing decision inside the existing `bedrock` provider. The new provider gives operators a dedicated configuration surface for Claude (native Anthropic Messages API), OpenAI-compatible models (gpt-*), and Gemma models served through Mantle, without requiring a full Bedrock setup.

## Changes

- Added `schemas.BedrockMantle` (`"bedrock_mantle"`) as a new `ModelProvider` constant and registered it in `StandardProviders`, `dynamicallyConfigurableProviders`, `CanProviderKeyValueBeEmpty`, and `isKeySkippingAllowed`.
- Added `BedrockMantleKeyConfig` to the `Key` struct, carrying AWS credentials and region for SigV4 auth against the `bedrock-mantle` service. The existing `BedrockKeyConfig` is unchanged.
- Introduced the `core/providers/bedrockmantle` package implementing the full `Provider` interface. Chat, streaming chat, Responses, and streaming Responses dispatch by model family: Anthropic-family models use the native Anthropic Messages surface (`/anthropic/v1/messages`); all others use the OpenAI-compatible surface (`/v1` or `/openai/v1`). All other operations return unsupported-operation errors.
- Refactored `signAWSRequest` in the `bedrock` package to accept a `*BedrockKeyConfig` instead of individual credential fields, eliminating the now-redundant `signAWSRequestFromKey` wrapper. All call sites updated accordingly.
- Exported `SignMantleV4Headers` (previously `mantleSigV4Headers`, a method on `BedrockProvider`) so the new `bedrockmantle` package can sign requests without depending on the internal Bedrock provider struct. The function now supports both `BedrockKeyConfig` and `BedrockMantleKeyConfig` by mapping the latter into a synthetic `BedrockKeyConfig` for signing, and correctly handles GET requests (nil body) for the list-models path.
- Extended the Anthropic chat and Responses request builders to convert native structured outputs to tool calls for `BedrockMantle`, matching the existing `Vertex` workaround.
- Added `BedrockMantle` to the comprehensive LLM test harness (`ComprehensiveTestAccount`) with key config, provider config, and a full test file covering the supported scenarios (chat, streaming, tool calls, vision, structured outputs, prompt caching, reasoning, list models) and explicitly disabling unsupported ones.
- Marked `isMantleModel` in `bedrock/mantle.go` as deprecated in favour of the new provider.

## Type of change

- [ ] Bug fix
- [x] Feature
- [ ] Refactor
- [ ] Documentation
- [ ] Chore/CI

## Affected areas

- [x] Core (Go)
- [ ] Transports (HTTP)
- [x] Providers/Integrations
- [ ] Plugins
- [ ] UI (React)
- [ ] Docs

## How to test

Set AWS credentials and run the new provider test:

```sh
export AWS_ACCESS_KEY_ID=...
export AWS_SECRET_ACCESS_KEY=...
export AWS_SESSION_TOKEN=...   # optional, for temporary credentials
export AWS_REGION=us-east-1

go test ./core/providers/bedrockmantle/... -v -run TestBedrockMantle
```

To run the full suite (skips Bedrock Mantle automatically when credentials are absent):

```sh
go test ./...
```

Configure a `bedrock_mantle` provider by supplying a `BedrockMantleKeyConfig` (or a Bearer API key in `Value`) with the desired region. The region can also be embedded as a prefix in the model ID (e.g. `us-west-2/anthropic.claude-haiku-4-5`) or set at the alias level via `AliasConfig.Region`.

## Breaking changes

- [ ] Yes
- [x] No

The `signAWSRequest` signature change is internal to the `bedrock` package and does not affect any public API. The `isMantleModel` function is deprecated but not removed.

## Security considerations

AWS credentials for `BedrockMantleKeyConfig` follow the same `SecretVar` resolution pattern used by `BedrockKeyConfig` (env-var references, never inlined literals). SigV4 signing is performed per-request on the exact body bytes that are sent, so the signature always covers what is transmitted. When a Bearer API key is present it takes precedence and no AWS credentials are required.

## Checklist

- [ ] I read `docs/contributing/README.md` and followed the guidelines
- [x] I added/updated tests where appropriate
- [ ] I updated documentation where needed
- [x] I verified builds succeed (Go and UI)
- [ ] I verified the CI pipeline passes locally if applicable

* feat: add `bedrock_mantle` provider with SigV4 key config, DB migration, and UI support (#4737)

Adds `bedrock_mantle` as a first-class provider, enabling Bifrost to route requests to AWS Bedrock through a Mantle proxy endpoint. The provider supports the same SigV4 credential options as the existing Bedrock provider (inherited IAM role, explicit access/secret key, session token, AssumeRole) as well as a Bearer API key authentication mode.

- Added `BedrockMantle` to the Anthropic passthrough allowlist in `clearAnthropicPassthroughForNonNativeProvider` so raw request bodies are preserved when routing through Bedrock Mantle.
- Added `BedrockMantleKeyConfig` redaction logic in `clientconfig.go`, mirroring the existing Bedrock redaction pattern.
- Added a new `migrationAddBedrockMantleKeyColumns` database migration that introduces seven `bedrock_mantle_*` SigV4 credential columns to the `config_keys` table.
- Extended `TableKey` with the seven Bedrock Mantle credential fields, along with `BeforeSave` serialization and `AfterFind` reconstruction hooks.
- Updated `mergeUpdatedKey` in the HTTP handler to correctly restore redacted Bedrock Mantle credential fields during key updates.
- Fixed `isClaudeModel` in the Anthropic integration to recognize `bedrock_mantle` (previously incorrectly matched `bedrock`) as a provider that can serve Claude models.
- Included `BedrockMantleKeyConfig` in the key hash inputs used by `mergeProviderKeys` and `reconcileProviderKeys` for config file/DB reconciliation.
- Added Bedrock Mantle credential redaction to `GetAllKeys`.
- Extended `config.schema.json` with `bedrock_mantle_key` and `provider_with_bedrock_mantle_config` definitions and registered `bedrock_mantle` as a valid provider name throughout the schema.
- Added UI support: provider icon (reusing the Bedrock SVG mark with a distinct gradient ID), model placeholder text, `isKeyRequiredByProvider` entry, label, form schema (`BedrockMantleKeyConfigSchema`), type definitions (`BedrockMantleKeyConfig`, `DefaultBedrockMantleKeyConfig`), and a full authentication method tab UI (IAM Role / Explicit Credentials / API Key) matching the Bedrock provider UX.
- Added `bedrock_mantle` to the Anthropic beta-headers provider family and the provider config sheet's Anthropic family list.
- Stripped the internal `_auth_type` field from `bedrock_mantle_key_config` before submitting the form payload.

- [ ] Bug fix
- [x] Feature
- [ ] Refactor
- [ ] Documentation
- [ ] Chore/CI

- [x] Core (Go)
- [x] Transports (HTTP)
- [x] Providers/Integrations
- [ ] Plugins
- [x] UI (React)
- [ ] Docs

```sh
go version
go test ./...

cd ui
pnpm i || npm i
pnpm test || npm test
pnpm build || npm run build
```

Configure a `bedrock_mantle` provider in `config.json` or via the UI with one of the three auth methods:

- **IAM Role (Inherited):** set only `region`; leave access/secret key empty.
- **Explicit Credentials:** set `access_key`, `secret_key`, and `region`; optionally set `session_token`, `role_arn`, `external_id`, and `session_name`.
- **API Key:** set `region` and provide a Bearer token as the key `value`.

Send a request targeting a Claude model through the `bedrock_mantle` provider and verify the response is returned correctly and that credentials are redacted in the UI and API responses.

_Add before/after screenshots of the new Bedrock Mantle provider form and icon in the UI._

- [ ] Yes
- [x] No

_Link related issues and discussions._

- All seven Bedrock Mantle credential fields (`access_key`, `secret_key`, `session_token`, `region`, `role_arn`, `external_id`, `role_session_name`) are stored as `SecretVar` and are redacted in API responses and the UI, consistent with the existing Bedrock provider handling.
- The `_auth_type` discriminator field is stripped from the payload before it is persisted or transmitted.

- [ ] I read `docs/contributing/README.md` and followed the guidelines
- [ ] I added/updated tests where appropriate
- [ ] I updated documentation where needed
- [ ] I verified builds succeed (Go and UI)
- [ ] I verified the CI pipeline passes locally if applicable

* feat: add AWS Bedrock Mantle provider docs (#4738)

## Summary

Adds documentation for the AWS Bedrock Mantle provider, a distinct AWS endpoint (`bedrock-mantle.{region}.api.aws`) that exposes Claude models via the native Anthropic Messages API and OpenAI-family/Gemma models via an OpenAI-compatible API — all addressable through a single `bedrock_mantle/<model>` prefix in Bifrost.

## Changes

- Added a new `bedrock-mantle.mdx` provider page covering model ID formats, supported operations, all three authentication modes (SigV4 with explicit credentials, IAM role/inherited credentials, and Bearer API key), IAM role assumption via `role_arn`, and usage examples.
- Added the Bedrock Mantle configuration block to the `providers.mdx` config reference, with tabs for Static Credentials, IAM Role, and API Key (Bearer) auth modes.
- Added Bedrock Mantle to the provider capability matrix in `overview.mdx`.
- Registered `bedrock-mantle` in `docs.json` so it appears in the sidebar navigation.

## Type of change

- [ ] Bug fix
- [ ] Feature
- [ ] Refactor
- [x] Documentation
- [ ] Chore/CI

## Affected areas

- [ ] Core (Go)
- [ ] Transports (HTTP)
- [ ] Providers/Integrations
- [ ] Plugins
- [ ] UI (React)
- [x] Docs

## How to test

Navigate to the Bedrock Mantle provider page and config reference in the rendered docs and verify:

- The sidebar entry for `bedrock-mantle` appears between `bedrock` and `cerebras`.
- All three auth tabs (Static Credentials, IAM Role, API Key) render correctly in both the provider page and the config reference.
- The capability matrix row for `bedrock_mantle/<model>` is present and accurate.
- Cross-links between the provider page and the config reference resolve correctly.

## Breaking changes

- [x] No

## Security considerations

Authentication credentials (`access_key`, `secret_key`, `session_token`, API keys) are documented using the `env.*` indirection pattern, consistent with how other providers handle secrets. No credentials are hardcoded in examples.

## Checklist

- [ ] I read `docs/contributing/README.md` and followed the guidelines
- [ ] I added/updated tests where appropriate
- [x] I updated documentation where needed
- [ ] I verified builds succeed (Go and UI)
- [ ] I verified the CI pipeline passes locally if applicable

* tests: add `bedrock_mantle` provider capabilities and Postman environment config (#4739)

Adds E2E test configuration and capability definitions for the `bedrock_mantle` provider, enabling it to be tested through the Bifrost V1 API test suite.

- Added `bedrock_mantle` to `provider-capabilities.json` with `chat_completions`, `chat_completions_with_tools`, `responses`, `responses_with_tools`, and `list_models` enabled
- Added a new Postman environment file (`bifrost-v1-bedrock-mantle.postman_environment.json`) configured to use `anthropic.claude-opus-4-8` as the default model and `us-east-1` as the default region, with secret placeholders for API key, access key, secret key, and session token

- [ ] Bug fix
- [x] Feature
- [ ] Refactor
- [ ] Documentation
- [ ] Chore/CI

- [ ] Core (Go)
- [ ] Transports (HTTP)
- [x] Providers/Integrations
- [ ] Plugins
- [ ] UI (React)
- [ ] Docs

Run the E2E test suite targeting the `bedrock_mantle` provider using the new Postman environment:

```sh

newman run tests/e2e/api/bifrost-v1.postman_collection.json \
  -e tests/e2e/api/provider_config/bifrost-v1-bedrock-mantle.postman_environment.json \
  --env-var "bedrock_mantle_api_key=<your_api_key>" \
  --env-var "bedrock_mantle_access_key=<your_access_key>" \
  --env-var "bedrock_mantle_secret_key=<your_secret_key>"
```

Expected outcome: chat completions, tool-use, responses, and model listing tests pass; all unsupported capability tests are skipped or return expected errors.

N/A

- [ ] Yes
- [x] No

N/A

The Postman environment file stores API key, access key, secret key, and session token as `secret` type fields with empty default values, ensuring credentials are not committed to the repository.

- [ ] I read `docs/contributing/README.md` and followed the guidelines
- [x] I added/updated tests where appropriate
- [ ] I updated documentation where needed
- [x] I verified builds succeed (Go and UI)
- [ ] I verified the CI pipeline passes locally if applicable

* chunking_strategy as extra params for openai models (#4741)

* fix: fix the provider governance form UI to only show calendar aligned toggle when budget is alignable (#4724)

## Summary

The calendar alignment toggle in the provider governance form was previously shown whenever any budget existed. This PR restricts its visibility and submission to only when at least one budget uses a calendar-alignable reset period (day, week, month, or year).

## Changes

- Introduced a `showCalendarAlignment` derived boolean that checks whether any configured budget has a reset duration supported by `supportsCalendarAlignment`.
- Replaced the previous condition (`watchedBudgets.length > 0`) with `showCalendarAlignment` to control rendering of the calendar alignment toggle.
- Updated the form submission payload so that `calendar_aligned` is only set to `true` when at least one budget actually supports calendar alignment — preventing the flag from being submitted for incompatible budget configurations.

## Type of change

- [x] Bug fix
- [ ] Feature
- [ ] Refactor
- [ ] Documentation
- [ ] Chore/CI

## Affected areas

- [ ] Core (Go)
- [ ] Transports (HTTP)
- [ ] Providers/Integrations
- [ ] Plugins
- [x] UI (React)
- [ ] Docs

## How to test

1. Navigate to a provider's governance settings in the UI.
2. Add a budget with a reset duration that does **not** support calendar alignment (e.g., hourly). Verify the calendar alignment toggle does **not** appear.
3. Add or change a budget to use a calendar-alignable period (e.g., daily, weekly, monthly, yearly). Verify the toggle **does** appear.
4. Enable the toggle and save. Confirm `calendar_aligned: true` is included in the submitted payload.
5. Remove all calendar-alignable budgets and save. Confirm `calendar_aligned` is not set to `true` in the payload.

```sh
cd ui
pnpm i || npm i
pnpm test || npm test
pnpm build || npm run build
```

## Screenshots/Recordings

_Before:_ Calendar alignment toggle appears whenever any budget is present, regardless of reset period.

_After:_ Calendar alignment toggle only appears when at least one budget uses a day/week/month/year reset period.

## Breaking changes

- [x] No

## Related issues

## Security considerations

None.

## Checklist

- [ ] I read `docs/contributing/README.md` and followed the guidelines
- [ ] I added/updated tests where appropriate
- [ ] I updated documentation where needed
- [ ] I verified builds succeed (Go and UI)
- [ ] I verified the CI pipeline passes locally if applicable

* fix: custom providers with space in names could not set budget (#4725)

## Summary

Fixes a bug where updating or deleting provider-level governance for a custom provider whose name contains a space (e.g. `"OpenRouter Base"`) would return a 404. The UI percent-encodes the provider name in the URL path (`OpenRouter%20Base`), but the handler was comparing the raw encoded string directly against the stored provider name, causing the lookup to fail. Closes #4689

## Changes

- `updateProviderGovernance` and `deleteProviderGovernance` now call `url.PathUnescape` on the `provider_name` path parameter before using it, matching the decoded name against what is stored in the config store.
- Returns a `400` if the path parameter contains an invalid percent-encoding sequence.
- Added a regression test (`TestProviderGovernance_DecodesEncodedProviderName`) that seeds a provider with a space in its name, issues a PUT and DELETE using the percent-encoded path param, and asserts both succeed and persist correctly.
- Added a guard test (`TestProviderGovernance_UnknownProviderStill404`) to confirm that a genuinely unknown provider still returns 404 after the decode change.

## Type of change

- [x] Bug fix
- [ ] Feature
- [ ] Refactor
- [ ] Documentation
- [ ] Chore/CI

## Affected areas

- [ ] Core (Go)
- [x] Transports (HTTP)
- [ ] Providers/Integrations
- [ ] Plugins
- [ ] UI (React)
- [ ] Docs

## How to test

```sh
go test ./transports/bifrost-http/handlers/... -run TestProviderGovernance
```

Expected output: all three `TestProviderGovernance_*` tests pass. Specifically:

- `TestProviderGovernance_DecodesEncodedProviderName` — PUT and DELETE with `OpenRouter%20Base` return `200`.
- `TestProviderGovernance_UnknownProviderStill404` — PUT with an unknown encoded name returns `404`.

## Breaking changes

- [ ] Yes
- [x] No

## Related issues

- Resolves #4689

## Security considerations

`url.PathUnescape` is used rather than `url.QueryUnescape` to correctly handle path-encoded characters. Invalid encoding sequences are rejected with a `400` rather than passed through, preventing malformed input from reaching the config store.

## Checklist

- [x] I read `docs/contributing/README.md` and followed the guidelines
- [x] I added/updated tests where appropriate
- [ ] I updated documentation where needed
- [x] I verified builds succeed (Go and UI)
- [ ] I verified the CI pipeline passes locally if applicable

* fix: pass through gs:// image URLs on Vertex Gemini closes #4402 (#4568)

Signed-off-by: Akshay Deo <akshay@akshaydeo.com>
Co-authored-by: Akshay Deo <akshay@akshaydeo.com>

* Fix mcp reconnect failure on startup (#4316)

* Fix mcp reconnect failure on startup

* test: assert failed MCP client cleanup

---------

Co-authored-by: Gowtham <692171+HackToHell@users.noreply.github.com>

* fix(responses): preserve codex tool_search_call/tool_search_output input items (#4121)

* fix(responses): preserve codex tool_search_call/tool_search_output input items

Bifrost's Responses input deserializer rejected codex's tool-search follow-up
request with HTTP 400 "openai responses request input is neither a string nor an
array of responses messages", which hung/failed the agent turn. The fix teaches
ResponsesMessage about the two tool_search item types and round-trips them
verbatim. Background, since tool_search is non-obvious:

How codex's tool_search works (the path that hits this bug)
-----------------------------------------------------------
codex normally sends every MCP tool inline in the request `tools[]` as
`{type:"function", ...}`. But when a model's catalog has
`supports_search_tool: true` AND the tool count crosses
DIRECT_MCP_TOOL_EXPOSURE_THRESHOLD (= 100) — e.g. an agent wired to several MCP
servers — codex stops sending them inline and "defers" them behind a discovery
tool:
  should_defer = supports_search_tool && (ToolSearchAlwaysDeferMcpTools || n >= 100)

The deferred flow is a two-request round-trip:

  1. Request 1: codex hides the deferred tools and instead declares one tool:
       {"type":"tool_search","execution":"client","description":"...",
        "parameters":{query, limit}}
     `execution:"client"` means the model does NOT run the search — codex does.

  2. The model emits a `tool_search_call` with `arguments` = {query, limit}.

  3. codex runs the search CLIENT-SIDE: a BM25 index over the deferred tool
     metadata (codex's ToolSearchHandler, core/src/tools/handlers/tool_search.rs,
     using the `bm25` crate). It picks the top-N matching tools.

  4. Request 2 (follow-up): codex appends two items to `input[]`:
       - {"type":"tool_search_call",   "call_id":..., "execution":"client",
          "arguments":{...}}
       - {"type":"tool_search_output", "call_id":..., "status":"completed",
          "execution":"client", "tools":[ {type:"function", ...the matches} ]}
     and also surfaces the discovered tools in `tools[]`. The model can now call
     them. This repeats as the model needs more tools.

Root cause
----------
ResponsesMessage (the element type of the Responses `input` array AND the
response `Output` array) doesn't model `tool_search_call` / `tool_search_output`:

  - The call's `arguments` is a JSON OBJECT, whereas function_call's `arguments`
    is a JSON STRING. So it cannot decode into ResponsesToolMessage.Arguments
    (*string) -> sonic.Unmarshal of the whole []ResponsesMessage errors ->
    OpenAIResponsesRequestInput.UnmarshalJSON falls through to the "neither a
    string nor an array" 400. The entire request dies before reaching OpenAI.
  - The output's `tools` array is also unmodeled (would be dropped/mangled,
    which OpenAI then rejects with "Missing input[N].tools[0].type").

OpenAI's Responses API supports both items natively (verified end-to-end against
the gateway: the tool_search tool spec is accepted and echoed; OpenAI validates
arguments-as-object and tools[].type). So this is purely a Bifrost modelling gap,
in the same family as the tool-type allowlist that already lists
ResponsesToolTypeToolSearch / ResponsesToolTypeNamespace — just a different code
path (input-item deserialization vs the request tools[] allowlist).

Fix
---
Add ResponsesMessageTypeToolSearchCall / ResponsesMessageTypeToolSearchOutput and
give ResponsesMessage custom (Un)MarshalJSON that preserves these two item types
verbatim (original bytes in, original bytes out), so the object `arguments` and
the `tools` array survive intact. Every other item type defers to the default
struct (de)coding, unchanged. One change covers both directions because request
input and response Output are both []ResponsesMessage.

Impact: unblocks codex tool-search deferral (multi-MCP-server / >=100-tool agents)
through Bifrost. Verified with a round-trip test reproducing the exact follow-up
payload, plus the existing providers/openai and schemas suites (no regressions).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(responses): reset ResponsesMessage receiver in UnmarshalJSON

Clear the receiver at the top of ResponsesMessage.UnmarshalJSON so a reused
instance never retains a stale rawToolSearch (or other field) from a prior
decode. Without this, unmarshalling a tool_search item and then a normal
message into the same value would leave the preserved bytes in place, and
MarshalJSON would re-emit them. Not reachable via the array-decode path (each
element starts zero), but a cheap, defensive correctness fix.

Addresses CodeRabbit review on PR #4121.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(openapi): correct prompt_cache_retention enum to in_memory

The chat schema declared the enum as [in-memory, 24h], but OpenAI's
actual accepted values are in_memory (underscore) and 24h. The hyphenated
form was a typo from when the enum was first added and never matched
OpenAI, so spec-generated clients produced Literal['in-memory', '24h']
and rejected the valid value with a pydantic literal_error.

The Go runtime treats prompt_cache_retention as a pass-through *string,
so no behavior changes — only the spec enum, the regenerated openapi.json,
and the doc comment are corrected.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Signed-off-by: Akshay Deo <akshay@akshaydeo.com>
Co-authored-by: Suresh Kumar Ponnusamy <suresh@atomicwork.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Akshay Deo <akshay@akshaydeo.com>

* fix: build fix in vertex (#4765)

## Summary

Fixes a bug in the Vertex provider where the code path for handling Gemini/Gemma model families was duplicated, with the non-streaming branch incorrectly using `ToGeminiChatCompletionRequest` (without image URL scheme support) while the streaming branch used `ToGeminiChatCompletionRequestWithImageURLSchemes`. This consolidates the logic so both paths use the image URL scheme-aware converter.

## Changes

- Replaced `ToGeminiChatCompletionRequest` with `ToGeminiChatCompletionRequestWithImageURLSchemes` in the non-streaming Gemini/Gemma branch, making it consistent with the streaming branch
- Removed the duplicate non-streaming Gemini/Gemma and OpenAI handler blocks that had been incorrectly separated from the streaming path, consolidating them into a single unified code path

## Type of change

- [x] Bug fix
- [ ] Feature
- [ ] Refactor
- [ ] Documentation
- [ ] Chore/CI

## Affected areas

- [ ] Core (Go)
- [ ] Transports (HTTP)
- [x] Providers/Integrations
- [ ] Plugins
- [ ] UI (React)
- [ ] Docs

## How to test

Send a chat completion request to the Vertex provider using a Gemini or Gemma model with image URL content. Verify that image URLs are correctly processed in both streaming and non-streaming modes.

```sh
go test ./core/providers/vertex/...
```

## Breaking changes

- [ ] Yes
- [x] No

## Related issues

## Security considerations

None.

## Checklist

- [ ] I read `docs/contributing/README.md` and followed the guidelines
- [ ] I added/updated tests where appropriate
- [ ] I updated documentation where needed
- [ ] I verified builds succeed (Go and UI)
- [ ] I verified the CI pipeline passes locally if applicable

* fix: responses.go build fix (#4766)

## Summary

Consolidates the two separate `UnmarshalJSON` implementations on `ResponsesMessage` into a single method that handles both the verbatim `tool_search` preservation and the `arguments` normalization logic. Previously, the file contained a duplicate `UnmarshalJSON` definition — the first handled `tool_search` items and fell back to a plain `sonic.Unmarshal`, while the second (the correct one) handled argument normalization. The duplicate caused the normalization path to be unreachable for non-`tool_search` items, meaning `tool_search_call` items with object-typed `arguments` would silently fail mid-stream and hang streaming clients.

## Changes

- Removed the redundant first `UnmarshalJSON` that short-circuited to `sonic.Unmarshal` without normalizing `arguments`, leaving only the correct implementation that handles both the `rawToolSearch` early-return and the `arguments` object-to-string normalization.
- Relocated `MarshalJSON` to follow `UnmarshalJSON` for logical grouping.
- The fix ensures `tool_search_call` items whose `arguments` field is a JSON object (e.g. `{}` while in-progress, `{"query":"...","limit":10}` when completed) are correctly stringified into the `*string` field expected by `ResponsesToolMessage`, preventing decode failures that previously dropped items silently.

## Type of change

- [x] Bug fix
- [ ] Feature
- [ ] Refactor
- [ ] Documentation
- [ ] Chore/CI

## Affected areas

- [x] Core (Go)
- [ ] Transports (HTTP)
- [ ] Providers/Integrations
- [ ] Plugins
- [ ] UI (React)
- [ ] Docs

## How to test

```sh
go test ./core/schemas/...
```

Validate by sending a request that triggers `tool_search_call` streaming events and confirming that items with both `{}` (in-progress) and `{"query":"...","limit":10}` (completed) `arguments` values are decoded without error and do not hang the streaming client.

## Breaking changes

- [ ] Yes
- [x] No

## Related issues

## Security considerations

None.

## Checklist

- [ ] I read `docs/contributing/README.md` and followed the guidelines
- [ ] I added/updated tests where appropriate
- [ ] I updated documentation where needed
- [ ] I verified builds succeed (Go and UI)
- [ ] I verified the CI pipeline passes locally if applicable

* feat: add server/client_id filter to MCP clients list (#4767)

## Summary

Adds a `server` (client ID) filter to the MCP clients list endpoint and UI, allowing users to filter the MCP clients table to a specific server. The filter state is persisted in the URL via query parameters, enabling shareable and bookmarkable filtered views.

## Changes

- Added `ClientID` field to `MCPClientsQueryParams` and applied it as a `WHERE client_id = ?` clause in `GetMCPClientsPaginated`
- Exposed the filter via a new `server` query parameter on `GET /api/mcp/clients`
- Migrated the MCP registry page from local `useState` to `nuqs` `useQueryStates`, storing `search`, `server`, and `offset` in the URL
- Added `server` prop and `onServerFilterClear` callback to `MCPClientsTable`, rendering a dismissible "Server filter" badge/button when the filter is active
- Extended `GetMCPClientsParams` type and the RTK Query API call to pass the `server` parameter through to the backend
- `hasActiveFilters` now accounts for both `debouncedSearch` and `server`, preventing the empty state from showing while a server filter is active

## Type of change

- [ ] Bug fix
- [x] Feature
- [ ] Refactor
- [ ] Documentation
- [ ] Chore/CI

## Affected areas

- [x] Core (Go)
- [x] Transports (HTTP)
- [ ] Providers/Integrations
- [ ] Plugins
- [x] UI (React)
- [ ] Docs

## How to test

1. Navigate to the MCP Registry page.
2. Confirm that `search`, `server`, and `offset` appear in the URL and survive a page refresh.
3. Set a `server` query param (e.g. `?server=<client_id>`) directly in the URL and verify the table filters to only that client.
4. Click the "Server filter" dismiss button and confirm the filter clears and the URL updates.
5. Verify that the empty state is not shown when a server filter is active but returns no results.

```sh
# Core/Transports
go test ./framework/configstore/... ./transports/bifrost-http/...

# UI
cd ui
pnpm i
pnpm build
```

## Screenshots/Recordings

_Add before/after screenshots showing the server filter badge and URL state._

## Breaking changes

- [ ] Yes
- [x] No

## Related issues

_Link related issues here._

## Security considerations

The `client_id` filter is applied as a parameterised query (`WHERE client_id = ?`), so there is no SQL injection risk. No secrets or PII are exposed through the new filter parameter.

## Checklist

- [ ] I read `docs/contributing/README.md` and followed the guidelines
- [ ] I added/updated tests where appropriate
- [ ] I updated documentation where needed
- [ ] I verified builds succeed (Go and UI)
- [ ] I verified the CI pipeline passes locally if applicable

* fix: redact decoder details from invalid request payload errors (#4770)

## Summary

Error messages returned on JSON decode failures were leaking internal decoder details (e.g., field names, Go type information, and `cannot unmarshal` messages) back to API callers. This replaces all such messages with a single, generic `"Invalid request payload"` string to avoid exposing implementation internals.

## Changes

- Replaced all `fmt.Sprintf("invalid/Invalid request format: %v", err)` and similar patterns across handlers (`config`, `featureflags`, `governance`, `inference`, `mcp`, `mcp_per_user_headers`, `mcpinference`, `provider_keys`, `providers`, `session`) with the static string `"Invalid request payload"`.
- Removed now-unused `fmt` import from `mcpinference.go`.
- Added `requestpayload_test.go` with two tests that assert the generic message is returned and that decoder internals (`cannot unmarshal`, field names, Go struct details) are not present in the response body or error string.
- Updated the existing `governance_test.go` assertion for the unknown-field case to expect `"Invalid request payload"` instead of `"unknown field"`.

## Type of change

- [ ] Bug fix
- [ ] Feature
- [ ] Refactor
- [ ] Documentation
- [ ] Chore/CI

## Affected areas

- [ ] Core (Go)
- [x] Transports (HTTP)
- [ ] Providers/Integrations
- [ ] Plugins
- [ ] UI (React)
- [ ] Docs

## How to test

```sh
go test ./transports/bifrost-http/handlers/...
```

Confirm that:
- `TestSessionLoginInvalidPayloadDoesNotExposeDecoderDetails` passes and the response body contains `"Invalid request payload"` with no decoder internals.
- `TestPrepareRequestInvalidPayloadDoesNotExposeDecoderDetails` passes and the returned error is exactly `"invalid request payload"`.
- `TestComplexityAnalyzerConfigPutRejectsInvalidPayloads` passes with the updated `"Invalid request payload"` expectation for the unknown-field case.

## Breaking changes

- [ ] Yes
- [x] No

## Security considerations

Decoder error messages from Go's `encoding/json` and `sonic` can expose internal struct field names, type information, and value details. Returning these verbatim in HTTP responses constitutes an information disclosure risk. This change ensures all parse-failure responses return a fixed, opaque message regardless of the underlying decode error.

## Checklist

- [ ] I read `docs/contributing/README.md` and followed the guidelines
- [x] I added/updated tests where appropriate
- [ ] I updated documentation where needed
- [x] I verified builds succeed (Go and UI)
- [ ] I verified the CI pipeline passes locally if applicable

* complexity router : analyzer's no-signal fallback + few improvements (#4708)

## Summary

Briefly explain the purpose of this PR and the problem it solves.

## Changes

- What was changed and why
- Any notable design decisions or trade-offs

## Type of change

- [ ] Bug fix
- [ ] Feature
- [ ] Refactor
- [ ] Documentation
- [ ] Chore/CI

## Affected areas

- [ ] Core (Go)
- [ ] Transports (HTTP)
- [ ] Providers/Integrations
- [ ] Plugins
- [ ] UI (React)
- [ ] Docs

## How to test

Describe the steps to validate this change. Include commands and expected outcomes.

```sh
# Core/Transports
go version
go test ./...

# UI
cd ui
pnpm i || npm i
pnpm test || npm test
pnpm build || npm run build
```

If adding new configs or environment variables, document them here.

## Screenshots/Recordings

If UI changes, add before/after screenshots or short clips.

## Breaking changes

- [ ] Yes
- [ ] No

If yes, describe impact and migration instructions.

## Related issues

Link related issues and discussions. Example: Closes #123

## Security considerations

Note any security implications (auth, secrets, PII, sandboxing, etc.).

## Checklist

- [ ] I read `docs/contributing/README.md` and followed the guidelines
- [ ] I added/updated tests where appropriate
- [ ] I updated documentation where needed
- [ ] I verified builds succeed (Go and UI)
- [ ] I verified the CI pipeline passes locally if applicable

* complexity: stemming support along w exact keyword match in complexity analyzer's logic (#4791)

## Summary

Briefly explain the purpose of this PR and the problem it solves.

## Changes

- What was changed and why
- Any notable design decisions or trade-offs

## Type of change

- [ ] Bug fix
- [ ] Feature
- [ ] Refactor
- [ ] Documentation
- [ ] Chore/CI

## Affected areas

- [ ] Core (Go)
- [ ] Transports (HTTP)
- [ ] Providers/Integrations
- [ ] Plugins
- [ ] UI (React)
- [ ] Docs

## How to test

Describe the steps to validate this change. Include commands and expected outcomes.

```sh
# Core/Transports
go version
go test ./...

# UI
cd ui
pnpm i || npm i
pnpm test || npm test
pnpm build || npm run build
```

If adding new configs or environment variables, document them here.

## Screenshots/Recordings

If UI changes, add before/after screenshots or short clips.

## Breaking changes

- [ ] Yes
- [ ] No

If yes, describe impact and migration instructions.

## Related issues

Link related issues and discussions. Example: Closes #123

## Security considerations

Note any security implications (auth, secrets, PII, sandboxing, etc.).

## Checklist

- [ ] I read `docs/contributing/README.md` and followed the guidelines
- [ ] I added/updated tests where appropriate
- [ ] I updated documentation where needed
- [ ] I verified builds succeed (Go and UI)
- [ ] I verified the CI pipeline passes locally if applicable

* refactor: rename MCP handler files to remove underscores and drop redundant prefixes (#4747)

## Summary

Renames several handler files in the `bifrost-http` transport package to use a consistent naming convention, removing underscores in favor of camelCase-style concatenated names.

## Changes

- `mcp_per_user_headers.go` → `mcpheaders.go`
- `oauth2.go` → `mcpoauth2.go`
- `mcp_sessions.go` → `mcpsessions.go`
- `mcp_sessions_test.go` → `mcpsessions_test.go`
- `temp_token_scopes.go` → `temptokens.go`

## Type of change

- [ ] Bug fix
- [ ] Feature
- [ ] Refactor
- [ ] Documentation
- [x] Chore/CI

## Affected areas

- [ ] Core (Go)
- [x] Transports (HTTP)
- [ ] Providers/Integrations
- [ ] Plugins
- [ ] UI (React)
- [ ] Docs

## How to test

```sh
go version
go test ./...
```

## Breaking changes

- [ ] Yes
- [x] No

## Related issues

## Security considerations

None.

## Checklist

- [ ] I read `docs/contributing/README.md` and followed the guidelines
- [ ] I added/updated tests where appropriate
- [ ] I updated documentation where needed
- [ ] I verified builds succeed (Go and UI)
- [ ] I verified the CI pipeline passes locally if applicable

* feat: add OAuth2 AS discovery endpoints, signing key management, and `MCPServerAuthMode` config (#4505)

## Summary

This PR introduces the foundational OAuth 2.1 authorization server infrastructure for Bifrost's `/mcp` endpoint. It adds a configurable `mcp_server_auth_mode` that controls how inbound MCP clients are authenticated, enabling Bifrost to act as a spec-compliant OAuth 2.1 AS with RFC-mandated discovery endpoints, a JWKS endpoint, and a persistent RS256 signing keypair.

## Changes

- **`MCPServerAuthMode`** — new `varchar` column on `TableClientConfig` with three modes:
  - `headers` (default): existing VK/api-key/session header auth only; discovery endpoints return 404.
  - `both`: accepts both header credentials and Bifrost-issued JWTs; discovery endpoints are live.
  - `oauth`: Bifrost JWTs only; header credentials are rejected on `/mcp`. **Breaking for existing VK-based MCP integrations.**
- **`OAuth2ServerConfig`** — new JSON blob column on `TableClientConfig` holding AS-specific settings (`IssuerURL`, `AuthCodeTTL`, `AccessTokenTTL`). Serialized via `BeforeSave`/`AfterFind` hooks. Only meaningful when mode is `both` or `oauth`.
- **`OAuth2SigningKey`** — RS2048 keypair generated on first use and persisted in `governance_config` under `oauth2_signing_key`. The private key PEM is encrypted at rest via `framework/encrypt` when encryption is enabled.
- **`GetOAuth2SigningKey`** — new `ConfigStore` interface method that lazily generates and persists the signing keypair on first call, always returning a usable key.
- **`OAuth2DiscoveryHandler`** — serves the three well-known discovery endpoints:
  - `GET /.well-known/oauth-protected-resource[/{path}]` (RFC 9728)
  - `GET /.well-known/oauth-authorization-server[/{path}]` (RFC 8414)
  - `GET /.well-known/jwks.json` (RFC 7517)
  All three return 404 when `MCPServerAuthMode` is `headers`. Routes are always registered; the mode flag is the feature toggle.
- **`oauth2IssuerURL` / `oauth2ServerCfg`** — utility helpers that resolve the effective issuer URL (configured `IssuerURL` or request-derived fallback) and AS config defaults.
- **`OAuth2ConsentScopeName`** — new temp-token scope for binding browser sessions to pending authorization requests on the public consent page.
- **Database migration** `add_oauth2_server_tables` — adds `mcp_server_auth_mode` and `oauth2_server_config_json` columns to `config_client`.
- **Config schema** — `mcp_server_auth_mode` and `oauth2_server_config` added to `config.schema.json` with full descriptions and validation.
- **`IsMCPOAuthDiscoveryEnabled`** — helper on `ClientConfig` that returns true when mode is `both` or `oauth`.

## Type of change

- [ ] Bug fix
- [x] Feature
- [ ] Refactor
- [ ] Documentation
- [ ] Chore/CI

## Affected areas

- [x] Core (Go)
- [x] Transports (HTTP)
- [ ] Providers/Integrations
- [ ] Plugins
- [ ] UI (React)
- [ ] Docs

## How to test

```sh
go test ./framework/configstore/... ./transports/bifrost-http/...
```

1. Start Bifrost with `mcp_server_auth_mode` unset (or `"headers"`). Confirm `GET /.well-known/oauth-authorization-server` returns 404.
2. Set `mcp_server_auth_mode` to `"both"` and restart. Confirm:
   - `GET /.well-known/oauth-authorization-server` returns a valid JSON document with `issuer`, `authorization_endpoint`, `token_endpoint`, etc.
   - `GET /.well-known/oauth-protected-resource` returns a document pointing to `/mcp`.
   - `GET /.well-known/jwks.json` returns a JWKS with one RS256 key entry.
3. Confirm the signing keypair is persisted in `governance_config` and survives a restart (same `kid` returned).
4. Set `mcp_server_auth_mode` to `"oauth"` and confirm header-credential MCP requests are rejected.

**New config fields:**

| Field | Type | Default | Description |
|---|---|---|---|
| `mcp_server_auth_mode` | `"headers"` \| `"both"` \| `"oauth"` | `"headers"` | Inbound MCP client auth mode |
| `oauth2_server_config.issuer_url` | string / env var | _(request host)_ | Stable AS issuer URL |
| `oauth2_server_config.auth_code_ttl` | int (seconds) | `600` | Authorization code lifetime |
| `oauth2_server_config.access_token_ttl` | int (seconds) | `600` | JWT Bearer token lifetime |

## Breaking changes

- [x] Yes
- [ ] No

Setting `mcp_server_auth_mode` to `"oauth"` disables VK/api-key/session header authentication on `/mcp`. Existing virtual-key MCP integrations will stop working. Use `"both"` for a non-breaking migration path that accepts both credential types simultaneously.

## Security considerations

- The RS256 private key is encrypted at rest using `framework/encrypt` when encryption is enabled. The plaintext key is only held in memory during the signing operation.
- The `OAuth2ConsentScopeName` temp token is the sole credential binding a browser session to a pending authorization request on the public (unauthenticated) consent page — it must be treated as a short-lived secret.
- Refresh tokens have no timer-based expiry; they are invalidated only by rotation on use, subject liveness checks, explicit revocation, or enforcement policy changes.
- Multi-host / reverse-proxy deployments must set a stable `issuer_url`; omitting it causes token verification failures when the `Host` header differs across nodes.

## Checklist

- [ ] I read `docs/contributing/README.md` and followed the guidelines
- [ ] I added/updated tests where appropriate
- [ ] I updated documentation where needed
- [ ] I verified builds succeed (Go and UI)
- [ ] I verified the CI pipeline passes locally if applicable

* feat: add OAuth2 issuance endpoints (DCR, authorize, token) with PKCE and refresh token rotation (#4506)

## Summary

This PR implements the downstream OAuth2 token issuance flow, enabling Bifrost to act as a full OAuth2 authorization server. It adds the three core RFC-compliant endpoints (Dynamic Client Registration, Authorization, and Token), the backing database tables and store methods, and the consent temp-token scope needed to bind the consent UI to a specific authorization request.

## Changes

- **New DB tables** (`oauth2_clients`, `oauth2_authorize_requests`, `oauth2_refresh_tokens`) added via a new `add_oauth2_issuance_tables` migration step, with GORM model definitions in `tables/oauth2_issuance.go`. `TableOAuth2Client` serializes `redirect_uris` and `grant_types` as JSON columns with `BeforeSave`/`AfterFind` hooks.
- **`ConfigStore` interface** extended with methods for creating/fetching OAuth2 clients, managing authorize requests (including code-hash lookup and expiry sweeping), and refresh token operations (`GetOAuth2RefreshTokenByHash`, `ConsumeOAuth2AuthorizeRequest`, `RotateOAuth2RefreshToken`).
- **`RDBConfigStore`** implements all new interface methods. `ConsumeOAuth2AuthorizeRequest` and `RotateOAuth2RefreshToken` are wrapped in transactions so failures leave the grant in a retryable state.
- **`OAuth2IssuanceHandler`** (`handlers/oauth2_issuance.go`) wires three public routes:
  - `POST /oauth2/register` — RFC 7591 DCR; only public clients (`token_endpoint_auth_method=none`) are accepted.
  - `GET /oauth2/authorize` — PKCE-S256 (RFC 7636) + resource indicator (RFC 8707); creates a pending authorize request and redirects to the consent UI with an optional scoped temp token.
  - `POST /oauth2/token` — handles `authorization_code` and `refresh_token` grants; issues RS256 JWTs signed with the persisted signing key and rotates refresh tokens on every use.
- **Stolen-token detection** via `FamilyID` on refresh tokens: all tokens descended from the same authorization grant share a family ID, enabling full family revocation when a revoked token is re-presented (RFC 9700 §2.2.2).
- **Loopback redirect URI matching** follows RFC 8252 §7.3 (port-agnostic for `localhost`/`127.0.0.1`).
- **`oauth2ConsentScope`** temp-token scope registered at startup, binding consent-page API calls to a single authorize request ID via path substitution.
- **Server bootstrap** pre-warms the OAuth2 signing key when MCP OAuth discovery is enabled, so JWKS and JWT signing are ready before the first request.
- `github.com/golang-jwt/jwt/v5` promoted from indirect to a direct dependency.

## Type of change

- [ ] Bug fix
- [x] Feature
- [ ] Refactor
- [ ] Documentation
- [ ] Chore/CI

## Affected areas

- [x] Core (Go)
- [x] Transports (HTTP)
- [ ] Providers/Integrations
- [ ] Plugins
- [ ] UI (React)
- [ ] Docs

## How to test

```sh
go test ./framework/configstore/... ./transports/bifrost-http/...
```

1. Start Bifrost with MCP OAuth discovery enabled and a configured config store.
2. Register a client:
   ```sh
   curl -X POST http://localhost:PORT/oauth2/register \
     -H "Content-Type: application/json" \
     -d '{"client_name":"test","redirect_uris":["http://localhost:8080/callback"]}'
   ```
3. Initiate an authorization request via `GET /oauth2/authorize` with `response_type=code`, `code_challenge` (S256), `resource`, and the returned `client_id`.
4. Complete the consent flow and exchange the auth code at `POST /oauth2/token` with `grant_type=authorization_code` and the PKCE verifier.
5. Refresh the access token using `grant_type=refresh_token` and verify the old refresh token is revoked and a new one is issued.

## Breaking changes

- [x] No

## Security considerations

- Refresh tokens are stored as SHA-256 hashes only; plaintext is returned to the client once and never persisted.
- Auth codes are single-use: `ConsumeOAuth2AuthorizeRequest` atomically transitions the request to `code_issued` and creates the refresh token in one transaction.
- Refresh token rotation is atomi…
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug] chunking_strategy is dropped for OpenAI-compatible transcription requests

2 participants