diff --git a/core/changelog.md b/core/changelog.md index dd13b8feefe..c7d9abe97ec 100644 --- a/core/changelog.md +++ b/core/changelog.md @@ -1,19 +1,54 @@ - fix: forward OpenCode Responses requests directly to /v1/responses [@mohammadrezwankhan](https://github.com/mohammadrezwankhan) -- fix: preserve max_tokens for OpenCode-compatible chat endpoints [@Alex-wangyang](https://github.com/Alex-wangyang) -- fix: stop treating HuggingFace as a provider that omits the `[DONE]` marker - it does send one, and listing it as not sending one made the shared OpenAI streaming loop `break` on the first `finish_reason`, discarding the trailing usage-only chunk (`choices: []` plus top-level `usage`) that several router inference providers emit after it. Those streams completed with zero tokens and therefore zero cost, while non-streaming calls to the same models priced correctly [@elliottrabac](https://github.com/elliottrabac) -- fix: default `stream_options.include_usage` on HuggingFace chat streaming - the shared OpenAI streaming handler sets it on its own request path but returns early when a provider supplies a custom request converter, so HuggingFace never sent it. Several router inference providers then omit the terminal usage chunk (and Together nests usage under `choices[]` for some models, which the shared accumulator does not read), so streams completed with zero tokens and therefore zero cost while non-streaming calls on the same models priced correctly. An explicit `stream_options` from the caller still wins [@elliottrabac](https://github.com/elliottrabac) -- feat: support Gemini's server-side `toolCall`/`toolResponse` parts with `thoughtSignature` round-trip fidelity - server-side search rounds now surface as `web_search_call` items carrying their own call ID and queries, unmapped tool types are preserved on the native round-trip instead of being dropped, and each `thoughtSignature` appears exactly once across the reconstructed parts so Gemini accepts the replayed turn -- feat: async 3D generation on Runware via `/videos` plus a raw `/runware_passthrough` route - `taskType` is now read from extra_params so any Runware async task can be driven through `/videos` (the 16:9 1080p width/height defaults now apply only to `videoInference`), `outputs.files[].url` is surfaced as `VideoOutput` URLs with the content type derived from the file extension, and the passthrough route forwards raw task arrays for capabilities with no first-class Bifrost surface such as upscaling and background removal -- feat: surface Runware's provider-reported per-task `cost` across image, video/3D and passthrough so pricing uses the exact figure verbatim instead of a datasheet estimate - this matters for task types like 3D that have no datasheet rate; when no cost is reported the behavior is unchanged -- feat: send `s3://` image and document references to Bedrock Converse as the `s3Location` source member instead of downloading the bytes and re-uploading them - Converse resolves the object itself, which skips a round trip and the 25 MiB inline cap entirely. Image format is derived from the object extension since nothing is fetched and there is no `Content-Type` to read, and an extension-less object is rejected up front rather than producing an opaque 400 -- feat: resolve Vertex URL sources per model family rather than inlining everything - a `gs://` URI is now forwarded to Gemini/Gemma as `fileData.fileUri` (the documented form, resolved under the caller's own project IAM, and the only thing that keeps multi-hundred-MB video inputs viable) and read from Cloud Storage with the request key's own Google credentials for Claude-on-Vertex, which accepts base64 sources only. `http(s)` is still always fetched: forwarding one was measured against the harness and Vertex rejected every endpoint shape with `URL_REJECTED-REJECTED_FC_TOO_MANY_PENDING` -- fix: allow Vertex AI to send function declarations and a Google Search tool in the same request without `includeServerSideToolInvocations` - Vertex accepts the combination natively, so Google Search was being dropped for no reason, and search localization via `RetrievalConfig.LatLng` is now preserved when both tool types are present -- fix: prefer function declarations over Google Search when tool combination is disabled on Gemini - Google's tool combination is Preview and Gemini 3 only (https://ai.google.dev/gemini-api/docs/generate-content/tool-combination), so on every other model one of the two tool types has to be dropped. Function declarations now win: they carry the caller's own tools, or the ones Bifrost's MCP gateway synthesized from their connected servers, and dropping those leaves the model unable to invoke them at all while it answers as though the capabilities never existed. Dropping Google Search only costs grounding, so the model still answers, just without citations. One is disabled, the other is degraded. Set `include_server_side_tool_invocations` to send both. A lone Google Search tool still converts back correctly, and `retrievalConfig` is only emitted when a search tool actually survived conversion +- feat: add the `VideoEdit` operation with `BifrostVideoEditRequest`, `VideoEditInput` and `VideoEditParameters` for prompt-driven edits, upscaling and background removal on an existing video supplied as bytes, a URL or a provider video ID; implemented for OpenAI (`/v1/videos/edits`) and Runware (`videoInference`, `upscale`, `removeBackground`), with the model optional when the source is a video ID and the prompt optional for asset-driven task types (#6270) +- feat: batch accounting: `MergeBifrostLLMUsage` promoted to `schemas`, `Endpoint` on `BifrostBatchResultsResponse`, `BatchResultItem.Failed()`, `BatchRequestCountsFromResults`, `BatchRequestCounts.IsZero()`, raw-JSON Gemini batch result parsing and `custom_id` validation in `ConvertRequestsToJSONL`; the settlement engine (`AccountBatchResults` with runner-ID ownership fencing, idempotent aggregate log writes and governance reporting) and a sweeper that polls due jobs with capped, jittered backoff; aggregate log entries carry a `bifrost/` user agent via `BifrostContextKeyRuntimeVersion` (thanks [@SahilChoudhary22](https://github.com/SahilChoudhary22)!) (#5291, #5294, #6474) +- feat: Claude-on-Vertex batch support: `ToVertexBatchCreateRequest` resolves Anthropic families to `publishers/anthropic/models/...`, `vertexConvertRequestsToJSONL` emits Claude-on-Vertex instances, `custom_id` round-trips through `batchResultsByKey`, and `GeminiBatchGenerateContentRequest` keeps `tools`, `toolConfig`, `cachedContent`, `labels` and the display name (#5368) +- feat: add `HTTPTransportPreAuthHook` to the `HTTPTransportPlugin` interface, a phase that runs before transport authentication; `HTTPTransportPreHook` now runs after it (#6375) - Breaking on the Gemini API surface: a request carrying both function declarations and Google Search without `include_server_side_tool_invocations` previously kept Google Search and dropped the function declarations. It now does the opposite. Set `include_server_side_tool_invocations` to `true` to send both, which is supported on Gemini 3 models. Vertex is unaffected, since it accepts the combination natively and drops neither. + Breaking for plugin authors: Go plugins implementing `HTTPTransportPlugin` must add `HTTPTransportPreAuthHook` (`.so` plugins that predate it are skipped for that phase), and any plugin that injected a credential such as `x-bf-vk` or `Authorization` from `HTTPTransportPreHook` must move that work to `HTTPTransportPreAuthHook`, since the pre-hook no longer runs before auth. -- fix: always emit a Gemini candidate carrying its finish reason on `generateContent`, even when nothing visible was generated - a thinking model that spends its whole output budget before emitting a token is a successful 200 with an empty answer, but `Candidates` is `omitempty`, so dropping that candidate produced a body with no `candidates` key at all and left a lone `usageMetadata` object that every Gemini-shaped client dereferences blind -- fix: drop payload-free Gemini parts when assembling a candidate - every `Part` field is `omitempty`, so such a part marshals to exactly `{}`; the harness observed one on the wire when a transcription request for an unintelligible tone came back as `parts:[{}]`, where it is noise a client will try to read and it masks the contentless case by making the parts slice look non-empty -- fix: accept a bare model identifier on Bedrock rerank by synthesizing the foundation-model ARN from the resolved region - Rerank is the one Bedrock surface that names its model by ARN rather than by bare ID, so all three rerank drop-ins in the provider harness 400'd on `amazon.rerank-v1:0`. The partition is derived from the region (`aws`, `aws-cn`, `aws-us-gov`) so GovCloud and China build a correct ARN, and an explicit ARN still passes through untouched -- fix: stop stripping `file_url` from OpenAI-shaped chat file blocks on marshal - dropping it produced `{"type":"file","file":{}}` and an upstream complaint about a missing `file_id`, which hid the fact that a source had been discarded. Providers that cannot take a URL now say so by name, and any OpenAI-compatible endpoint that does accept one keeps working without a Bifrost change -- fix: leave URL content sources Bifrost cannot download in place on the OpenAI and native-Anthropic paths instead of failing the request - only `http(s)` is fetched, and whether a `gs://`, `s3://` or scheme-less reference is usable is the provider's call, so the source now travels as `{"type":"url"}` and the platform answers for itself +- feat: add `semaphore_size` and `inject_timeout` to `PluginConfig` so observability `Inject` calls are context-bounded per plugin (#6341) +- feat: Runware provider expansion: chat completions, streaming and Responses through its OpenAI-compatible `/v1/chat/completions` endpoint (Responses muxed via `ToChatRequest()`), `ListModels` sweeping the curated `modelSearch` catalog with the AIR as the model ID, image upscale via `/v1/images/edits` (`type=upscale`) and image-to-3D via `/v1/videos` (`type=3d`), a shared `settings` extra-param coercion for multipart and JSON callers, prompt-optional asset-driven operations, and input handling for edit, upscale and video task shapes (#6260, #6372, #6208) +- feat: OpenAI `ultrafast` service tier: `BifrostServiceTierUltrafast`, capability-gated forwarding via `serviceTierForModel` on chat, Responses and compaction, and `ultrafast` preserved through `WithDefaults` (#6396) +- feat: JSON bodies on `/v1/images/edits`: `ImageInput` accepts a bare string or `{ "url", "image" }`, typed extra params reach providers with their real types, and `images` is a known field (#6418) +- feat: `EmbeddingData.EncodingFormat` with typed `int8`, `uint8`, `binary`, `ubinary` and `base64` vectors; Bedrock Titan V2 `embeddingTypes` and Cohere `embedding_types` on Converse, native invoke and LangChain `BedrockEmbeddings` compatibility (#6381) +- feat: rerank: `RerankDocument.Data` for structured documents, `RerankResult.ID`, `RerankParameters.NextToken`, `ReturnDocuments` forwarded to Cohere and Vertex, `ToCohereError` for Cohere-shaped errors, `/genai/v1/rank` served cross-provider via `x-model-provider`, cross-provider responses converted back to the caller's wire shape with `ToBedrockRerankResponse`, `ToCohereRerankResponse` and `ToVertexRankResponse`, and rerank cost accounting for Bedrock and Cohere (#6301, #6328) +- feat: datasheet-backed compatibility flows: Anthropic, Bedrock, Cohere and Gemini request shaping (adaptive-only thinking, adaptive thinking, native effort, disable-reasoning, mid-conversation system turns, computer-use and text-editor tool generations, default max output tokens, tool validation, thinking-budget zeroing) is resolved through `schemas.ResolveModelCaps` instead of hardcoded model-name checks (#6281, #6492) +- feat: Gemini 3 per-model `thinkingLevel` support table (`geminiThinkingLevelSupport`) with `clampThinkingLevel` snapping requested levels to the nearest rung (ties break upward) and `lowestThinkingLevel` for `reasoning_effort: "none"`, so `setThinkingBudgetZeroIfSupported` sets the floor level on Gemini 3+ instead of zeroing `thinkingBudget` (#6280) +- feat: Bifrost overhead latency accounting: `upstream_latency` and `overhead_latency` on `BifrostResponseExtraFields` (`PopulateOverheadLatency`, `BifrostContextKeyRequestStartTime`, `populateLatencyExtraFields` so logging plugins see both at hook time); per-phase overhead spans across the request pipeline (`queue-wait`, `attribute-population`, `convertor`, `request-marshal`, `response-parse`, `handle-setup`, `pipeline-pre`, `pipeline-post`, `worker-setup`, `key-pool`, Bedrock `request-sign` and `credentials-fetch`, `response-finalize`) with `StampWorkerHandoff` on `ChannelMessage.sentAt`; lock-free stream overhead accumulators for per-chunk parse, conversion and backpressure installed via `ResetStreamOverhead`, `StampStreamTransport` for the outbound marshal and client-write time, and `defaultSSEDataReader.ReadDataLine` attributing socket reads to upstream; and `IsOverheadBreakdownSpan`, `WithoutOverheadBreakdownSpans` and the `OverheadSpanConsumer` interface so breakdown spans stay out of connectors that do not opt in (#5533, #6388, #6389, #6433, #6470, #6495) +- feat: input/output/additional cost split (`BifrostCost`) on inference usages, extended to speech, transcription and OCR usages +- feat: `Notification`, `NotificationInput`, `NotificationSeverity`, `NotificationAudience` and the `NotificationPublisher` function type for the dashboard notification center (#6207) +- feat: `BifrostContextKeySkipModelCheck` short-circuits the virtual key model allowlist for evaluate-only requests such as `/inspect` while keeping every other governance rule (#6479) +- feat: `HarnessSessionHeaders` and `MaxSessionIDLength` so Claude Code, Codex CLI and OpenCode session headers can fall back into the session ID (#6333) +- feat: `RedactSensitiveHeaders`, with `IsSensitiveHeader` extended to Cloudflare Access (`cf-access-*`), AWS ALB OIDC (`x-amzn-oidc-*`) and generic `jwt`/`assertion` headers (#6371) +- feat: `ResponsesResponseError.Type` and a shared Responses stream-error normalizer so terminal `error`/`response.failed` events inside an HTTP 200 Azure SSE stream surface as errors with their nested type, code and message on both create-stream and retrieve-stream paths (thanks [@dani29](https://github.com/dani29)!) (#6302) +- feat: `ServiceTier` on `StreamAccumulatorResult`, with Anthropic's `service_tier` from `message_start` latched onto the final chunk of chat and Responses streams (#6236) +- feat: OpenRouter speech and transcription through the OpenAI-compatible audio handlers instead of returning unsupported-operation errors (#5734) +- fix: preserve `max_tokens` for OpenCode-compatible chat endpoints (thanks [@Alex-wangyang](https://github.com/Alex-wangyang)!) (#6458) +- fix: HuggingFace chat streaming completed with zero tokens and therefore zero cost while non-streaming calls on the same models priced correctly, for two reasons: HuggingFace was listed as a provider that omits the `[DONE]` marker (it sends one), which made the shared OpenAI streaming loop `break` on the first `finish_reason` and discard the trailing usage-only chunk that several router inference providers emit; and `stream_options.include_usage` never reached the router because the shared streaming handler returns early when a provider supplies a custom request converter. Both are corrected, and an explicit `stream_options` from the caller still wins (thanks [@elliottrabac](https://github.com/elliottrabac)!) (#6478) +- fix: preserve the caller's JSON Schema key order for structured outputs - `ChatParameters.UnmarshalJSON` holds `response_format` as raw bytes and the new `ChatResponseFormat` reader splices them verbatim into OpenAI, Anthropic, Bedrock, Gemini (unless a union `type` array needs normalizing) and Cohere requests, and `ResponsesTextConfigFormatJSONSchema` re-encodes in the decoded key sequence, because OpenAI structured outputs generate fields in the declared order and a re-sorted schema silently changes model behavior (#6235) +- fix: open reasoning stream items that carry both an encrypted payload and a visible summary as `thinking` blocks instead of `redacted_thinking` on the Anthropic egress, with `isReasoningItem` and `reasoningPayloadAndSummary` shared by the native-reasoning and misclassified-function-call branches (#6292) +- fix: replayed thinking blocks through the Anthropic ingress with a `bedrock/` model prefix: content-less `tool_result` blocks are kept, interleaved text/tool-use/thinking order is preserved by the grouped converter, `incomplete` maps to `error` on Converse `toolResult.status`, and buffered reasoning is consumed by the item that owns it, so multi-turn tool use no longer wedges (#6346) +- fix: Gemini/Vertex HTTP 400s on Claude Code traffic routed through `/anthropic/v1/messages`: trailing assistant prefills are trimmed on both the Responses and chat paths, mid-conversation `system` messages are inlined in place instead of hoisted into `systemInstruction`, and `AnthropicMessageResponse` gains `ExtraFields` (#6363) +- fix: alias Bedrock `toolUseId`/`toolResultId` values longer than 64 characters or outside `[a-zA-Z0-9_.:-]` (such as Gemini thought-signature IDs) with a deterministic hash prefix, applied identically on `tool_use` and `tool_result` in both the Responses and chat converters (#6300) +- fix: route Grok (`xai.`) models through the `openai/v1` Mantle path on Bedrock and Bedrock Mantle, since they have no Converse equivalent (#6022) +- fix: register Bedrock Mantle in `ProviderSendsDoneMarker` so its streams end after `finish_reason` instead of waiting for a `[DONE]` marker (#6021) +- fix: include OpenRouter embedding models from `/v1/embeddings/models` in `ListModels`, merged case-insensitively and best-effort (#6264) +- fix: force `reasoning.effort` to `"none"` for models that reason by default but do not support reasoning with tool calls when they advertise `supports_none_reasoning_effort`, instead of dropping `reasoning` outright (#6293) +- fix: backfill upscale output resolution on Replicate from the `target`/`factor` params and `metrics.resolution_target` bands so resolution-tiered pricing bills the real output size (#6083) +- fix: filter forwarded `Accept-Encoding` to the codecs `CheckAndDecodeBody` can decode (`gzip`, `x-gzip`, `deflate`, `br`, `zstd`, `identity`), restrict streaming endpoints to `gzip`/`identity` via `SetPassthroughHeadersForStreaming`, and decode chained content encodings in reverse order (#6360) +- fix: `tool_sync_interval` handling: negative values are rejected (the "disable sync" semantic is gone now that the connection checker drives discovery and liveness together), `ResolveToolSyncInterval` follows the global setting for sub-second values, a fresh per-call checker starts on `EnableClient` and on a sticky-to-per-call flip, and `MCPManager.UpdateToolSyncInterval`, `ConnectionCheckerManager.SetGlobalInterval`/`ApplyGlobalInterval`/`RetimeClient` and `ClientConnectionChecker.SetHealthyInterval` hot-reload the global cadence and re-time running checkers in place; `GetMCPConfig` carries the stored global interval (#6502) +- fix: `SetClientTools` and `UpdateClientCredentials` replace the MCP tool map instead of `maps.Copy`-merging into it, so a tool removed upstream is evicted from memory once the database has dropped it (#6484) +- fix: per-call shared-credential MCP clients (`oauth`, `headers`, `none`) refresh tools synchronously on credential update instead of returning `ErrMCPReconnectNotApplicable`; disabled per-call clients and per-user auth types keep the sentinel (#6483) +- fix: park a failed `EnableClient` dial at `Disabled` instead of `Unstable`, add `ErrMCPEnableConnectFailed` so callers do not roll back the persisted `disabled` flag, and guard `isEnableable` on both state and config so the admin can retry (#6431) +- fix: `output_item.done` replaces server-side tool item shells (`web_search_call`, `code_interpreter_call`, `image_generation_call`) in the Responses streaming accumulator so their full payload survives (#6475) +- feat: send `s3://` image and document references to Bedrock Converse as the `s3Location` source member instead of downloading the bytes and re-uploading them - Converse resolves the object itself, which skips a round trip and the 25 MiB inline cap entirely. Image format is derived from the object extension since nothing is fetched and there is no `Content-Type` to read, and an extension-less object is rejected up front rather than producing an opaque 400 (#6239) +- feat: resolve Vertex URL sources per model family rather than inlining everything - a `gs://` URI is now forwarded to Gemini/Gemma as `fileData.fileUri` (the documented form, resolved under the caller's own project IAM, and the only thing that keeps multi-hundred-MB video inputs viable) and read from Cloud Storage with the request key's own Google credentials for Claude-on-Vertex, which accepts base64 sources only. `http(s)` is still always fetched: forwarding one was measured against the harness and Vertex rejected every endpoint shape with `URL_REJECTED-REJECTED_FC_TOO_MANY_PENDING` (#6239) +- fix: always emit a Gemini candidate carrying its finish reason on `generateContent`, even when nothing visible was generated - a thinking model that spends its whole output budget before emitting a token is a successful 200 with an empty answer, but `Candidates` is `omitempty`, so dropping that candidate produced a body with no `candidates` key at all and left a lone `usageMetadata` object that every Gemini-shaped client dereferences blind (#6239) +- fix: drop payload-free Gemini parts when assembling a candidate - every `Part` field is `omitempty`, so such a part marshals to exactly `{}`; the harness observed one on the wire when a transcription request for an unintelligible tone came back as `parts:[{}]`, where it is noise a client will try to read and it masks the contentless case by making the parts slice look non-empty (#6239) +- fix: accept a bare model identifier on Bedrock rerank by synthesizing the foundation-model ARN from the resolved region - Rerank is the one Bedrock surface that names its model by ARN rather than by bare ID, so all three rerank drop-ins in the provider harness 400'd on `amazon.rerank-v1:0`. The partition is derived from the region (`aws`, `aws-cn`, `aws-us-gov`) so GovCloud and China build a correct ARN, and an explicit ARN still passes through untouched (#6239) +- fix: stop stripping `file_url` from OpenAI-shaped chat file blocks on marshal - dropping it produced `{"type":"file","file":{}}` and an upstream complaint about a missing `file_id`, which hid the fact that a source had been discarded. Providers that cannot take a URL now say so by name, and any OpenAI-compatible endpoint that does accept one keeps working without a Bifrost change (#6239) +- fix: leave URL content sources Bifrost cannot download in place on the OpenAI and native-Anthropic paths instead of failing the request - only `http(s)` is fetched, and whether a `gs://`, `s3://` or scheme-less reference is usable is the provider's call, so the source now travels as `{"type":"url"}` and the platform answers for itself (#6239) +- perf: JSON serialization on the hot path: shared MCP tools cache their serialized bytes on `ChatTool` (`EnsureSerialized`, `precomputeToolSerialization`) so catalog tools are marshalled once, and `OrderedMap.MarshalJSON` writes compact JSON directly into a buffer with an inline HTML-safe string escaper instead of re-routing every nested map through `MarshalSorted`, pinned by a byte-identity fuzz harness (#6241, #6242) +- perf: allocation and tracing reductions on the request path: `StartSpanID` and `SpanFromHandle` on the tracer with `Span.SetAttributes` for bulk writes, a resolved-once attribute block in `executeRequestWithRetries`, reusable worker delivery timers, `Span.Reset` keeping map capacity, `reservedKeys` as a set, pre-sized `userValues`, logging context reads deferred to the final chunk, no redundant `fmt.Sprintf` in logger calls, cached plugin span names, compact JSON request bodies, `math/rand/v2` in `GetRandomString`, and a `HasPluginLogs` guard before draining plugin logs (#5657, #5956, #5957, #6211) +- chore: remove the legacy `gen_ai.*`-namespaced Bifrost-internal attribute constants, `AttrPromptTokens`/`AttrCompletionTokens`, `AttrLegacyRetryCount` and the nanosecond `AttrTimeToFirstToken` in favor of the canonical `bifrost.*` keys (#6403) +- chore: build with Go 1.26.6 (#6269) diff --git a/core/version b/core/version index 36c5cb9efa2..27f9cd322bb 100644 --- a/core/version +++ b/core/version @@ -1 +1 @@ -1.7.13 +1.8.0 diff --git a/framework/changelog.md b/framework/changelog.md index 515b3eb56a7..9d5861b2872 100644 --- a/framework/changelog.md +++ b/framework/changelog.md @@ -1,30 +1,33 @@ -- feat: add `cost_per_request` flat-fee pricing field across DB, cost engine, overrides and docs (#6079) -- feat(modelcatalog): resolve pricing overrides for catalog rows (#6055) -- feat: add `use_idp_credentials` to token-exchange config (#6068) -- feat: bedrock vpc endpoints support (#6064) -- feat: add additional metadata in S3 log export (#6070) -- feat: make log recalculation task cancellable backend (#5801) -- feat: add `roots_only` filter to collapse fallback chains with child aggregates (#5737) -- feat: support matview_refresh_interval "off" to disable logstore matview maintenance (thanks [@jeremym-tanium](https://github.com/jeremym-tanium)!) (#5693) -- feat: persist and resync MCP tool discoveries uniformly across all client types via a hash-gated core callback -- feat: add VK and Users filters to the OAuth Grants and MCP Auth Sessions sidebars -- feat: generalize TokenRefreshWorker's auth-mode scope and allow gating OAuthTokenRefreshWorker sweeps -- feat(mcp-guardrails): add MCP log redaction changes (#5744) -- feat: add plugin logs to mcp logs (#5746) -- fix: combine `offline_access` with `/.default` for Entra OBO instead of replacing it (#6078) -- fix: don't treat a CAS loss to a still-active concurrent refresh as a dead credential -- fix: propagate ctx through headerCredentialCache.Fill and userTokenCache.Fill so a canceled request unblocks instead of waiting on an unrelated leader -- fix: add per-entry version to the LRU cache so a rejected stale Get cannot evict a concurrently-updated value -- fix: make the OAuth flow claim atomic against concurrent reauth, close a leaked sqlDB in flows-table perf setup -- fix: route pending token_exchange clients through the verify-exchange confirm dialog -- chore: dependabot dependency updates (#6040) +- feat: batch accounting: the `batch_jobs` table and its lifecycle store API (`UpsertBatchJob`, `GetBatchJob`, `ListDueBatchJobs`, `ClaimBatchJob`, `MarkBatchJobAggregateLogWritten`, `MarkBatchJobGovernanceReported`, `CompleteBatchJob`, `MarkBatchJobUnpriceable`, `FailBatchJob`) with runner fencing on `claimed_at` and `user_id`, `team_id`, `customer_id` and `source_log_id` attribution so settlement carries the creating request's identity; `batch_debug` on logs; batch pricing in the model catalog (`computeBatchTextCost` with catalog batch rates and a 0.5 default ratio, `CalculateBatchCostDetailsForUsage`, `BatchResultsRequest` routed through the batch path); and `persistRecalcOutcomes` shared by foreground and background recalculation (thanks [@SahilChoudhary22](https://github.com/SahilChoudhary22)!) (#5292, #5293, #6505) +- feat: input/output/additional cost split: denormalized `input_cost`, `output_cost` and `additional_cost` columns on logs, carried through matviews, ClickHouse, the hybrid store, cost recalculation and the quota API, populated on fallback billing paths, with semantic cache cost folded into additional cost +- feat: Bifrost overhead latency: `upstream_latency` and `overhead_latency` columns on logs with avg, p90, p95 and p99 overhead aggregates in `mv_logs_hourly`, ClickHouse, Postgres `percentile_cont` and the Go-side SQLite/MySQL histograms, the `overhead_breakdown` column for the per-span self-time decomposition, and `CompleteAndFlushTrace` handing connectors a copy of the trace without breakdown spans unless the plugin implements `OverheadSpanConsumer` (#5533, #5534, #6388, #6389) +- feat: `video_edit_input` column on logs for the new video edit request type (#6270) +- feat: new pricing columns and cost computation: megapixel-tier image fields (`output_cost_per_image_above_{4,8,16,32,64}_megapixels`) with a unified pixel-count tier ladder in `computeImageOutputCost`; per-size and joint size+quality image rates for 1024x1536 and 1536x1024 with a priority chain of size+quality, quality-only, size-only, then flat per-image rate, and `parseImageDimensions` so portrait and landscape sizes with equal pixel counts price correctly; `input_cost_per_query` for rerank; and `ultrafast` service tier rates (#6082, #6379, #6396) +- feat: notifications store: `TableNotification`, `NotificationStore`, `CreateNotification` and `ListNotifications` with JSON-serialized role IDs (#6207) +- feat: `gencache` generation-stamped memo cache, with `GetProvidersForModel` and `GetModelsForProvider` memoized until any backing store advances its write generation (#5641, #6224) +- feat: `DimensionScope` in `queryscope` and `applyDimensionCeiling` on rankings, histograms and key-pair queries so grouped analytics only expose organisation ids the caller may see; `getAvailableFilterData` no longer passes an empty id list to the redaction lookups, which returned every row (#6262) +- feat: `ObservabilityLimits` (per-plugin semaphore size and inject timeout) with context-bounded `Inject` calls and `DeadlineExceeded` accounting (#6341) +- feat: `GetSharedOauthTokensByConfigIDs` batch lookup on the config store so shared-OAuth MCP clients project `needs_reauth` when their token row is invalidated (#6429) +- feat: `mcp_library_sync_interval: 0` disables MCP library sync (`MCPLibrarySyncDisabled`), `file://` catalog URLs resolve through `datasheet.FilePathFromURL` without retry backoff, and `ResolveFrameworkPricingConfig` no longer backfills a zero interval (#6195) +- feat: `ReloadComplexityAnalyzerConfig` on `ServerCallbacks` for the routing handler (#6146) +- feat: `service_tier` copied from the processed stream response into `StreamAccumulatorResult` in `ProcessStreamingChunk` (#6236) +- fix: resolve runtime provider `together` (and variants such as `together_ai`, matched with `strings.Contains`) to the datasheet identity for catalog reads and price configured aliases through `AliasConfig.ModelName`, then `ModelID`, then the alias key (thanks [@dani29](https://github.com/dani29)!) (#6257, #6320) +- fix: escape every RediSearch special character in TAG query values in the Redis vector store, iterating bytes rather than runes (thanks [@AdityaPainuli](https://github.com/AdityaPainuli)!) (#5351) +- fix: redact sensitive and identity-aware-proxy request headers at `SetTraceRequestHeaders` so every connector (Datadog, OTEL, BigQuery, Kafka, Pub/Sub) receives redacted values (#6371) +- fix: `supports_none_reasoning_effort` datasheet flag wired through `extractSupportedParams` and `dropUnsupportedParams` so models that reason by default get `reasoning.effort: "none"` instead of losing `reasoning` (#6293) +- fix: reject negative `tool_sync_interval` at `UpdateMCPClientConfig` and on config file load, treat the value as whole minutes, and carry the stored global interval into `GetMCPConfig` (#6409, #6502) +- perf: `spanHandle` carries the `*Span` pointer so `EndSpan`, `SetAttribute` and `SpanFromHandle` skip the per-call trace and span scan, alongside bulk span attribute writes and reusable delivery timers in the tracing hot path (#5657, #5956, #6387) +- chore: remove legacy `gen_ai.*` attribute emission from the tracer in favor of the canonical `bifrost.*` keys (#6403) +- chore: close leaked Postgres pools and a stale hardcoded date in logstore tests (#6351) +- chore: build with Go 1.26.6 (#6269) +- chore: upgraded core to v1.8.0 -This release adds 18 database migrations. `merge_oauth_token_tables`, `drop_oauth_config_pkce_columns`, `drop_oauth_config_token_id_column` and `mcp_tool_logs_add_redaction_mapping_column` are non-reversible. Back up your database before upgrading. +This release adds 13 database migrations (7 configstore, 6 logstore). All are additive and reversible: each rollback drops the column, table or index it created, and `logs_recreate_matviews_with_cost_breakdown` is a no-op both ways because `repairMatViewShapes` rebuilds `mv_logs_hourly` on the next startup. **High-throughput deployments: run the logstore migrations during a low-activity window.** -All eight logstore migrations in this release alter `logs` or `mcp_tool_logs`, the two highest-insert tables in Bifrost, and several also build indexes on them. On a busy instance those index builds block concurrent log inserts until they complete. Schedule the upgrade for a low-traffic period, or expect elevated log-write latency while the migrations run. +Five of the six logstore migrations alter `logs`, the highest-insert table in Bifrost, and the hourly matview is rebuilt against the full table on the first boot after upgrading. Schedule the upgrade for a low-traffic period, or expect elevated log-write latency while the migrations run. diff --git a/framework/vectorstore/redis_test.go b/framework/vectorstore/redis_test.go index db90b1b612b..4d8219b1f03 100644 --- a/framework/vectorstore/redis_test.go +++ b/framework/vectorstore/redis_test.go @@ -1515,16 +1515,28 @@ func TestRedisStore_VectorSearch(t *testing.T) { require.NoError(t, err) } - time.Sleep(500 * time.Millisecond) - + // Poll rather than sleep a fixed interval: RediSearch makes a new document + // searchable slightly after the write returns, and under `go test ./...` + // (many packages sharing one Redis) a fixed 500ms was occasionally too short, + // so the search returned nothing. A query error still fails immediately, + // since an unescaped value produces a syntax error, not a delay. for _, doc := range specialDocs { queries := []Query{ {Field: "model", Operator: QueryOperatorEqual, Value: doc.model}, } - results, err := setup.Store.GetNearest(setup.ctx, TestNamespace, doc.embedding, queries, []string{"type", "model"}, 0.1, 10) - require.NoError(t, err, "search must not fail for model %q", doc.model) - require.Len(t, results, 1, "expected exactly the doc tagged %q", doc.model) - assert.Equal(t, doc.model, results[0].Properties["model"]) + deadline := time.Now().Add(5 * time.Second) + for { + results, err := setup.Store.GetNearest(setup.ctx, TestNamespace, doc.embedding, queries, []string{"type", "model"}, 0.1, 10) + require.NoError(t, err, "search must not fail for model %q", doc.model) + if len(results) == 1 { + assert.Equal(t, doc.model, results[0].Properties["model"]) + break + } + if time.Now().After(deadline) { + require.Failf(t, "tagged doc not searchable", "expected exactly the doc tagged %q, got %d results after 5s", doc.model, len(results)) + } + time.Sleep(100 * time.Millisecond) + } } }) } diff --git a/framework/version b/framework/version index f0ed37967c0..dc1e644a101 100644 --- a/framework/version +++ b/framework/version @@ -1 +1 @@ -1.5.10 +1.6.0 diff --git a/plugins/compat/changelog.md b/plugins/compat/changelog.md index bb1a0e1ca9a..2966d0ad9bc 100644 --- a/plugins/compat/changelog.md +++ b/plugins/compat/changelog.md @@ -1 +1,4 @@ -- chore: upgraded core to v1.7.11 and framework to v1.5.9 +- fix: force `reasoning.effort` to `"none"` in `dropUnsupportedParams` when a model supports reasoning but not `reasoning_with_tool_calls` and advertises `supports_none_reasoning_effort`; models without the flag still have `reasoning` dropped (#6293) +- fix: clone `json.RawMessage` values (such as a raw `response_format`) in the request copier so the compat clone never shares a backing array with the original request (#6235) +- feat: add a no-op `HTTPTransportPreAuthHook` for the new pre-authentication transport phase (#6375) +- chore: upgraded core to v1.8.0 and framework to v1.6.0 diff --git a/plugins/compat/version b/plugins/compat/version index 072d0fa39ef..0ea3a944b39 100644 --- a/plugins/compat/version +++ b/plugins/compat/version @@ -1 +1 @@ -0.1.36 +0.2.0 diff --git a/plugins/governance/changelog.md b/plugins/governance/changelog.md index 35d560a2ebd..fbcf53e3272 100644 --- a/plugins/governance/changelog.md +++ b/plugins/governance/changelog.md @@ -1,3 +1,5 @@ -- fix: skip list models call for budgets and rate-limits (#6051) -- feat: honor the auth-skip context path in the governance resolver -- chore: upgraded core to v1.7.11 and framework to v1.5.9 +- feat: routing rules and the complexity router are extracted into the dedicated `routing` plugin: governance now runs at priority 4 and routing at 5, `PublishRoutingAllowlist` and `LoadBalanceProvider` are exported on `GovernancePlugin` and `BaseGovernancePlugin` and are called from the routing plugin after rule evaluation instead of from governance's `PreRequestHook`, `runPreRequestRouting` is removed, and `ReloadRoutingRule`, `RemoveRoutingRule` and the routing rule and complexity analyzer handlers leave `GovernanceManager` and `GovernanceHandler` for `RoutingHandler` under `/api/routing/*` (with deprecated `/api/governance/*` aliases) (#6144, #6145, #6146) +- feat: batch usage reporting: `ReportBatchUsage` applies settled batch cost, tokens and requests to every budget and rate limit on a `BatchUsageReport` exactly once per request ID via a claim/release marker with a 7-day TTL, and charges the creating user's tiers when `UserID` is present; `BumpBudgetUsage` and `BumpRateLimitUsage` are added to `GovernanceStore`; governance IDs, including VK-scoped, user-scoped and global wildcard budgets and rate limits, are collected for batch-create requests that carry no model (thanks [@SahilChoudhary22](https://github.com/SahilChoudhary22)!) (#5295, #6410, #6505) +- feat: honor `BifrostContextKeySkipModelCheck` in `EvaluateVirtualKeyRequest` so evaluate-only requests such as `/inspect` bypass the model allowlist while budgets, rate limits and provider checks still apply (#6479) +- feat: add a no-op `HTTPTransportPreAuthHook` for the new pre-authentication transport phase (#6375) +- chore: upgraded core to v1.8.0 and framework to v1.6.0 diff --git a/plugins/governance/version b/plugins/governance/version index 5577648bd3d..bd8bf882d06 100644 --- a/plugins/governance/version +++ b/plugins/governance/version @@ -1 +1 @@ -1.6.14 +1.7.0 diff --git a/plugins/jsonparser/changelog.md b/plugins/jsonparser/changelog.md index bb1a0e1ca9a..7a7dc123de2 100644 --- a/plugins/jsonparser/changelog.md +++ b/plugins/jsonparser/changelog.md @@ -1 +1,2 @@ -- chore: upgraded core to v1.7.11 and framework to v1.5.9 +- feat: add a no-op `HTTPTransportPreAuthHook` for the new pre-authentication transport phase (#6375) +- chore: upgraded core to v1.8.0 and framework to v1.6.0 diff --git a/plugins/jsonparser/version b/plugins/jsonparser/version index dc848ff4c54..dc1e644a101 100644 --- a/plugins/jsonparser/version +++ b/plugins/jsonparser/version @@ -1 +1 @@ -1.5.37 +1.6.0 diff --git a/plugins/logging/changelog.md b/plugins/logging/changelog.md index 1f4ac482081..0b992116aef 100644 --- a/plugins/logging/changelog.md +++ b/plugins/logging/changelog.md @@ -1,7 +1,8 @@ -- feat: make log recalculation task cancellable backend (#5801) -- feat: add `roots_only` filter to collapse fallback chains with child aggregates (#5737) -- feat: add plugin logs in mcp logs (#5746) -- feat(mcp-guardrails): add MCP log redaction changes (#5744) -- feat: video requests info in logs ui (#5946) -- feat: cost for prompt guardrails (#4931) -- chore: upgraded core to v1.7.11 and framework to v1.5.9 +- feat: batch accounting: `recordBatchJobLifecycle` persists batch state on create and retrieve, `accountBatchResults` settles costs inline on the `/results` path under a 30-second bound, `StartBatchAccountingSweeper` re-drives jobs that timed out with a per-runner ownership identity, `EmitBatchAggregateLog` writes the aggregate cost entry with the creating request's identity and a `bifrost/` user agent, `calculateBatchAggregateCost` reprices `Model="mixed"` rows per model breakdown during cost recalculation (foreground and background, via the shared `persistRecalcOutcomes`), and `batch_debug` is included in list queries; `Init` takes a `batchStore` (nil disables batch accounting) (thanks [@SahilChoudhary22](https://github.com/SahilChoudhary22)!) (#5296, #6121, #6474, #6505) +- feat: input/output/additional cost split persisted on every log and surfaced as `cost_breakdown` in the log detail API, with cached-read, reasoning, guardrail, MCP and semantic cache detail; fallback billing paths, speech, transcription and OCR usages populate the split, and legacy total-only rows are attributed to input cost (#6511) +- feat: Bifrost overhead latency: `upstream_latency` and `overhead_latency` forwarded from `PostLLMHook` and backfilled from the root span's authoritative attributes in `Inject`, stamped only on the terminal entry per trace; `computeOverheadBreakdown` walks the span tree, computes self-time per span, groups overhead-side spans into buckets (serialization, middleware, plugins, queue wait, key selection, convertor, networking, client delivery, scheduling, worker hand-off, provider-internal) and persists them to `overhead_breakdown`, with streaming traces folding parse, convert, backpressure, transport CPU and client-write time into their own buckets and using the measured sum as overhead; `ConsumesOverheadSpans` returns true so the plugin keeps receiving breakdown spans that other connectors no longer see (#5533, #6388, #6389, #6433, #6470, #6495) +- feat: `service_tier` from streamed Anthropic responses flows through `convertToProcessedStreamResponse` into the log entry so repricing uses the served tier (#6236) +- feat: video edit requests are logged with their input (#6270) +- perf: identity and governance context reads are deferred past the non-final-chunk gate in `PostLLMHook`, and JSON encoding in HTTP helpers uses sonic (#5957, #6268) +- feat: add a no-op `HTTPTransportPreAuthHook` for the new pre-authentication transport phase (#6375) +- chore: upgraded core to v1.8.0 and framework to v1.6.0 diff --git a/plugins/logging/version b/plugins/logging/version index 1df3b822c71..bd8bf882d06 100644 --- a/plugins/logging/version +++ b/plugins/logging/version @@ -1 +1 @@ -1.6.10 +1.7.0 diff --git a/plugins/maxim/changelog.md b/plugins/maxim/changelog.md index bb1a0e1ca9a..eccbf9706f9 100644 --- a/plugins/maxim/changelog.md +++ b/plugins/maxim/changelog.md @@ -1 +1,4 @@ -- chore: upgraded core to v1.7.11 and framework to v1.5.9 +- feat: `addLatencyTags` forwards `upstream_latency_ms` and `overhead_latency_ms` as tags on both the generation and the trace, leaving unmeasured values unreported (#6345) +- fix: sensitive and identity-aware-proxy request headers are redacted in `PostLLMHook` before export (#6371) +- feat: add a no-op `HTTPTransportPreAuthHook` for the new pre-authentication transport phase (#6375) +- chore: upgraded core to v1.8.0 and framework to v1.6.0 diff --git a/plugins/maxim/version b/plugins/maxim/version index f886f96703f..bd8bf882d06 100644 --- a/plugins/maxim/version +++ b/plugins/maxim/version @@ -1 +1 @@ -1.6.37 +1.7.0 diff --git a/plugins/mocker/changelog.md b/plugins/mocker/changelog.md index bb1a0e1ca9a..7a7dc123de2 100644 --- a/plugins/mocker/changelog.md +++ b/plugins/mocker/changelog.md @@ -1 +1,2 @@ -- chore: upgraded core to v1.7.11 and framework to v1.5.9 +- feat: add a no-op `HTTPTransportPreAuthHook` for the new pre-authentication transport phase (#6375) +- chore: upgraded core to v1.8.0 and framework to v1.6.0 diff --git a/plugins/mocker/version b/plugins/mocker/version index dc848ff4c54..dc1e644a101 100644 --- a/plugins/mocker/version +++ b/plugins/mocker/version @@ -1 +1 @@ -1.5.37 +1.6.0 diff --git a/plugins/modelcatalogresolver/changelog.md b/plugins/modelcatalogresolver/changelog.md index bb1a0e1ca9a..f640b0af0a4 100644 --- a/plugins/modelcatalogresolver/changelog.md +++ b/plugins/modelcatalogresolver/changelog.md @@ -1 +1 @@ -- chore: upgraded core to v1.7.11 and framework to v1.5.9 +- chore: upgraded core to v1.8.0 and framework to v1.6.0 diff --git a/plugins/modelcatalogresolver/version b/plugins/modelcatalogresolver/version index f8f3c08725d..9084fa2f716 100644 --- a/plugins/modelcatalogresolver/version +++ b/plugins/modelcatalogresolver/version @@ -1 +1 @@ -1.0.18 +1.1.0 diff --git a/plugins/otel/changelog.md b/plugins/otel/changelog.md index d5dd03d220b..e913d748ded 100644 --- a/plugins/otel/changelog.md +++ b/plugins/otel/changelog.md @@ -1,3 +1,7 @@ -- feat: add separate headers support for traces and metrics in OTEL collector (#5940) -- feat: add support for a separate metrics tab independent of traces for OTEL (#5939) -- chore: upgraded core to v1.7.11 and framework to v1.5.9 +- feat: `bifrost_overhead_latency_microseconds` histogram derived from the root span's overhead attribute, with a microsecond-scale bucket set (#6345) +- chore: remove the legacy `gen_ai.*`-namespaced Bifrost-internal attributes, `gen_ai.usage.prompt_tokens`/`completion_tokens` and the nanosecond `time_to_first_token` attribute; `buildSpanAttrs` and `entitySetFromAttrs` read the canonical `bifrost.*` keys and `time_to_first_chunk` directly (#6403) + + Dashboards and alerts that read the legacy `gen_ai.*` Bifrost-internal attributes or the nanosecond TTFT attribute must migrate to the `bifrost.*` keys and `time_to_first_chunk` (seconds). + +- feat: add a no-op `HTTPTransportPreAuthHook` for the new pre-authentication transport phase (#6375) +- chore: upgraded core to v1.8.0 and framework to v1.6.0 diff --git a/plugins/otel/version b/plugins/otel/version index 4ea2b1f403a..bc80560fad6 100644 --- a/plugins/otel/version +++ b/plugins/otel/version @@ -1 +1 @@ -1.4.9 +1.5.0 diff --git a/plugins/prompts/changelog.md b/plugins/prompts/changelog.md index bb1a0e1ca9a..7a7dc123de2 100644 --- a/plugins/prompts/changelog.md +++ b/plugins/prompts/changelog.md @@ -1 +1,2 @@ -- chore: upgraded core to v1.7.11 and framework to v1.5.9 +- feat: add a no-op `HTTPTransportPreAuthHook` for the new pre-authentication transport phase (#6375) +- chore: upgraded core to v1.8.0 and framework to v1.6.0 diff --git a/plugins/prompts/version b/plugins/prompts/version index f6c5328c425..9084fa2f716 100644 --- a/plugins/prompts/version +++ b/plugins/prompts/version @@ -1 +1 @@ -1.0.37 +1.1.0 diff --git a/plugins/routing/changelog.md b/plugins/routing/changelog.md index e69de29bb2d..ef329733cf2 100644 --- a/plugins/routing/changelog.md +++ b/plugins/routing/changelog.md @@ -0,0 +1,2 @@ +- feat: initial release: the routing rules engine (`rules/`) and the complexity router (`complexity/`) are extracted from the governance plugin into a dedicated routing plugin that depends on governance through a small `Governance` interface; it runs at priority 5, after governance has stamped the virtual key scope, and calls `PublishRoutingAllowlist` and `LoadBalanceProvider` after rule evaluation so both act on the post-rule model (the `HasRules` early return moved into `applyRoutingRules` so provider materialization still runs with no rules configured); routing rules and complexity analyzer config endpoints are served by `RoutingHandler` at `/api/routing/rules` and `/api/routing/complexity-analyzer-config`, with the legacy `/api/governance/*` paths registered as deprecated aliases on the same handlers (#6144, #6145, #6146) +- feat: complexity routing extracts text from mixed-modality user turns (text plus image, file or audio blocks) instead of skipping the turn, and still produces no input for turns with no text at all (#6253) diff --git a/plugins/semanticcache/changelog.md b/plugins/semanticcache/changelog.md index 69e709aad3d..7a7dc123de2 100644 --- a/plugins/semanticcache/changelog.md +++ b/plugins/semanticcache/changelog.md @@ -1,2 +1,2 @@ -- feat: account for prompt guardrail cost in cache search (#4931) -- chore: upgraded core to v1.7.11 and framework to v1.5.9 +- feat: add a no-op `HTTPTransportPreAuthHook` for the new pre-authentication transport phase (#6375) +- chore: upgraded core to v1.8.0 and framework to v1.6.0 diff --git a/plugins/semanticcache/version b/plugins/semanticcache/version index dc848ff4c54..dc1e644a101 100644 --- a/plugins/semanticcache/version +++ b/plugins/semanticcache/version @@ -1 +1 @@ -1.5.37 +1.6.0 diff --git a/plugins/telemetry/changelog.md b/plugins/telemetry/changelog.md index bb1a0e1ca9a..752552bf6d1 100644 --- a/plugins/telemetry/changelog.md +++ b/plugins/telemetry/changelog.md @@ -1 +1,7 @@ -- chore: upgraded core to v1.7.11 and framework to v1.5.9 +- feat: `bifrost_overhead_latency_microseconds` histogram measured across the HTTP transport hooks as total request time minus time blocked on upstream provider sockets, with a microsecond-scale bucket set (#6345) +- chore: `x-bf-prom-*` request headers are no longer consumed as Prometheus label dimensions in `collectPrometheusKeyValues` and `applyCustomLabels` (the prefix is still stripped from forwarded requests), and legacy `gen_ai.*` attribute emission is removed (#6403) + + Deployments that relied on `x-bf-prom-*` request headers to add Prometheus label dimensions lose those labels; use the supported custom label configuration instead. + +- feat: add a no-op `HTTPTransportPreAuthHook` for the new pre-authentication transport phase (#6375) +- chore: upgraded core to v1.8.0 and framework to v1.6.0 diff --git a/plugins/telemetry/version b/plugins/telemetry/version index dc848ff4c54..dc1e644a101 100644 --- a/plugins/telemetry/version +++ b/plugins/telemetry/version @@ -1 +1 @@ -1.5.37 +1.6.0 diff --git a/transports/changelog.md b/transports/changelog.md index 5cb525bfab4..0b71e510371 100644 --- a/transports/changelog.md +++ b/transports/changelog.md @@ -1,52 +1,171 @@ + +v2.0.0 is the first stable release on the 2.0 line. This changelog rolls up `2.0.0-prerelease1` (based on [v1.6.3](https://docs.getbifrost.ai/changelogs/v1.6.3)), `2.0.0-prerelease2`, `2.0.0-prerelease3` and the final release window, so it is the complete delta for a deployment upgrading from any v1.6.x release. Fixes that also shipped on the v1.6.x line after v1.6.3 are listed once here. + + + +**Breaking changes.** Read the [v2.0.0 migration guide](https://docs.getbifrost.ai/migration-guides/v2.0.0) before upgrading. + +- **Custom plugin downloads are SSRF-protected** - a plugin `path` pointing at an http(s) URL is rejected if it resolves to a loopback, private, CGNAT, link-local or otherwise non-public address, and every custom plugin path is re-verified on each restart, including ones defined in `config.json`. +- **Custom plugin create and update require admin authentication** - `POST /api/plugins` and `PUT /api/plugins/{name}` reject a custom `path` when the caller only got through because dashboard auth is disabled or unconfigured. +- **Governance APIs moved under `/api/governance/*`** - `/api/teams`, `/api/users`, `/api/roles`, `/api/audit-logs` and other top-level governance paths moved under one namespace; Team and User lists use `limit`/`offset` pagination. Routing rules and the complexity analyzer moved from `/api/governance/*` to `/api/routing/rules` and `/api/routing/complexity-analyzer-config`; the old paths remain as deprecated aliases. +- **`HTTPTransportPreHook` now runs after authentication** - the pipeline is `HTTPTransportPreAuthHook -> auth -> HTTPTransportPreHook -> handler`. Plugins that inject a credential (`x-bf-vk`, `Authorization`, `x-api-key`) must move that work to the new `HTTPTransportPreAuthHook`, and Go plugins implementing `HTTPTransportPlugin` must add the method (`.so` plugins that predate it are skipped for that phase). +- **Legacy telemetry attributes removed** - the `gen_ai.*`-namespaced Bifrost-internal span attributes, `gen_ai.usage.prompt_tokens`/`completion_tokens`, the nanosecond `time_to_first_token` attribute and `x-bf-prom-*` request-header Prometheus dimensions are gone from the OTel and Prometheus connectors. Dashboards should read the `bifrost.*` keys and `time_to_first_chunk`. +- **Gemini tool preference** - a Gemini API request carrying both function declarations and Google Search without `include_server_side_tool_invocations` now keeps the function declarations and drops Google Search (previously the opposite). Set `include_server_side_tool_invocations: true` to send both on Gemini 3 models. Vertex is unaffected. + + ## โœจ Features -- **MCP Per-User OAuth** - MCP clients can hold per-user OAuth credentials and per-user headers, configurable from `config.json` as well as the UI, with a documented shared vs per-identity token lookup contract and VK/Users filters on the OAuth Grants and MCP Auth Sessions sidebars -- **Token Exchange IDP Credentials** - New `use_idp_credentials` on `token_exchange` reuses SSO login app credentials for providers that require it, such as Microsoft Entra ID; `client_id` becomes optional when it is set (#6068, #6069) +- **Batch Accounting** - Provider batch jobs are tracked in a new `batch_jobs` table and settled asynchronously: results are priced per model from catalog batch rates (0.5 default ratio) on the `/results` path, one aggregate cost log is written idempotently with the creating request's identity, a background sweeper with ownership fencing re-drives jobs that timed out, settled usage is charged exactly once to the creating user's budgets and rate limits (including unscoped virtual key budgets on model-less batch-create requests), mixed-model batch rows are repriced during cost recalculation, and the log detail view shows a Batch Details block with per-state request counts and the settled cost (#5291, #5292, #5293, #5294, #5295, #5296, #6109, #6121, #6376, #6410, #6474, #6505) +- **Claude-on-Vertex Batches** - Vertex batch jobs route Anthropic models to `publishers/anthropic/...`, build Claude-on-Vertex JSONL instances, round-trip `custom_id`, and preserve `tools`, `toolConfig`, `cachedContent`, `labels` and `display_name` on Gemini/Vertex batch requests (#5368) +- **Input / Output Cost Split** - Every log carries `input_cost`, `output_cost` and `additional_cost` (guardrails, semantic cache, MCP) next to the total, across the RDB, ClickHouse, matviews, recalculation and the quota API; speech, transcription and OCR usages carry `BifrostCost`; the log detail view shows the split with per-category detail (#6511) +- **Bifrost Overhead Latency** - `upstream_latency` and `overhead_latency` are recorded on every log, aggregated (avg, p90, p95, p99) in the dashboard's new Bifrost Overhead chart and shown in the log detail view; the overhead is decomposed by span self-time into serialization, conversion, plugins, middleware, key selection, queue wait, networking, client delivery and scheduling buckets (including streaming per-chunk parse, conversion and backpressure and the worker hand-off), persisted to `overhead_breakdown` and rendered as a stacked bar in the log detail view; a `bifrost_overhead_latency_microseconds` histogram is exported to Prometheus and OpenTelemetry and `upstream_latency_ms`/`overhead_latency_ms` tags to Maxim, while breakdown spans are kept out of observability connectors (#5533, #5534, #5535, #6345, #6388, #6389, #6433, #6470, #6495) +- **Notification Center** - Role-targeted dashboard notifications stored in the database, delivered over WebSocket and surfaced in a topbar tray via `GET/POST /api/notifications` (#6207, #6227, #6324) +- **Topbar and Responsive Dashboard** - Persistent topbar with page titles, theme toggle, external links, user menu and version; responsive layouts across all views with truncation and tooltips for long values and icon-only buttons; version-skew detection with an auto-reloading upgrading screen (#6196, #6105, #6126, #6204, #6232, #6330, #6370, #6476, #6485, #6493) +- **Video Edits** - `POST /v1/videos/edits` applies prompt-driven edits, upscaling and background removal to an existing video supplied as bytes, a URL or a provider video ID, on OpenAI and Runware (#6270) +- **Runware Chat, Catalog and Media Operations** - Chat completions, streaming and Responses via Runware's OpenAI-compatible endpoint, `ListModels` from the curated catalog, image upscale via `/v1/images/edits` (`type=upscale`), image-to-3D and async 3D generation via `/v1/videos` (`type=3d`), provider-reported per-task cost, and a raw `/runware_passthrough` route (#6260, #6372, #6208, #6075) +- **JSON Image Edits** - `POST /v1/images/edits` accepts JSON bodies with URL or base64 images and typed extra params in addition to multipart (#6418) +- **OpenAI Ultrafast Service Tier** - `service_tier: "ultrafast"` is forwarded only to models that support it and billed at dedicated ultrafast rates, with matching custom pricing override fields (#6396, #6399) +- **Service Tier on Logs** - Logs record the tier actually served, including Anthropic's `service_tier` from `message_start` on streams, with a Service Tier column and detail field so repricing uses the served tier (#6233, #6236) +- **Pricing Fields** - New per-request flat fee (`cost_per_request`), megapixel-based image tiers (4/8/16/32/64 MP), per-size and joint size+quality image rates for `gpt-image-1`-style models, and `input_cost_per_query` for rerank flow through datasheet sync, the cost engine, custom overrides, the API and the UI override form; upscale output resolution is backfilled from `target`/`factor` on Replicate so tiered rates bill the real output size (#6079, #6082, #6083, #6379, #6380) +- **Model Catalog Pricing and Overrides** - Pricing data in the model catalog (thanks [@johnbrett](https://github.com/johnbrett)!), with resolved pricing overrides exposed on `/api/models/details` and on catalog rows, shown in the dashboard (#6055, #6056, #6058) +- **Typed Embeddings on Bedrock** - Titan V2 `embeddingTypes` and Cohere `embedding_types` on Converse, the native invoke route and LangChain `BedrockEmbeddings` (#6381) +- **Rerank Upgrades** - Structured JSON documents, `return_documents`, `next_token` pagination, caller document IDs preserved in every result, Cohere-shaped errors, cross-provider responses converted back to the caller's wire shape, and `/genai/v1/rank` served cross-provider (#6328, #6301, #6432) +- **OpenRouter Speech, Transcription and Embeddings** - TTS and STT through OpenRouter's audio endpoints, and embedding models included in `ListModels` (#5734, #6264) +- **Grok on Bedrock Mantle** - `xai.` models route through the `openai/v1` Mantle path (#6022) +- **Gemini 3 Thinking Levels** - A per-model `thinkingLevel` support table clamps requested levels to the rungs each model implements; `reasoning_effort: "none"` sets the model's floor level instead of zeroing `thinkingBudget` (#6280) +- **Datasheet-Backed Compatibility** - Anthropic, Bedrock, Cohere and Gemini request shaping (adaptive thinking, native effort, disable-reasoning, mid-conversation system turns, computer-use and text-editor tool generations, default max output tokens, tool validation) is resolved from model capabilities instead of hardcoded model-name checks (#6281, #6492) +- **Reasoning Effort None** - Models that reason by default but do not support reasoning with tool calls get `reasoning.effort: "none"` when they advertise `supports_none_reasoning_effort`, instead of losing `reasoning` entirely (#6293) +- **HTTP Transport Pre-Auth Hook** - New `HTTPTransportPreAuthHook` plugin phase runs before transport authentication so plugins can inject credentials such as `x-bf-vk`; a `virtual-key-from-config` native plugin example ships alongside it (#6375, #6373) +- **Plugin Inject Limits** - Per-plugin `semaphore_size` and `inject_timeout` on `PluginConfig` bound observability `Inject` calls so a hung connector releases its slot (#6341) +- **Harness Session Autodetection** - Claude Code, Codex CLI and OpenCode session headers populate the session ID when `x-bf-session-id` is absent (#6333) +- **Auth and Model Check Skip Paths** - Context keys let trusted internal callers bypass auth resolution, and let evaluate-only requests such as `/inspect` bypass the virtual key provider and model allowlists while budgets and rate limits still apply (#6124, #6479) +- **Passthrough Encoding Negotiation** - Forwarded `Accept-Encoding` is filtered to decodable codecs (gzip, deflate, brotli, zstd; gzip and identity for streams) and chained content encodings are decoded (#6360) +- **Routing Plugin** - Routing rules and the complexity router live in a dedicated `routing` plugin that runs after governance so rules evaluate on the fully stamped context; endpoints moved to `/api/routing/rules` and `/api/routing/complexity-analyzer-config` with deprecated `/api/governance/*` aliases; complexity routing now reads the text of mixed text+image turns (#6144, #6145, #6146, #6147, #6253) +- **Dimension Scope Ceiling** - Grouped log analytics (rankings, histograms, key pairs) are bounded to the customer, team, business unit, user and virtual key ids the caller may see (#6262) +- **MCP Per-User OAuth and Token Exchange** - MCP clients can hold per-user OAuth credentials and per-user headers, configurable from `config.json` as well as the UI, with a documented shared vs per-identity token lookup contract, `oauth_config.resource` (RFC 8707), VK/Users filters on the OAuth Grants and MCP Auth Sessions sidebars and one shared create/install client form; `token_exchange` gains `use_idp_credentials` to reuse SSO login app credentials for providers such as Microsoft Entra ID (`client_id` becomes optional) and combines `offline_access` with `/.default` for Entra OBO; shared-OAuth clients show `needs_reauth` when their token row is invalidated, `Reauthorize` is limited to shared clients, the OAuth flow claim is atomic against concurrent reauth, stored scopes survive a decode failure, and credential caches propagate cancellation and version their entries (#6068, #6069, #6078, #6411, #6428, #6429, #6504) +- **MCP Connection Lifecycle and Tool Discovery** - Discovered tools persist and resync uniformly across all client types through a hash-gated core callback, surviving restarts and propagating across a cluster; connections use make-before-break reconnects with ephemeral clients rebuilt across the whole connect+init retry, last-known tool maps preserved, connect attempts bound to entry identity and background reconnects deduped; `needs_session_stickiness` is pinned across `config.json` reconciliation; updating static headers on a sticky client pre-flight verifies the new credential and swaps it onto the live connection, per-call shared-credential clients refresh tools synchronously, and a failed enable parks the client at `Disabled` so it can be retried; the global `tool_sync_interval` hot-reloads and re-times running checkers; state badges render with spaces and the `disconnected` filter bucket is now `unstable` (#6409, #6430, #6431, #6483, #6502) +- **Air-Gapped MCP Catalog** - `mcp_library_sync_interval: 0` disables catalog sync and `file://` URLs load the MCP server library from disk (#6195) +- **MCP Log Redaction and Plugin Logs** - MCP tool logs carry redaction mappings and plugin logs (#5744, #5746) +- **Splunk Connector Configuration** - `config.schema.json`, Helm values and dashboard entries for the Splunk HEC observability connector (#6296, #6091, #6099) +- **Helm Broker Clustering** - `bifrost.cluster.type: broker` with broker address, port and TLS settings alongside the existing mesh transport (#6398) +- **HTTP/2 Ping Interval in the UI** - Provider network configuration exposes `http2_ping_interval_in_seconds` (#6228) +- **Status Code Badges** - Error and passthrough logs show the upstream HTTP status code in the log detail header (#5536) +- **Server-Side Tool Calls in Logs** - `web_search_call`, `code_interpreter_call` and similar Responses items render their full payload in the log detail view (#6475) +- **Gemini Server-Side Tool Calls** - Gemini `toolCall`/`toolResponse` parts surface as `web_search_call` items with their own call ID and queries, unmapped tool types are preserved on the native round-trip, and each `thoughtSignature` appears exactly once on replay (#6071) - **Bedrock VPC Endpoints** - AWS Bedrock keys can target VPC endpoints (#6064) -- **Per-Request Flat-Fee Pricing** - New `cost_per_request` field flows through datasheet sync, the cost engine, custom overrides and the UI override form (#6079) -- **Pricing Overrides in the Model Catalog** - `/api/models/details` exposes resolved pricing overrides, and catalog rows resolve overrides server-side (#6055, #6056) -- **MCP Tool Discovery Persistence** - Discovered MCP tools persist and resync uniformly across all client types through a hash-gated core callback, surviving restarts and propagating across a cluster - **W3C Trace ID Propagation** - Requests carry a W3C trace ID on the context (#5945) -- **Cancellable Log Cost Recalculation** - Log cost recalculation tasks can be cancelled from the backend (#5801) +- **Durable Background Jobs** - New `sidekiq` background-job table, store methods, and runner with recovery and reaper; cost recalculation migrated to a durable, resumable and cancellable job with polling instead of SSE (#5800, #5801) - **Separate OTEL Metrics Pipeline** - The OTEL collector supports a metrics tab independent of traces, plus separate headers for traces and metrics (#5939, #5940) -- **Roots-Only Log Filter** - New `roots_only` filter collapses fallback chains into their root entry with child aggregates (#5737) -- **MCP Log Redaction and Plugin Logs** - MCP tool logs carry redaction mappings and plugin logs (#5744, #5746) -- **User Agent and App Attribution in Logs** - Logs and MCP tool logs record user agent, app, source, decision, app key and device ID +- **Grouped Logs View** - The logs table groups fallback chains under expandable roots backed by the new `roots_only` filter with child aggregates, and the model catalog persists tab, search and provider in the URL (#5522, #5737, #6059) +- **User Agent and App Attribution** - Logs and MCP tool logs record user agent, app, source, decision, app key and device ID, with custom user-agent mapping and dashboard dimension rankings; MCP tool logs observed by the Bifrost Edge agent can be ingested with device, app key, decision and source attribution - **S3 Log Export Metadata** - Additional metadata is written alongside S3 log exports (#6070) - **Matview Maintenance Off Switch** - `matview_refresh_interval` accepts `"off"` to disable logstore matview maintenance entirely (thanks [@jeremym-tanium](https://github.com/jeremym-tanium)!) (#5693) - **Video Request Info in Logs UI** - Video requests surface their details in the logs UI (#5946) - **Shell Rewriter Hook** - The UI handler exposes a `ShellRewriter` hook for pre-hydration HTML rewriting (#5807) -- **Auth Skip Path** - Adds a context path letting trusted internal callers bypass auth resolution -- - **Runware passthrough** - Adds `runware_passthrough` path for handling passthrough mode for Runware provider +- **Custom Branding** - Logo and icon branding support with an OSS fallback stub, cached in localStorage to prevent a logo flash on load (#5806, #6096) +- **User Assignment on Virtual Keys** - Users can be assigned from the virtual key sheet (#5863) +- **Quarterly Budgets** - Quarterly budget windows with a configurable fiscal year start for customers and virtual key provider configs, surfaced in budget labels (#5996, #5997, #5999, #6115, #6116) +- **Sarvam AI Provider** - Added Sarvam AI as a first-class provider with chat, text-to-speech, and speech-to-text support (thanks [@Purvi09](https://github.com/Purvi09)!) +- **ElevenLabs Sound Effects** - Added text-to-sound generation support via `/v1/sound-generation` (thanks [@SecretSun](https://github.com/SecretSun)!) +- **Bedrock Project Scoping** - Added optional `project_id` to Bedrock and Bedrock Mantle key configs with per-alias overrides for Bedrock, Bedrock Mantle, and Vertex, plus UI support +- **Trace Redaction** - Phase-scoped redaction and revealing, transient redaction data field for guardrails, and trace content redaction before connector export +- **Audit Log Object Storage** - S3/GCS object storage config schema for audit log archival +- **Alerting Configuration** - Alerting schema in `config.schema.json` with declarative channels and CEL-based rules, Helm chart support, and enterprise fallback pages +- **Canonical Model Names** - Dashboard model rankings now show canonical model names instead of inference-profile IDs (thanks [@satyamkrishna](https://github.com/satyamkrishna)!) +- **OAuth2 Hardening** - Allowlist for private-use redirect URI schemes (RFC 8252 ยง7.1) and a `shouldSweep` gate on the OAuth2 sweep worker +- **Mirrored Schema Support** - `schema_url` / `BIFROST_SCHEMA_URL` for mirrored schema locations in isolated deployments +- **Vertex Single-Region Config** - Enforce single-region configuration in Vertex key config +- **Helm Chart Updates** - `bifrost.alerting`, audit-log object storage, `postgresql.external.port` string support, and `bifrost.mcp.toolGroups[*].id` +- **ChatGPT Passthrough** - Added a ChatGPT passthrough route on the OpenAI integration with dedicated request handling +- **Edge Fallback Pages** - Added fallback pages for Bifrost Edge control views (config, devices, inventory) backed by governance resolver support +- **Agent Handover View** - Added an agent handover page with seeded end-to-end data support +- **First-Time Setup Token** - A setup token gates first-time setup so a fresh deployment is not open to the world, and the onboarding checklist is back, completing its dashboard auth step on SSO deployments (#5759, #5784, #6322) ## ๐Ÿž Fixed -- **GenAI SSE Heartbeats** - GenAI streams delimit heartbeat comments so Google SDK clients preserve the following event (thanks [@dani29](https://github.com/dani29)!) (#6240) +- **Structured Output Schema Order** - `response_format` JSON schemas are forwarded byte-for-byte to OpenAI, Anthropic, Bedrock, Gemini and Cohere so the model generates fields in the caller's declared order instead of a re-sorted one (#6235) +- **Thinking Block Typing on Streams** - Reasoning items carrying both an encrypted payload and a visible summary open as `thinking` blocks instead of `redacted_thinking` (#6292) +- **Replayed Thinking Blocks via `bedrock/` Prefix** - Content-less `tool_result` blocks are kept, interleaved block order is preserved, `incomplete` maps to `error` on Converse, and pending reasoning is consumed by its owning item, so multi-turn tool use no longer wedges (#6346) +- **Gemini 400s on Claude Code Traffic** - Trailing assistant prefills are trimmed and mid-conversation system turns are inlined for Gemini/Vertex; `extra_fields` is echoed on `/anthropic/v1/messages` (#6363) +- **Bedrock Tool Use IDs** - IDs longer than 64 characters or outside Bedrock's charset (such as Gemini thought-signature IDs) are aliased deterministically on both `tool_use` and `tool_result` (#6300) +- **Azure Responses Stream Errors** - Terminal `error` and `response.failed` events inside an already-open HTTP 200 SSE stream are surfaced as errors with their nested type, code and message (thanks [@dani29](https://github.com/dani29)!) (#6302) +- **GenAI SSE Heartbeats** - GenAI streams delimit heartbeat comments so Google SDK clients preserve the following event, while older openai-go clients keep the bare heartbeat (thanks [@dani29](https://github.com/dani29)!) (#6252) +- **OpenCode max_tokens** - `max_tokens` is preserved for OpenCode-compatible chat endpoints (thanks [@Alex-wangyang](https://github.com/Alex-wangyang)!) (#6458) +- **HuggingFace Streaming Usage** - HuggingFace is no longer listed as omitting the `[DONE]` marker, and `stream_options.include_usage` defaults on its chat streaming path, so streamed calls stop reporting zero tokens and zero cost (thanks [@elliottrabac](https://github.com/elliottrabac)!) (#6478) +- **Provider Key Name on Update** - A key PUT that omits `name` no longer clears it, and already-exists errors keep their constraint detail (thanks [@cpsc](https://github.com/cpsc)!) (#6417) +- **Bedrock Mantle Streaming** - Bedrock Mantle is registered in `ProviderSendsDoneMarker` so streams end after `finish_reason` (#6021) +- **URL-Sourced Files and Images** - `gs://` URIs go to Gemini/Gemma as `fileData.fileUri` and are read from Cloud Storage for Claude-on-Vertex, `s3://` references go to Bedrock Converse as `s3Location`, Bedrock rerank synthesizes the foundation-model ARN from a bare model ID, OpenAI file blocks keep `file_url`, non-http schemes pass through on the OpenAI and native-Anthropic paths, and Gemini always emits a candidate with its finish reason and drops payload-free parts (#6239) +- **Together and Alias Pricing** - The management catalog resolves runtime provider `together` to the datasheet identity and prices configured aliases through their target model (thanks [@dani29](https://github.com/dani29)!) (#6257, #6320) +- **Redis Vector Store TAG Escaping** - All RediSearch special characters are escaped in TAG query values (thanks [@AdityaPainuli](https://github.com/AdityaPainuli)!) (#5351) +- **MCP Tool Sync Interval Corruption** - Toggling an MCP client's enable/disable switch no longer corrupts `tool_sync_interval`; the value is a whole number of minutes, negative values are rejected instead of silently disabling sync, and re-enabling a per-call client restarts its discovery cycle (#6409, #6502) +- **MCP Tool Map Staleness** - `SetClientTools` replaces the in-memory tool map instead of merging, so tools removed upstream leave memory once the database has dropped them (#6484) +- **SSE Reconnect Identity** - `OnConnectionLost` on SSE MCP clients is gated on connection identity so a stale connection cannot tear down its replacement +- **Connector Header Redaction** - `Authorization`, `x-api-key`, Cloudflare Access and AWS ALB OIDC headers are redacted before export to every observability backend (#6371) +- **Vertex Mixed Tools** - Vertex AI accepts function declarations and Google Search in the same request without `includeServerSideToolInvocations`, and search localization via `retrievalConfig.latLng` is preserved (#6066) +- **Gemini Tool Preference** - When tool combination is disabled, function declarations win over Google Search so the model can still call the caller's tools (#6065) +- **Bedrock Stop Reasons** - Bedrock `content_filter` and `guardrail_intervened` stop reasons map to `incomplete` status with a `content_filter` reason +- **Encrypted Reasoning on Compaction** - The fail-soft that strips `encrypted_content` before retrying a rejected request also covers `/v1/responses/compact` and count-tokens requests, and recognizes Anthropic's `redacted_thinking` rejection (#6041, #5960) +- **DAC-Scoped VK Reads** - `from_memory` virtual key reads are blocked for DAC-scoped callers - **Path Normalization Auth Bypass** - Fixed a path normalization flaw that allowed auth to be bypassed (#5763) - **Minimal Reasoning Effort on GPT-5 Models** - `reasoning_effort: "minimal"` is preserved for GPT-5-family OpenAI models instead of being downgraded to `low` (thanks [@jitokim](https://github.com/jitokim)!) (#6046) - **Gemini Truncated Response Finish Reason** - Truncated Gemini responses report `MAX_TOKENS` instead of `OTHER` (thanks [@AdityaPainuli](https://github.com/AdityaPainuli)!) (#5979) - **Null Tool-Call Function Name on Streaming** - Streaming continuation deltas no longer materialize an absent tool-call function name as `null` (thanks [@AdityaPainuli](https://github.com/AdityaPainuli)!) (#5966) - **Bedrock Document Uploads** - Fixed Bedrock file handling in inference so office and PDF documents sent as OpenAI `type: "file"` are accepted (#5947) - **xAI Usage Cost** - Fixed USD cost ticks for xAI usage (#5950) -- **Anthropic Encrypted Reasoning** - Added an Anthropic error branch when stripping encrypted reasoning content -- **MCP Reconnect and Lock Ordering** - Broke a lock-order inversion in `ConnectionCheckerManager`, rebuilt ephemeral clients across the whole connect+init retry, preserved last-known tool maps across close-first reconnects, bound connect attempts to entry identity, deduped background reconnects and gated SSE `OnConnectionLost` on connection identity -- **MCP OAuth Session Correctness** - Restricted `Reauthorize` to shared OAuth clients, rejected inactive tokens in `ValidateToken`, made the OAuth flow claim atomic against concurrent reauth, stopped dropping stored scopes on decode failure, and closed a verify-headers double-submit race that also dropped TLS, timeout and per-user-header fields -- **Session Stickiness Reconciliation** - `needs_session_stickiness` is pinned across `config.json` reconciliation, so an unrelated file edit can no longer silently revert a client to per-call -- **Credential Cache Cancellation** - `headerCredentialCache.Fill` and `userTokenCache.Fill` propagate context so a cancelled request unblocks instead of waiting on an unrelated leader; LRU entries carry a version so a rejected stale `Get` cannot evict a concurrently-updated value - **Governance List-Models Call** - Budgets and rate limits no longer trigger a list-models call (#6051) - **Realtime Response Create Input** - Guarded `response.create` input (#6050) -- **HTTP Server Timeouts** - Configured bounded `http.Server` timeouts and a request-body limit -- **MCP Client State Badges** - State badges render with spaces instead of underscores, and the state filter bucket was renamed from `disconnected` to `unstable` -- **Entra OBO Scope** - `offline_access` is combined with `/.default` for Entra OBO instead of replacing it (#6078) +- **Governance Rate-Limit Reset CPU** - Guards against invalid reset timeouts, parallelized resting-budget flows only when absolutely required, and fixed the calendar-based alignment qualifier +- **Masked Key Persistence** - Never persist masked provider key previews to config storage (thanks [@eyeveil](https://github.com/eyeveil)!) +- **OpenShift Arbitrary UIDs** - Build-time group-0 ownership with no runtime chown (thanks [@eyeveil](https://github.com/eyeveil)!) +- **Passthrough Virtual Key Attribution** - Passthrough calls via the Azure `api-key` header now attribute to the virtual key (thanks [@eyeveil](https://github.com/eyeveil)!) +- **Rerank for Custom Providers** - `/v1/rerank` now works with custom OpenAI-compatible providers (thanks [@eyeveil](https://github.com/eyeveil)!) +- **Responses Stream Usage** - Persist stream usage when providers omit or reuse sequence numbers (thanks [@eyeveil](https://github.com/eyeveil)!) +- **Wildcard allowed_models Repair** - Repair bare wildcard `allowed_models` rows that broke admin provider updates (thanks [@eyeveil](https://github.com/eyeveil)!) +- **Streaming Error Panic** - Nil-safe tracing span lookup prevents panics on streaming errors (thanks [@eyeveil](https://github.com/eyeveil)!) +- **Anthropic Tool ID Sanitization** - Sanitize `tool_use`/`tool_result` ids to Anthropic's charset (thanks [@Shaik-Sirajuddin](https://github.com/Shaik-Sirajuddin)!) +- **Realtime Transcription Sessions** - Support GA transcription-type sessions in `POST /v1/realtime/client_secrets` (thanks [@Shaik-Sirajuddin](https://github.com/Shaik-Sirajuddin)!) +- **Diarized Transcription** - Support `diarized_json` segments and ElevenLabs speaker passthrough (thanks [@Shaik-Sirajuddin](https://github.com/Shaik-Sirajuddin)!) +- **Model Discovery** - Skip disabled keys when scheduling model-discovery fetches (thanks [@Shaik-Sirajuddin](https://github.com/Shaik-Sirajuddin)!) +- **MCP Timeout Placeholder** - Show the real global default in the MCP tool execution timeout placeholder (thanks [@Shaik-Sirajuddin](https://github.com/Shaik-Sirajuddin)!) +- **Redacted Thinking Round-Trip** - Round-trip Anthropic `redacted_thinking` blocks on the Responses surface (thanks [@fus3r](https://github.com/fus3r)!) +- **Streaming Accumulation** - Preserve citation annotations and `finish_reason` in the accumulated streaming response (thanks [@fus3r](https://github.com/fus3r)!) +- **Gemini Grounded Streaming** - Reset web-search flag when recycling pooled stream state so `web_search_call` items keep emitting (thanks [@fus3r](https://github.com/fus3r)!) +- **Bedrock Truncation Signal** - Signal `max_output_tokens` truncation on the Responses API (thanks [@jeremym-tanium](https://github.com/jeremym-tanium)!) +- **Bedrock Reasoning Config** - Preserve `reasoning_config` on cross-provider translation so fallbacks keep extended thinking (thanks [@Purvi09](https://github.com/Purvi09)!) +- **Anthropic tool_search** - Forward and rebuild server-side `tool_search` on the Responses path (thanks [@ws4charlie](https://github.com/ws4charlie)!) +- **OpenAI Responses Input** - Strip `role` from non-message input items (thanks [@nettee](https://github.com/nettee)!) and serialize compaction request `input` correctly (thanks [@mcclurmc](https://github.com/mcclurmc)!) +- **additional_tools Support** - Added `additional_tools` message type support, preserving nested tool types on `/v1/responses` +- **Plugin Stream Errors** - Emit structured plugin stream errors on integration routes (thanks [@jeffhos](https://github.com/jeffhos)!) +- **Pooled Object Hygiene** - Zero pooled ChannelMessage references on release and sweep orphaned deferred spans in trace store TTL cleanup (thanks [@citrocat](https://github.com/citrocat)!) +- **Hybrid Log Token Usage** - Rebuild token usage from denormalized columns in hybrid log list (thanks [@G-XD](https://github.com/G-XD)!) +- **MCP Tool Ordering** - Deterministic MCP tool ordering for prompt cache stability +- **MCP Inline-Auth Links** - Warn callers not to truncate the `#t=` temp-token fragment (thanks [@MarcusPeng](https://github.com/MarcusPeng)!) +- **Gemini Fixes** - Web search options map to Google Search grounding, file upload MIME types preserved, and video reference fields map to instances (thanks [@vojthor](https://github.com/vojthor)!) +- **OpenAI Parameters** - Honor service tier in chat completion and cap max reasoning effort +- **Anthropic Costing** - Correct inference geo cost and cache rate for fast mode +- **SecretVar Parsing** - Parse `SecretVar` JSON with `ref`/`env_var` fields even when `value` is absent +- **Telemetry** - Forward request id and trace id, reduce metrics cardinality explosion risk, and send status codes on OTEL metrics +- **Dashboard** - Preserve active time period when applying dimension filters, adjust bucket size thresholds for month-range durations, show user popover with `preferred_username` fallback, filter provider-level keys from the prompt manager selector (thanks [@rlex](https://github.com/rlex)!), skip password validation for redacted credentials, and improve `ModelMultiselect` empty and error states +- **API Key Provider Selection** - Fixed provider selection for API keys +- **Azure Auth Headers** - Pass Azure auth headers in helpers +- **Stream Delta Schema** - Added `ExtraContent` to `ChatStreamResponseChoiceDelta` (thanks [@nghodkicisco](https://github.com/nghodkicisco)!) +- **API Auth Bypass** - Stopped `/api/devices` bypassing auth via the `/api/dev` prefix +- **Bedrock Error Types** - Surface the AWS exception type (`X-Amzn-Errortype`) on non-streaming Bedrock error responses instead of dropping it ## ๐Ÿ”ง Maintenance +- **Hot-Path Performance** - Cached serialization for shared MCP tools, a direct `OrderedMap` JSON writer, bulk span attribute writes with cached span pointers, reusable worker delivery timers, retained span attribute maps, generation-stamped memoization of `GetProvidersForModel` and `GetModelsForProvider` via the new `gencache` package, sonic-based JSON responses, and a plugin-log existence check before draining (#6242, #6241, #5956, #5957, #5657, #6387, #5641, #6224, #6268, #6211) +- **Go Toolchain** - Modules build with Go 1.26.6 and the Nix flake pins 1.26.7 (#6269, #6385) +- **Dependency Upgrades** - Dependabot updates across all modules, newman 6.2.2 with pinned transitive overrides, module path fixes and `openai_config` referenced from every provider config schema (#6040, #5864, #6267, #6305, #6275) +- **Test Coverage** - vLLM instances provisioned on RunPod in the release pipeline, Runware harness coverage including `/v1/images/edits` and `/v1/videos`, batch and pricing-override lifecycle harness cases, an Anthropic `message_start` usage regression test, LangChain rerank and embedding integration tests, and e2e fixes for dashboard auth, budget reset and MCP state (#5541, #6303, #6319, #6299, #6327, #6432, #6351) +- **Documentation** - v2.0.0 migration guide with the governance namespace mapping and a v1.5.x downgrade guide for `prerelease3` deployments, v2.0.0 availability callouts, routing API namespace docs, Bedrock application inference profiles, Splunk connector docs, config.schema.json and Datadog env var reference fixes, and Discord badge fixes (thanks [@Swpn0neel](https://github.com/Swpn0neel)!) (#6332, #6374, #6420, #6147, #6203, #6099, #5938, #6019, #6425, #6448) +- **Helm** - Chart releases v2.1.35 and v2.1.36 (#6129, #6249) - **Governance Route Families** - Editions can override governance route families (#5839) -- **Dependency Upgrades** - Dependabot updates across all modules, plus module path fixes (#6040, #5864) -- **Documentation** - config.schema.json doc fixes and Datadog env var reference fixes in the helm chart docs (#5938, #6019) ## ๐Ÿ—„๏ธ Database Migrations +All migrations below are new relative to v1.6.11. Deployments on an older v1.6.x release should also review the intermediate v1.6.x changelogs. + **configstore:** - **add_mcp_client_pending_oauth_config_json_column** - Adds `pending_oauth_config_json` to `config_mcp_clients`. Reversible: drops the added column. @@ -59,6 +178,13 @@ - **add_needs_session_stickiness_column** - Adds `needs_session_stickiness` to `config_mcp_clients`. Reversible: drops the added column. - **add_bedrock_endpoints_columns** - Adds Bedrock VPC endpoint columns to the keys table. Reversible: drops the added columns. - **add_cost_per_request_pricing_column** - Adds `cost_per_request` to model pricing. Reversible: drops the added column. +- **add_notifications_table** - Creates the `notifications` table for the dashboard notification center. Reversible: drops the table. +- **add_batch_jobs_table** - Creates `batch_jobs` with a unique `(provider, batch_id)` identity index, a sweeper scan index and a runner-id index. Reversible: drops the table. +- **add_image_megapixel_tier_pricing_columns** - Adds the five `output_cost_per_image_above_{4,8,16,32,64}_megapixels` columns to model pricing. Reversible: drops the added columns. +- **add_input_cost_per_query_column** - Adds `input_cost_per_query` to model pricing for rerank. Reversible: drops the added column. +- **add_ultrafast_pricing_columns** - Adds the four `*_ultrafast` token rate columns to model pricing. Reversible: drops the added columns. +- **add_image_size_quality_pricing_columns** - Adds the 14 per-size and size+quality image output rate columns to model pricing. Reversible: drops the added columns. +- **add_batch_jobs_attribution_columns** - Adds `user_id`, `team_id`, `customer_id` and `source_log_id` to `batch_jobs` plus a `user_id` index. Reversible: drops the index and the four columns. **logstore:** @@ -66,9 +192,15 @@ - **mcp_tool_logs_add_redaction_mapping_column** - Adds the redaction mapping column to MCP tool logs. **Non-reversible**: rollback is a no-op because dropping the column would permanently destroy reveal data for already-redacted MCP logs. - **logs_add_user_agent_column** - Adds user agent and app columns, their indexes, and a `UserAgentMapping` table. Reversible: drops the indexes and the mapping table. - **mcp_tool_logs_add_user_agent_column** - Adds user agent and app columns plus indexes to MCP tool logs. Reversible: drops both indexes and the `app` column. +- **logs_recreate_matviews_with_app_column** - Recreates the log materialized views to include the user agent and app columns. Rollback is a no-op because `ensureMatViews` recreates them on next startup. - **mcp_tool_logs_add_endpoint_columns** - Adds `source`, `decision`, `app_key` and `device_id` to MCP tool logs. Reversible: drops all four columns. - **mcp_tool_logs_add_plugin_logs_column** - Adds `plugin_logs` to MCP tool logs. Reversible: drops the added column. -- **logs_recreate_matviews_with_user_agent_column** and **logs_recreate_matviews_with_app_column** - Recreate the log materialized views to include the new columns. Rollback is a no-op because `ensureMatViews` recreates them on next startup. +- **logs_add_video_edit_input_column** - Adds `video_edit_input` to logs. Reversible: drops the added column. +- **logs_add_upstream_and_overhead_latency_columns** - Adds `upstream_latency` and `overhead_latency` to logs. Reversible: drops both columns. +- **logs_add_batch_debug_column** - Adds `batch_debug` to logs. Reversible: drops the added column. +- **logs_add_cost_breakdown_columns** - Adds `input_cost`, `output_cost` and `additional_cost` to logs. Reversible: drops the three columns. +- **logs_recreate_matviews_with_cost_breakdown** - Marks the hourly matview for rebuild with the cost split columns; `repairMatViewShapes` drops and recreates `mv_logs_hourly` on the next startup. Rollback is a no-op because `ensureMatViews` recreates it on next startup. +- **logs_add_overhead_breakdown_column** - Adds `overhead_breakdown` to logs. Reversible: drops the added column. **High-throughput deployments: run the logstore migrations during a low-activity window.** @@ -83,7 +215,48 @@ Every logstore migration above alters `logs` or `mcp_tool_logs`, the two highest ## ๐Ÿ™ Closed GitHub Issues - [#123](https://github.com/maximhq/bifrost/issues/123) - Files API Support +- [#2347](https://github.com/maximhq/bifrost/issues/2347) - MCP tool ordering is non-deterministic, breaking prefix-based prompt caching +- [#3455](https://github.com/maximhq/bifrost/issues/3455) - Segfault/nil dereference panic in Bedrock provider +- [#4318](https://github.com/maximhq/bifrost/issues/4318) - allowed_models persisted as bare "*" string blocks subsequent provider updates +- [#4353](https://github.com/maximhq/bifrost/issues/4353) - config.db corruption from masked-key preview in provider_configs JSON column +- [#4367](https://github.com/maximhq/bifrost/issues/4367) - Image incompatible with OpenShift arbitrary UIDs +- [#4402](https://github.com/maximhq/bifrost/issues/4402) - Vertex provider drops image blocks whose URL uses gs:// scheme +- [#4477](https://github.com/maximhq/bifrost/issues/4477) - Passthrough calls using a Virtual Key log as actual key +- [#4679](https://github.com/maximhq/bifrost/issues/4679) - Bedrock Responses API does not signal max_output_tokens truncation +- [#4689](https://github.com/maximhq/bifrost/issues/4689) - Custom providers cannot set budget +- [#4712](https://github.com/maximhq/bifrost/issues/4712) - ElevenLabs sound effects (/v1/sound-generation) +- [#4780](https://github.com/maximhq/bifrost/issues/4780) - Anthropic server-side tool_search results are dropped on /v1/responses +- [#4834](https://github.com/maximhq/bifrost/issues/4834) - /v1/rerank is not available with custom providers +- [#4846](https://github.com/maximhq/bifrost/issues/4846) - Responses stream usage present in response.completed but not persisted in LLM Logs +- [#4851](https://github.com/maximhq/bifrost/issues/4851) - Governance rate-limit reset causes high CPU in BumpRateLimitUsage +- [#4870](https://github.com/maximhq/bifrost/issues/4870) - Pooled ChannelMessage retains request body, context, and undelivered response while idle +- [#4940](https://github.com/maximhq/bifrost/issues/4940) - Show canonical model names instead of Bedrock inference-profile IDs in Model Rankings +- [#4963](https://github.com/maximhq/bifrost/issues/4963) - Streaming finish_reason dropped from the accumulated (logged) response +- [#5002](https://github.com/maximhq/bifrost/issues/5002) - gpt-4o-transcribe-diarize transcription fails due to string segment IDs +- [#5013](https://github.com/maximhq/bifrost/issues/5013) - OpenAI /responses/compact input serialized as a JSON object causing 400 +- [#5026](https://github.com/maximhq/bifrost/issues/5026) - [Bug]: Toggling an MCP client's enable/disable switch corrupts its tool_sync_interval (nanoseconds resent as minutes) +- [#5027](https://github.com/maximhq/bifrost/issues/5027) - MCP Tool Execution Timeout placeholder shows 0 instead of real global default +- [#5036](https://github.com/maximhq/bifrost/issues/5036) - Plugin StreamInterceptionError is flattened on integration routes +- [#5037](https://github.com/maximhq/bifrost/issues/5037) - Disabled keys break provider model discovery +- [#5051](https://github.com/maximhq/bifrost/issues/5051) - Add Sarvam AI provider (chat + TTS/STT) +- [#5061](https://github.com/maximhq/bifrost/issues/5061) - Streaming responses drop citation annotations from the accumulated message +- [#5093](https://github.com/maximhq/bifrost/issues/5093) - Streaming /v1/responses drops Anthropic redacted_thinking blocks +- [#5097](https://github.com/maximhq/bifrost/issues/5097) - Anthropic rejects replayed tool_use/tool_result ids from non-conforming upstream providers +- [#5100](https://github.com/maximhq/bifrost/issues/5100) - additional_tools loses nested tool types on /v1/responses +- [#5101](https://github.com/maximhq/bifrost/issues/5101) - Chat-to-Responses tool replay sends role on function_call input items +- [#5108](https://github.com/maximhq/bifrost/issues/5108) - Bedrock reasoning_config silently dropped on cross-provider translation +- [#5113](https://github.com/maximhq/bifrost/issues/5113) - Gemini/Vertex streaming stops emitting web_search_call items after first grounded request +- [#5432](https://github.com/maximhq/bifrost/issues/5432) - Add TTS and STT support for OpenRouter - [#5472](https://github.com/maximhq/bifrost/issues/5472) - [Bug]: Bedrock rejects office/PDF document uploads via OpenAI `type:"file"` - "The PDF specified was not valid" +- [#5871](https://github.com/maximhq/bifrost/issues/5871) - [Bug]: AWS Bedrock Mantle streaming is broken +- [#5874](https://github.com/maximhq/bifrost/issues/5874) - [Bug]: SSE heartbeat frame aborts streams for openai-go ssestream consumers (< v3.43.0) with "unexpected end of JSON input" +- [#5885](https://github.com/maximhq/bifrost/issues/5885) - [Bug]: v1.6.8 omits message_start.message.usage on Bedrock-backed providers, breaking @ai-sdk/anthropic streaming - [#5900](https://github.com/maximhq/bifrost/issues/5900) - [Bug]: Streaming continuation chunks materialize omitted tool-call metadata as null - [#5978](https://github.com/maximhq/bifrost/issues/5978) - [Bug]: Gemini egress reports truncated responses as FinishReason OTHER, IncompleteDetails switch matches a string that never occurs - [#6044](https://github.com/maximhq/bifrost/issues/6044) - [Bug]: normalizeOpenAIReasoningEffort maps 'minimal' to 'low' for ALL OpenAI models, even ones that natively support 'minimal' +- [#6240](https://github.com/maximhq/bifrost/issues/6240) - [Bug]: GenAI SSE heartbeat framing causes @google/genai to silently drop the following data event +- [#6248](https://github.com/maximhq/bifrost/issues/6248) - [Bug]: OpenRouter embedding models missing from Semantic Cache dropdown +- [#6334](https://github.com/maximhq/bifrost/issues/6334) - [Bug]: Gemini/Vertex provider fails on Claude Code assistant prefills and mid-conversation system turns (Gemini 3.6 Flash & 3.7 Flash HTTP 400) +- [#6342](https://github.com/maximhq/bifrost/issues/6342) - [Bug]: Anthropic ingress with bedrock/ prefix restructures replayed thinking blocks, wedging multi-turn tool use on claude-opus-4-8 +- [#6416](https://github.com/maximhq/bifrost/issues/6416) - [Bug]: Provider key update silently clears "name" when omitted, then the unique-name index 409s subsequent updates +- [#6457](https://github.com/maximhq/bifrost/issues/6457) - [Bug]: OpenCode chat endpoints drop max completion limit diff --git a/transports/version b/transports/version index d29840733db..227cea21564 100644 --- a/transports/version +++ b/transports/version @@ -1 +1 @@ -2.0.0-prerelease3 +2.0.0