Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions Makefile
Original file line number Diff line number Diff line change
Expand Up @@ -468,6 +468,12 @@ _build-with-docker: # Internal target for Docker-based cross-compilation
exit 1; \
fi

# NOTE: transports/Dockerfile sets GOWORK=off and resolves the framework module
# from the published tag (go.mod), so plain `make docker-image` FAILS on branches
# whose local framework/ is ahead of the latest release (e.g. missing packages
# like framework/batchaccounting). Use `LOCAL=1 make docker-image` (Dockerfile.local,
# go-workspace build) on such branches — a stale-layer image from the plain target
# has already caused one silent-regression incident (2026-08-20).
docker-image: build-ui ## Build Docker image (LOCAL=1 to use Dockerfile.local)
@$(ECHO) "$(GREEN)Building Docker image...$(NC)"
$(eval GIT_SHA=$(shell git rev-parse --short HEAD))
Expand Down
9 changes: 9 additions & 0 deletions core/bifrost.go
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,7 @@ import (
"github.com/maximhq/bifrost/core/mcp"
"github.com/maximhq/bifrost/core/mcp/codemode/starlark"
"github.com/maximhq/bifrost/core/mcp/credstore"
"github.com/maximhq/bifrost/core/providers/alibaba"
"github.com/maximhq/bifrost/core/providers/anthropic"
"github.com/maximhq/bifrost/core/providers/azure"
"github.com/maximhq/bifrost/core/providers/bedrock"
Expand All @@ -33,6 +34,7 @@ import (
"github.com/maximhq/bifrost/core/providers/gemini"
"github.com/maximhq/bifrost/core/providers/groq"
"github.com/maximhq/bifrost/core/providers/huggingface"
"github.com/maximhq/bifrost/core/providers/kimi"
"github.com/maximhq/bifrost/core/providers/mistral"
"github.com/maximhq/bifrost/core/providers/nebius"
"github.com/maximhq/bifrost/core/providers/ollama"
Expand All @@ -51,6 +53,7 @@ import (
"github.com/maximhq/bifrost/core/providers/vllm"
"github.com/maximhq/bifrost/core/providers/wafer"
"github.com/maximhq/bifrost/core/providers/xai"
"github.com/maximhq/bifrost/core/providers/zhipu"
schemas "github.com/maximhq/bifrost/core/schemas"
"github.com/valyala/fasthttp"
)
Expand Down Expand Up @@ -4512,6 +4515,12 @@ func (bifrost *Bifrost) createBaseProvider(providerKey schemas.ModelProvider, co
return deepseek.NewDeepSeekProvider(config, bifrost.logger)
case schemas.Wafer:
return wafer.NewWaferProvider(config, bifrost.logger)
case schemas.Alibaba:
return alibaba.NewAlibabaProvider(config, bifrost.logger)
case schemas.Kimi:
return kimi.NewKimiProvider(config, bifrost.logger)
case schemas.Zhipu:
return zhipu.NewZhipuProvider(config, bifrost.logger)
case schemas.Gemini:
return gemini.NewGeminiProvider(config, bifrost.logger), nil
case schemas.OpenRouter:
Expand Down
2 changes: 2 additions & 0 deletions core/changelog.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,3 +16,5 @@
- fix: accept a bare model identifier on Bedrock rerank by synthesizing the foundation-model ARN from the resolved region - Rerank is the one Bedrock surface that names its model by ARN rather than by bare ID, so all three rerank drop-ins in the provider harness 400'd on `amazon.rerank-v1:0`. The partition is derived from the region (`aws`, `aws-cn`, `aws-us-gov`) so GovCloud and China build a correct ARN, and an explicit ARN still passes through untouched
- fix: stop stripping `file_url` from OpenAI-shaped chat file blocks on marshal - dropping it produced `{"type":"file","file":{}}` and an upstream complaint about a missing `file_id`, which hid the fact that a source had been discarded. Providers that cannot take a URL now say so by name, and any OpenAI-compatible endpoint that does accept one keeps working without a Bifrost change
- fix: leave URL content sources Bifrost cannot download in place on the OpenAI and native-Anthropic paths instead of failing the request - only `http(s)` is fetched, and whether a `gs://`, `s3://` or scheme-less reference is usable is the provider's call, so the source now travels as `{"type":"url"}` and the platform answers for itself
- fix: support GLM-5.2+ for max reasoning effort and clamp GLM-5.3+ values to max/high/low on every OpenAI-dialect mount (medium→high, none/minimal→low, xhigh→max), mirroring Z.ai's Coding Plan coercion — the GLM-5.3 API errors on every other value [@is911](https://github.com/is911)
- feat: add alibaba, kimi and zhipu providers — Alibaba Cloud Model Studio (Qwen/DashScope), Kimi (Moonshot AI) and Zhipu AI (GLM/Z.AI) — with default OpenAI-compatible mounts and optional Anthropic-compatible routing (per-key `use_anthropic_endpoints` through the shared Anthropic converters), per-vendor `reasoning_effort` shaping, and native Responses/embeddings where the vendor offers them. qwen3.8-max publishes `xhigh` as its top tier on the OpenAI-compatible mount (vendor enum: none/minimal/low/medium/high/xhigh), so the gateway forwards `xhigh` verbatim and clamps `max`→`xhigh` instead of forwarding a value the vendor 400'd until ~2026-08-22. On the Anthropic-compatible mounts the gateway now clamps `output_config.effort` per model family instead of the blanket `max`→`xhigh` for alibaba (the mount's own API page documents no effort field; live-verified it validates the value against its per-model chat-completions enum instead of mapping it server-side — per-model matrix from the Model Studio model-page docs supplied 2026-08-23: qwen3.8-max keeps the `max`→`xhigh` clamp; glm-5.3+ takes `max`/`high`/`low` with `xhigh`→`max`, `medium`→`high`, `minimal`/`none`→`low`; glm-5.2/5.1/5 and non-dated deepseek-v4-pro/flash take `high`/`max` with `xhigh`→`max`, `low`/`medium`/`minimal`/`none`→`high` (`low` is out of this family's enum, so the mildest tiers collapse onto the mildest valid value); the dated snapshots deepseek-v4-pro-0813/deepseek-v4-flash-0731 take `max`/`high`/`low` with `xhigh`/`medium`→`high`, `minimal`/`none`→`low`), sends the effort alone without a synthesized `thinking` field when one is set (Model Studio rejects `reasoning_effort` and `thinking_budget` together and engages thinking itself), and derives the mount base URL idempotently for all three vendors so a `base_url` already pointing at the Anthropic mount is used as-is instead of doubling the suffix — the path-rewriting suffix rules only apply on each vendor's own hosts (`*.aliyuncs.com`, `api.z.ai`/`open.bigmodel.cn`, and Kimi's kimi/moonshot hosts), so a custom or proxied base URL keeps its configured path and only gets the mount suffix appended [@is911](https://github.com/is911)
70 changes: 70 additions & 0 deletions core/internal/llmtests/account.go
Original file line number Diff line number Diff line change
Expand Up @@ -194,6 +194,9 @@ func (account *ComprehensiveTestAccount) GetConfiguredProviders() ([]schemas.Mod
schemas.Fireworks,
schemas.Sarvam,
schemas.Wafer,
schemas.Alibaba,
schemas.Kimi,
schemas.Zhipu,
ProviderOpenAICustom,
}, nil
}
Expand Down Expand Up @@ -485,6 +488,33 @@ func (account *ComprehensiveTestAccount) GetKeysForProvider(ctx context.Context,
UseForBatchAPI: bifrost.Ptr(true),
},
}, nil
case schemas.Alibaba:
return []schemas.Key{
{
Value: *schemas.NewSecretVar("env.ALIBABA_API_KEY"),
Models: []string{"*"},
Weight: 1.0,
UseForBatchAPI: bifrost.Ptr(true),
},
}, nil
case schemas.Kimi:
return []schemas.Key{
{
Value: *schemas.NewSecretVar("env.KIMI_API_KEY"),
Models: []string{"*"},
Weight: 1.0,
UseForBatchAPI: bifrost.Ptr(true),
},
}, nil
case schemas.Zhipu:
return []schemas.Key{
{
Value: *schemas.NewSecretVar("env.ZHIPU_API_KEY"),
Models: []string{"*"},
Weight: 1.0,
UseForBatchAPI: bifrost.Ptr(true),
},
}, nil
case schemas.Wafer:
return []schemas.Key{
{
Expand Down Expand Up @@ -872,6 +902,46 @@ func (account *ComprehensiveTestAccount) GetConfigForProvider(providerKey schema
BufferSize: 10,
},
}, nil
case schemas.Alibaba:
return &schemas.ProviderConfig{
NetworkConfig: schemas.NetworkConfig{
// Thinking models on 1M-context hosts need a generous timeout.
DefaultRequestTimeoutInSeconds: 180,
MaxRetries: 10,
RetryBackoffInitial: 5 * time.Second,
RetryBackoffMax: 3 * time.Minute,
},
ConcurrencyAndBufferSize: schemas.ConcurrencyAndBufferSize{
Concurrency: Concurrency,
BufferSize: 10,
},
}, nil
case schemas.Kimi:
return &schemas.ProviderConfig{
NetworkConfig: schemas.NetworkConfig{
DefaultRequestTimeoutInSeconds: 180,
MaxRetries: 10,
RetryBackoffInitial: 5 * time.Second,
RetryBackoffMax: 3 * time.Minute,
},
ConcurrencyAndBufferSize: schemas.ConcurrencyAndBufferSize{
Concurrency: Concurrency,
BufferSize: 10,
},
}, nil
case schemas.Zhipu:
return &schemas.ProviderConfig{
NetworkConfig: schemas.NetworkConfig{
DefaultRequestTimeoutInSeconds: 180,
MaxRetries: 10,
RetryBackoffInitial: 5 * time.Second,
RetryBackoffMax: 3 * time.Minute,
},
ConcurrencyAndBufferSize: schemas.ConcurrencyAndBufferSize{
Concurrency: Concurrency,
BufferSize: 10,
},
}, nil
case schemas.Wafer:
return &schemas.ProviderConfig{
NetworkConfig: schemas.NetworkConfig{
Expand Down
12 changes: 9 additions & 3 deletions core/internal/llmtests/chat_completion_stream.go
Original file line number Diff line number Diff line change
Expand Up @@ -209,8 +209,10 @@ func RunChatCompletionStreamTest(t *testing.T, client *bifrost.Bifrost, ctx cont

responseCount++

// Safety check to prevent infinite loops in case of issues
if responseCount > 500 {
// Safety check to prevent infinite loops in case of issues.
// Generous bound: token-by-token reasoning streams (e.g. GLM) can
// legitimately produce well over 500 chunks for a single response.
if responseCount > 2500 {
t.Fatal("Received too many streaming chunks, something might be wrong")
}

Expand Down Expand Up @@ -406,7 +408,11 @@ func RunChatCompletionStreamTest(t *testing.T, client *bifrost.Bifrost, ctx cont
}
}

if responseCount > 100 {
if responseCount > 1500 {
// Runaway stream: the tool-detection window is a safety bound,
// so tripping it must fail validation instead of letting a
// pathological stream pass on the strength of an earlier tool event.
streamErrors = append(streamErrors, "❌ Received too many streaming chunks in tool-call stream, something might be wrong")
goto toolStreamComplete
}

Expand Down
12 changes: 8 additions & 4 deletions core/internal/llmtests/responses_stream.go
Original file line number Diff line number Diff line change
Expand Up @@ -249,8 +249,10 @@ func RunResponsesStreamTest(t *testing.T, client *bifrost.Bifrost, ctx context.C

responseCount++

// Safety check to prevent infinite loops
if responseCount > 500 {
// Safety check to prevent infinite loops.
// Generous bound: token-by-token reasoning streams (e.g. GLM)
// can legitimately produce well over 500 chunks.
if responseCount > 2500 {
return ResponsesStreamValidationResult{
Passed: false,
Errors: []string{"❌ Received too many streaming chunks, something might be wrong"},
Expand Down Expand Up @@ -484,8 +486,10 @@ func RunResponsesStreamTest(t *testing.T, client *bifrost.Bifrost, ctx context.C
}
}

if responseCount > 100 {
goto toolStreamComplete
if responseCount > 1500 {
// Runaway stream: tripping the tool-detection safety bound must
// fail the test — an earlier tool event cannot mask it.
t.Fatalf("❌ Received too many streaming chunks in tool-call stream (%d), something might be wrong", responseCount)
}

case <-streamCtx.Done():
Expand Down
12 changes: 12 additions & 0 deletions core/internal/llmtests/validation_presets.go
Original file line number Diff line number Diff line change
Expand Up @@ -486,6 +486,18 @@ func ModifyExpectationsForProvider(expectations ResponseExpectations, provider s
expectations.ShouldHaveUsageStats = true
expectations.ShouldHaveLatency = true

case schemas.Alibaba:
expectations.ShouldHaveUsageStats = true
expectations.ShouldHaveLatency = true

case schemas.Kimi:
expectations.ShouldHaveUsageStats = true
expectations.ShouldHaveLatency = true

case schemas.Zhipu:
expectations.ShouldHaveUsageStats = true
expectations.ShouldHaveLatency = true

case schemas.Wafer:
expectations.ShouldHaveUsageStats = true
expectations.ShouldHaveLatency = true
Expand Down
Loading