Skip to content

fix: evict stale VK-scoped model configs from in-memory store on ReloadVirtualKey - #4043

Merged
akshaydeo merged 1 commit into
devfrom
06-04-fix_remove_stale_in_memory_model_configs
Jun 4, 2026
Merged

fix: evict stale VK-scoped model configs from in-memory store on ReloadVirtualKey#4043
akshaydeo merged 1 commit into
devfrom
06-04-fix_remove_stale_in_memory_model_configs

Conversation

@BearTS

@BearTS BearTS commented Jun 4, 2026

Copy link
Copy Markdown
Contributor

Summary

When a virtual key's governance model configs are deleted from the DB (e.g. when a standalone VK is adopted into an access profile and its VK-scoped model configs are removed), the in-memory store was never evicted of those stale entries. This caused deleted budget/rate-limit configs to continue enforcing against requests.

Changes

  • Added ScopedModelConfigIDs(scope, scopeID string) []string to GovernanceStore interface and implemented it on LocalGovernanceStore. It iterates the in-memory model config map and returns the IDs of all entries matching the given scope and scope ID.
  • Updated ReloadVirtualKey in the HTTP server to snapshot in-memory VK-scoped config IDs before upserting the DB results, then delete any IDs that were not present in the DB response, evicting stale entries.

Type of change

  • Bug fix
  • Feature
  • Refactor
  • Documentation
  • Chore/CI

Affected areas

  • Core (Go)
  • Transports (HTTP)
  • Providers/Integrations
  • Plugins
  • UI (React)
  • Docs

How to test

  1. Create a virtual key with a VK-scoped governance model config (e.g. a budget or rate limit).
  2. Adopt the virtual key into an access profile, causing the VK-scoped model config to be deleted from the DB.
  3. Trigger a ReloadVirtualKey for that VK.
  4. Verify that the previously enforced budget/rate limit is no longer applied to subsequent requests.
go test ./plugins/governance/... ./transports/bifrost-http/...

Breaking changes

  • Yes
  • No

Related issues

Security considerations

None. This change only affects in-memory eviction of governance configs and does not touch auth, secrets, or PII.

Checklist

  • I read docs/contributing/README.md and followed the guidelines
  • I added/updated tests where appropriate
  • I updated documentation where needed
  • I verified builds succeed (Go and UI)
  • I verified the CI pipeline passes locally if applicable

Summary by CodeRabbit

  • New Features

    • Enumerate scoped model configurations for targeted management.
    • Atomically apply incremental rate-limit usage adjustments.
  • Bug Fixes

    • Reload now evicts stale virtual-key-scoped configurations to prevent outdated settings persisting.
    • Rate-limit updates support optional transactional execution and consistent, hook-free writes.
  • Tests

    • Updated mocks to reflect transactional rate-limit update support.

@coderabbitai

coderabbitai Bot commented Jun 4, 2026

Copy link
Copy Markdown
Contributor

Review Change Stack

Caution

Review failed

The pull request is closed.

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 437d9940-8a6c-4e2f-8d74-98160c4f4df6

📥 Commits

Reviewing files that changed from the base of the PR and between 2f3a3cf and 9976675.

📒 Files selected for processing (5)
  • framework/configstore/rdb.go
  • framework/configstore/store.go
  • plugins/governance/store.go
  • transports/bifrost-http/lib/config_test.go
  • transports/bifrost-http/server/server.go

📝 Walkthrough

Walkthrough

Adds GovernanceStore.ScopedModelConfigIDs and LocalGovernanceStore.BumpRateLimitUsageBy; ReloadVirtualKey now evicts stale VK-scoped in-memory model configs after fetching DB results; ConfigStore.UpdateRateLimitUsage accepts an optional variadic *gorm.DB and RDB uses it when present.

Changes

Scoped configs, rate-limit usage, and VK reload

Layer / File(s) Summary
ConfigStore interface change
framework/configstore/store.go
ConfigStore.UpdateRateLimitUsage now accepts an optional variadic tx ...*gorm.DB parameter.
RDB implementation: optional tx
framework/configstore/rdb.go
RDBConfigStore.UpdateRateLimitUsage selects a provided transactional *gorm.DB when supplied and uses it for the update.
Mock UpdateRateLimitUsage signature
transports/bifrost-http/lib/config_test.go
MockConfigStore.UpdateRateLimitUsage updated to accept tx ...*gorm.DB (no-op implementation).
LocalGovernanceStore rate-limit bump helper
plugins/governance/store.go
Adds BumpRateLimitUsageBy to atomically apply token/request deltas via a CAS loop; fast-path no-op for zero deltas or missing rate limit.
Governance store scoped model config lookup
plugins/governance/store.go
Adds ScopedModelConfigIDs(scope, scopeID string) to list in-memory TableModelConfig IDs matching a (scope, scopeID) pair (treats nil ScopeID as "").
ReloadVirtualKey stale config eviction
transports/bifrost-http/server/server.go
ReloadVirtualKey pre-fetches VK-scoped model configs, upserts fetched records into in-memory state, snapshots prior scoped IDs, and deletes stale VK-scoped entries not present in the latest DB response.

Sequence Diagram(s)

sequenceDiagram
  participant Server
  participant GovernanceStore
  participant DB
  Server->>DB: Fetch VK-scoped model configs
  Server->>GovernanceStore: ScopedModelConfigIDs(scope, scopeID)
  Server->>GovernanceStore: Upsert/update in-memory configs for fetched IDs
  Server->>GovernanceStore: Delete scoped IDs not present in fetched set
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

  • maximhq/bifrost#3938: Modifies plugins/governance/store.go around in-memory governance model config selection and rate-limit usage updates.
  • maximhq/bifrost#3599: Related VK-scoped reload/stale-entry handling changes in transports/bifrost-http/server/server.go.
  • maximhq/bifrost#3940: Related scoped-model governance refactor used by these additions.

Suggested reviewers

  • roroghost17
  • akshaydeo
  • Pratham-Mishra04

Poem

🐇 I count the scoped configs one by one,
I nudge the tokens till the CAS is done.
I fetch fresh rows and trim what’s stale,
Cache and DB align without derail.
Hooray — a tidy sync, now hop and run!

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Title check ✅ Passed The title accurately and specifically describes the main change: evicting stale VK-scoped model configs from the in-memory store during ReloadVirtualKey, which is the core bug fix.
Description check ✅ Passed The description is complete, covering the problem, solution, type of change, affected areas, testing instructions, and security considerations; minor template items are unchecked but not required for functionality.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch 06-04-fix_remove_stale_in_memory_model_configs

Warning

There were issues while running some tools. Please review the errors and either fix the tool's configuration or disable the tool if it's a critical failure.

🔧 golangci-lint (2.12.2)

level=error msg="[linters_context] typechecking error: pattern ./...: directory prefix . does not contain main module or its selected dependencies"


Comment @coderabbitai help to get the list of available commands and usage tips.

@BearTS BearTS changed the title fix: remove stale in memory model configs fix: evict stale VK-scoped model configs from in-memory store on ReloadVirtualKey Jun 4, 2026

BearTS commented Jun 4, 2026

Copy link
Copy Markdown
Contributor Author

@BearTS
BearTS marked this pull request as ready for review June 4, 2026 02:09
@BearTS

BearTS commented Jun 4, 2026

Copy link
Copy Markdown
Contributor Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jun 4, 2026

Copy link
Copy Markdown
Contributor
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@transports/bifrost-http/server/server.go`:
- Around line 399-414: The race occurs because staleIDs is built after the DB
fetch so an interleaving ReloadVirtualKey can cause an older run to delete newer
in-memory configs; fix by taking the snapshot of current in-memory VK-scoped
config IDs from
governancePlugin.GetGovernanceStore().ScopedModelConfigIDs(tables.ModelConfigScopeVirtualKey,
id) into staleIDs before performing the DB fetch that produces mcs, then proceed
to delete entries removed from that snapshot only after you
UpdateModelConfigInMemory(ctx, &mcs[i]) for each returned config and finally
call DeleteModelConfigInMemory(ctx, mcID) for any remaining IDs — ensure you
reference the existing functions/vars (store, mcs, id, ctx,
UpdateModelConfigInMemory, DeleteModelConfigInMemory) and move the snapshot code
earlier so eviction is based on the pre-fetch view.
- Around line 396-415: The current ReloadVirtualKey path swallows errors from
GetModelConfigsByScopeAndScopeIDs causing a silent partial reload; change the
branch so that if GetModelConfigsByScopeAndScopeIDs returns an error you
propagate/return that error (or wrap it) from ReloadVirtualKey instead of
proceeding, so VK-scoped model-config in-memory updates
(store.UpdateModelConfigInMemory / store.DeleteModelConfigInMemory) only run
when the DB fetch succeeds; reference the call to
GetModelConfigsByScopeAndScopeIDs and ensure ReloadVirtualKey returns an error
on err != nil rather than treating it as a no-op.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 1433cef2-bf40-4937-b508-58fc48fc08f9

📥 Commits

Reviewing files that changed from the base of the PR and between 1e6aa2d and 99f52ff.

📒 Files selected for processing (2)
  • plugins/governance/store.go
  • transports/bifrost-http/server/server.go

Comment thread transports/bifrost-http/server/server.go Outdated
Comment thread transports/bifrost-http/server/server.go Outdated
@greptile-apps

greptile-apps Bot commented Jun 4, 2026

Copy link
Copy Markdown
Contributor

Confidence Score: 5/5

The core fix is safe to merge; the eviction logic correctly diffs in-memory state against the DB result and the DB fetch now gates all in-memory mutation.

The eviction logic in ReloadVirtualKey is straightforward and correct: snapshot, diff, upsert, delete. The DB fetch is moved before any mutation, so a DB error aborts cleanly. The only notable addition without a call site is BumpRateLimitUsageBy, which is unreachable via the GovernanceStore interface and has no callers, but it does not affect the correctness of existing paths.

plugins/governance/store.go — BumpRateLimitUsageBy is dead code and not wired into the interface.

Important Files Changed

Filename Overview
plugins/governance/store.go Adds ScopedModelConfigIDs to the GovernanceStore interface and LocalGovernanceStore, plus a new BumpRateLimitUsageBy method on LocalGovernanceStore that is not in the interface and has no callers in this PR.
transports/bifrost-http/server/server.go ReloadVirtualKey refactored to fetch DB model configs before mutating in-memory state, then evict stale VK-scoped entries. Logic is correct; early-exit on DB error prevents partial mutation.
framework/configstore/rdb.go UpdateRateLimitUsage gains an optional variadic tx parameter for transactional execution; existing callers are unaffected.
framework/configstore/store.go ConfigStore interface updated to match the new UpdateRateLimitUsage variadic signature.
transports/bifrost-http/lib/config_test.go MockConfigStore updated to match the new UpdateRateLimitUsage variadic signature; no functional change.

Reviews (6): Last reviewed commit: "fix: remove stale in memory model config..." | Re-trigger Greptile

Comment thread transports/bifrost-http/server/server.go Outdated
Comment thread plugins/governance/store.go
@BearTS
BearTS force-pushed the 06-04-fix_remove_stale_in_memory_model_configs branch 2 times, most recently from e8aa278 to ac4efb8 Compare June 4, 2026 12:48
@BearTS
BearTS force-pushed the 06-04-feat_add_txn_to_update_budget_usage branch from 1e6aa2d to b53c920 Compare June 4, 2026 15:59
@BearTS
BearTS force-pushed the 06-04-fix_remove_stale_in_memory_model_configs branch from ac4efb8 to 2f3a3cf Compare June 4, 2026 15:59

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@framework/configstore/rdb.go`:
- Around line 3011-3035: The VK cleanup currently only collects legacy BudgetID
from scopedModelConfigs and misses budgets owned via the Budgets relation,
leaving orphaned governance_budgets rows; update the loop over
scopedModelConfigs (the scopedModelConfigs []tables.TableModelConfig iteration)
to also extract IDs from each mc.Budgets collection (and from any nested budget
field names used in tables.TableModelConfig) and add them to budgetIDs,
de-duplicating (e.g., via a map) before running the Delete on TableBudget; keep
the existing handling for mc.BudgetID and mc.RateLimitID and ensure budgetIDs is
populated with both legacy and relation-owned budget IDs prior to the
txDB.Delete call.

In `@plugins/governance/store.go`:
- Around line 461-481: The BumpRateLimitUsageBy method currently accepts
"arbitrary token and request deltas" but doesn't state whether negatives are
allowed; add explicit validation or documentation: update the
BumpRateLimitUsageBy function to either reject negative deltas (e.g., check
tokenDelta < 0 || requestDelta < 0 and return a descriptive error) or add a
clear comment/docstring on BumpRateLimitUsageBy indicating that negative deltas
are permitted for corrections and ensure any subsequent logic (the clone
adjustments and CompareAndSwap on gs.rateLimits) safely supports decreasing
usage; pick one approach and apply consistently in the function and its callers.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 82cd2c48-d1a9-4cbd-9973-ea0571713aaa

📥 Commits

Reviewing files that changed from the base of the PR and between ac4efb8 and 2f3a3cf.

📒 Files selected for processing (5)
  • framework/configstore/rdb.go
  • framework/configstore/store.go
  • plugins/governance/store.go
  • transports/bifrost-http/lib/config_test.go
  • transports/bifrost-http/server/server.go
💤 Files with no reviewable changes (2)
  • transports/bifrost-http/server/server.go
  • transports/bifrost-http/lib/config_test.go

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Inline review comments failed to post. This is likely due to GitHub's internal server error or limits when posting large numbers of comments. If you are seeing this consistently it is likely a permissions issue. Please check "Moderation" -> "Code review limits" under your organization settings.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@framework/configstore/rdb.go`:
- Around line 3011-3035: The VK cleanup currently only collects legacy BudgetID
from scopedModelConfigs and misses budgets owned via the Budgets relation,
leaving orphaned governance_budgets rows; update the loop over
scopedModelConfigs (the scopedModelConfigs []tables.TableModelConfig iteration)
to also extract IDs from each mc.Budgets collection (and from any nested budget
field names used in tables.TableModelConfig) and add them to budgetIDs,
de-duplicating (e.g., via a map) before running the Delete on TableBudget; keep
the existing handling for mc.BudgetID and mc.RateLimitID and ensure budgetIDs is
populated with both legacy and relation-owned budget IDs prior to the
txDB.Delete call.

In `@plugins/governance/store.go`:
- Around line 461-481: The BumpRateLimitUsageBy method currently accepts
"arbitrary token and request deltas" but doesn't state whether negatives are
allowed; add explicit validation or documentation: update the
BumpRateLimitUsageBy function to either reject negative deltas (e.g., check
tokenDelta < 0 || requestDelta < 0 and return a descriptive error) or add a
clear comment/docstring on BumpRateLimitUsageBy indicating that negative deltas
are permitted for corrections and ensure any subsequent logic (the clone
adjustments and CompareAndSwap on gs.rateLimits) safely supports decreasing
usage; pick one approach and apply consistently in the function and its callers.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 82cd2c48-d1a9-4cbd-9973-ea0571713aaa

📥 Commits

Reviewing files that changed from the base of the PR and between ac4efb8 and 2f3a3cf.

📒 Files selected for processing (5)
  • framework/configstore/rdb.go
  • framework/configstore/store.go
  • plugins/governance/store.go
  • transports/bifrost-http/lib/config_test.go
  • transports/bifrost-http/server/server.go
💤 Files with no reviewable changes (2)
  • transports/bifrost-http/server/server.go
  • transports/bifrost-http/lib/config_test.go
🛑 Comments failed to post (2)
framework/configstore/rdb.go (1)

3011-3035: ⚠️ Potential issue | 🟠 Major | ⚡ Quick win

DeleteVirtualKey cleanup misses has-many model-config budgets.

This path deletes VK-scoped model configs but only collects legacy BudgetID. Model configs that own budgets through Budgets can leave orphaned governance_budgets rows after VK deletion.

💡 Suggested fix
 		var scopedModelConfigs []tables.TableModelConfig
 		if err := txDB.WithContext(ctx).
+			Preload("Budgets").
 			Where("scope = ? AND scope_id = ?", tables.ModelConfigScopeVirtualKey, id).
 			Find(&scopedModelConfigs).Error; err != nil {
 			return err
 		}
 		budgetIDs := make([]string, 0, len(scopedModelConfigs))
 		rateLimitIDs := make([]string, 0, len(scopedModelConfigs))
 		for _, mc := range scopedModelConfigs {
+			for i := range mc.Budgets {
+				budgetIDs = append(budgetIDs, mc.Budgets[i].ID)
+			}
 			if mc.BudgetID != nil {
 				budgetIDs = append(budgetIDs, *mc.BudgetID)
 			}
 			if mc.RateLimitID != nil {
 				rateLimitIDs = append(rateLimitIDs, *mc.RateLimitID)
 			}
 		}
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

		var scopedModelConfigs []tables.TableModelConfig
		if err := txDB.WithContext(ctx).
			Preload("Budgets").
			Where("scope = ? AND scope_id = ?", tables.ModelConfigScopeVirtualKey, id).
			Find(&scopedModelConfigs).Error; err != nil {
			return err
		}
		budgetIDs := make([]string, 0, len(scopedModelConfigs))
		rateLimitIDs := make([]string, 0, len(scopedModelConfigs))
		for _, mc := range scopedModelConfigs {
			for i := range mc.Budgets {
				budgetIDs = append(budgetIDs, mc.Budgets[i].ID)
			}
			if mc.BudgetID != nil {
				budgetIDs = append(budgetIDs, *mc.BudgetID)
			}
			if mc.RateLimitID != nil {
				rateLimitIDs = append(rateLimitIDs, *mc.RateLimitID)
			}
		}
		if err := txDB.WithContext(ctx).
			Where("scope = ? AND scope_id = ?", tables.ModelConfigScopeVirtualKey, id).
			Delete(&tables.TableModelConfig{}).Error; err != nil {
			return err
		}
		if len(budgetIDs) > 0 {
			if err := txDB.WithContext(ctx).Delete(&tables.TableBudget{}, "id IN ?", budgetIDs).Error; err != nil {
				return err
			}
		}
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@framework/configstore/rdb.go` around lines 3011 - 3035, The VK cleanup
currently only collects legacy BudgetID from scopedModelConfigs and misses
budgets owned via the Budgets relation, leaving orphaned governance_budgets
rows; update the loop over scopedModelConfigs (the scopedModelConfigs
[]tables.TableModelConfig iteration) to also extract IDs from each mc.Budgets
collection (and from any nested budget field names used in
tables.TableModelConfig) and add them to budgetIDs, de-duplicating (e.g., via a
map) before running the Delete on TableBudget; keep the existing handling for
mc.BudgetID and mc.RateLimitID and ensure budgetIDs is populated with both
legacy and relation-owned budget IDs prior to the txDB.Delete call.
plugins/governance/store.go (1)

461-481: ⚠️ Potential issue | 🟡 Minor | ⚡ Quick win

Clarify whether negative deltas are allowed or add validation.

The function comment mentions "arbitrary token and request deltas" but also states it's "used to fold accumulated usage carried from another rate limit," which suggests positive values. Should negative deltas be allowed for usage corrections, or should they be rejected? Consider either:

  1. Adding validation: if tokenDelta < 0 || requestDelta < 0 { return fmt.Errorf("deltas must be non-negative") }, or
  2. Documenting explicitly that negative deltas are allowed for corrections.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@plugins/governance/store.go` around lines 461 - 481, The BumpRateLimitUsageBy
method currently accepts "arbitrary token and request deltas" but doesn't state
whether negatives are allowed; add explicit validation or documentation: update
the BumpRateLimitUsageBy function to either reject negative deltas (e.g., check
tokenDelta < 0 || requestDelta < 0 and return a descriptive error) or add a
clear comment/docstring on BumpRateLimitUsageBy indicating that negative deltas
are permitted for corrections and ensure any subsequent logic (the clone
adjustments and CompareAndSwap on gs.rateLimits) safely supports decreasing
usage; pick one approach and apply consistently in the function and its callers.

@BearTS
BearTS force-pushed the 06-04-fix_remove_stale_in_memory_model_configs branch from 2f3a3cf to 8d85c06 Compare June 4, 2026 16:17

akshaydeo commented Jun 4, 2026

Copy link
Copy Markdown
Contributor

Merge activity

  • Jun 4, 4:31 PM UTC: A user started a stack merge that includes this pull request via Graphite.
  • Jun 4, 4:33 PM UTC: Graphite rebased this pull request as part of a merge.

@akshaydeo
akshaydeo changed the base branch from 06-04-feat_add_txn_to_update_budget_usage to graphite-base/4043 June 4, 2026 16:31
@akshaydeo
akshaydeo changed the base branch from graphite-base/4043 to dev June 4, 2026 16:31
@akshaydeo
akshaydeo force-pushed the 06-04-fix_remove_stale_in_memory_model_configs branch from 8d85c06 to 9976675 Compare June 4, 2026 16:32
@akshaydeo
akshaydeo merged commit 3ca2e70 into dev Jun 4, 2026
13 of 14 checks passed
@akshaydeo
akshaydeo deleted the 06-04-fix_remove_stale_in_memory_model_configs branch June 4, 2026 16:33
@akshaydeo akshaydeo mentioned this pull request Jun 5, 2026
18 tasks
akshaydeo added a commit that referenced this pull request Jun 6, 2026
## Summary

This PR bumps the Go toolchain version from `1.26.3` to `1.26.4` across all modules and CI workflows, and cuts a new release (`core` v1.5.17, `framework` v1.3.17, `transports` v1.5.9, `plugins/compat` v0.1.16, `plugins/governance` v1.5.17, and associated plugin versions) incorporating a large batch of features and fixes accumulated since the previous release.

## Changes

- **Go 1.26.4** — Updated `go-version` in all GitHub Actions workflows (`e2e-tests`, `helm-release`, `pr-tests`, `release-cli`, `release-pipeline`, `snyk`) and all `go.mod` files (core, framework, transports, cli, all plugins, examples, and test modules).
- **Core (v1.5.17)** — OpenAI compaction support, multi-customer logs and usage tracking, multiple team/business unit support, `request_headers` wildcard pattern capture for OTel and Maxim plugins, xAI `x_search` tool, fetch URL validation with SSRF hardening, `file://` pricing URL scheme, virtual key provider fan-out filtering, and a broad set of fixes including Anthropic prompt cache key, empty thinking block stripping, OpenAI stream usage event cleanup, Gemini numeric schema constraints, stale connection retries, Azure Claude diagnostic strip, and passthrough budget handling.
- **Framework (v1.3.17)** — Scope-aware budgets and limits wired from model configs, provider-level governance, multiple customer budget support with `calendar_aligned` windows, paginated virtual key fetch, `config.json` source-of-truth flow, FTS index cap reduction, sync worker drift fix, cascade deletes for model configs, and high-scale virtual key flow improvements.
- **Transports (v1.5.9)** — Full changelog covering all of the above plus UI improvements (log navigation, customer detail sheet, `BudgetDisplay` component, inline loading shell, materialized view alias), SCIM provisioning fields, Helm/config schema additions (`roles`, `per_user_oauth`), client IP resolution from forwarded headers, and dependency upgrades (`recharts` to 3.8.1, `golang.org/x` CVE remediation).
- **Plugins** — `governance` v1.5.17 adds team budget/rate-limit exporters, ghost node reconciliation fix, and VK double usage counting fix; `logging` v1.5.17 adds wildcard header capture and file attachment rendering; `otel` v1.2.17 adds `disable_content_logging` and multiple collectors support; `maxim` v1.6.17 adds `request_headers` wildcard capture; `compat` v0.1.16 fixes `max_tokens` preservation during param filtering.

## Type of change

- [ ] Bug fix
- [x] Feature
- [ ] Refactor
- [ ] Documentation
- [x] Chore/CI

## Affected areas

- [x] Core (Go)
- [x] Transports (HTTP)
- [x] Providers/Integrations
- [x] Plugins
- [x] UI (React)
- [ ] Docs

## How to test

```sh
# Verify Go version
go version  # should report go1.26.4

# Run core tests
cd core && go test ./...

# Run framework tests
cd framework && go test ./...

# Run transports tests
cd transports && go test ./...

# Run plugin tests
cd plugins/governance && go test ./...
cd plugins/logging && go test ./...
cd plugins/otel && go test ./...

# UI
cd ui
pnpm i
pnpm build
pnpm test
```

## Breaking changes

- [ ] Yes
- [x] No

## Related issues

#4053, #4066, #4041, #4012, #3976, #3947, #3991, #4045, #3957, #3938, #3937, #3939, #3981, #3998, #3997, #4092, #4091, #4079, #4080, #4086, #3929, #3994, #4028, #3970, #3919, #3861, #3664, #3999, #4088, #4070, #4051, #4043, #4057, #4023, #3941, #3955, #4024, #3956, #3967, #3925, #3992, #3900

## Security considerations

- Fetch URL validation hardened against SSRF by tightening IP checks for private networks and link-local addresses (#4092, #3947, #3991).
- Transitive `golang.org/x` dependencies (crypto, net, sys, text) bumped to address Docker Scout CVEs (#3900).

## Checklist

- [x] I read `docs/contributing/README.md` and followed the guidelines
- [x] I added/updated tests where appropriate
- [x] I updated documentation where needed
- [x] I verified builds succeed (Go and UI)
- [x] I verified the CI pipeline passes locally if applicable

<!-- This is an auto-generated comment: release notes by coderabbit.ai -->
## Summary by CodeRabbit

* **New Features**
  * OpenAI compaction, multi-customer/team logstore support, request-header wildcard capture, enhanced governance (provider-level & scope-aware limits), disable-content-logging option, support for multiple OpenTelemetry collectors, SSRF hardening and URL validation.

* **Chores**
  * Bumped Go toolchain across modules and updated component/plugin version releases.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
@akshaydeo akshaydeo mentioned this pull request Jun 7, 2026
akshaydeo added a commit that referenced this pull request Jun 7, 2026
## ✨ Features

- **OpenAI Compaction** — Added OpenAI conversation compaction support
across core, framework, logging, and the API surface (#4053)
- **Multi-Customer & Org Hierarchy** — Logs and usage tracking now
support multiple customers, teams, and business units, including
business unit CRUD, team assignment, and governance endpoints in the
OpenAPI spec (#4066, #4041, #4082)
- **Provider-Level Governance** — Budgets & limits are now scope-aware
and can be applied at the virtual-key top level and per provider, wired
from the model configs table, with UI filters for scope and providers
(#3938, #3937, #3939, #3981, #3962)
- **Customer Budgets** — Customers support multiple budgets and
`calendar_aligned` budget windows (#3998, #3997)
- **Virtual Key Attribution & Controls** — Added a `created_by` user
attribution column and a `blacklisted_models` column for virtual key
provider configs (#3672, #3653)
- **Request Header Capture** — OTel and Maxim observability plugins
capture `request_headers` by pattern, with wildcard support (e.g.
`x-custom-*`); logging gained the same wildcard header capture (#4012,
#3958)
- **OTel Content Controls & Collectors** — New `disable_content_logging`
option drops message/tool content from exported spans, plus support for
multiple OTel collectors (#4064, #3894)
- **xAI x_search** — Added xAI `x_search` tool support (#3976)
- **URL Validation** — Added fetch URL validation with private-network
configuration and link-local blocking (#3947, #3991)
- **File Scheme Pricing URLs** — Pricing source URLs now accept the
`file://` scheme for air-gapped and self-hosted deployments (#4045)
- **Paginated Virtual Keys** — Virtual key fetching is paginated to
handle deployments with very large numbers of keys (#3957)
- **Client IP Resolution** — Resolve client IP from
`X-Forwarded-For`/`X-Real-IP` headers
- **SCIM Provisioning** — Added `attributeType`/`attributeValue` SCIM
provisioning fields
- **Helm/Config Schema** — Added `roles` RBAC governance config and
`per_user_oauth` MCP auth to the Helm chart and config schema (#4004,
#4009)
- **Log Navigation UI** — Added a "View logs" menu item to customer,
team, and virtual key tables, clickable links in log detail views, a
customer detail sheet, and a reusable `BudgetDisplay` component (#4073,
#4054, #4026, #4055)
- **Faster First Paint** — Added an inline loading shell to `#root`
before React mounts (#4063)
- **Materialized View Alias** — Added an `alias` column to the
materialized view with filter support (#4078)

## 🐞 Fixed

- **Fetch URL IP Checks** — Hardened fetch URL IP checks against SSRF
(#4092)
- **Mantle Model Matching** — Broadened Mantle model matching to all
`gpt` variants (#4091)
- **Empty Thinking Blocks** — Strip thinking blocks when the signature
is empty (#4079)
- **OpenAI Stream Usage** — Removed usage from the `responses.created`
event in the OpenAI stream (#4080)
- **Prompt Cache Key** — Set the prompt cache key from the Anthropic
integration (#4086)
- **Upstream Failure Status** — Map upstream connection failures to 502
instead of 400 (#3929) (thanks
[@chris-colinsky](https://github.com/chris-colinsky)!)
- **Gemini Schema Constraints** — Accept numeric schema integer
constraints for Gemini (#3994) (thanks
[@yanhao98](https://github.com/yanhao98)!)
- **Files Provider Param** — Accept the `?provider=` query param on `GET
/v1/files` (#3971) (thanks [@alexef](https://github.com/alexef)!)
- **Optional Batch Model** — Made the `model` field optional on `POST
/v1/batches` (#3973) (thanks [@alexef](https://github.com/alexef)!)
- **Helm Azure Config** — Added missing `azure_key_config` fields to the
Helm schema (#3996) (thanks
[@axelray-dev](https://github.com/axelray-dev)!)
- **Text Completion Chunk Model** — Added the missing `Model` field to
`TextCompletionChunkResponse` (#3970) (thanks
[@kuishou68](https://github.com/kuishou68)!)
- **MCP Inline stdio Env** — MCP stdio server configs accept inline
environment variable assignments (#3861) (thanks
[@Shushmitaaaa](https://github.com/Shushmitaaaa)!)
- **Orphaned Tool Results** — Orphaned tool results in the OpenAI to
Anthropic conversion flow are no longer rejected by the Anthropic API
(#3919)
- **Node Usage Reconciliation** — Added a monotonic `inc_number` log
cursor so node usage reconciliation does not skip late async log writes
(#3664)
- **Bedrock Output Assessments** — Corrected the type of
`outputAssessments` in Bedrock responses (#4028)
- **Model Pool Pricing Reloads** — Preserve non-pricing model pool
entries across pricing reloads (#3999)
- **Ghost Node Reconciliation** — Replicate the VK hierarchy flow for
ghost node reconciliation (#4088)
- **VK Double Usage Counting** — Fixed double usage counting when
creating a virtual key (#4070)
- **Model Config Lifecycle** — Cascade deletes for model configs and
removal of stale in-memory model configs (#4051, #4043)
- **FTS Index Cap** — Reduced the FTS index `left()` cap from 800k to
250k chars to stay within the tsvector limit (#4057)
- **Sync Worker Drift** — Reduced the sync worker ticker period to 5m to
prevent threshold drift (#4023)
- **Passthrough** — Fixed passthrough budgets, gated passthrough models
per VK, model extraction for Azure passthrough, and restricted
fallbacks/provider selection to the VK boundary (#3941, #3988, #3983,
#3924)
- **Provider Response Headers** — Strip provider response headers and
add a content-type filter (#3955, #4024)
- **Stream Handling** — Drain non-SSE stream readers and retry stale
connections (#3956, #3967)
- **Azure Claude** — Strip Azure diagnostic property for Claude models
(#3925)
- **Compat max_tokens** — Preserve chat `max_tokens` during param
filtering (#3992)
- **Raw Request Flag** — Removed the raw request flag from providers
that don't support it (#4058)
- **UI Fixes** — Standardized page container layout, virtual key model
configs UI, and dashboard chart tooltips (#4046, #4052, #4044)

## 🔧 Maintenance

- **Dependency Upgrades** — Bumped transitive `golang.org/x`
dependencies (crypto, net, sys, text) for Docker Scout CVE remediation
and `recharts` to 3.8.1; cascaded version bumps across all modules
(#3900, #4003)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants