fix(billing): correct cache token double-counting in Claude relay to non-Claude providers - #3258
fix(billing): correct cache token double-counting in Claude relay to non-Claude providers#3258147API wants to merge 5401 commits into
Conversation
…d improved rate limiting
fix: /v1/chat/completions -> /v1/responses json_schema
将散落在多个文件中的预扣费/结算/退款逻辑抽象为统一的 BillingSession 生命周期管理: - 新增 BillingSettler 接口 (relay/common/billing.go) 避免循环引用 - 新增 FundingSource 接口 + WalletFunding / SubscriptionFunding 实现 (service/funding_source.go) - 新增 BillingSession 封装预扣/结算/退款原子操作 (service/billing_session.go) - 新增 SettleBilling 统一结算辅助函数,替换各 handler 中的 quotaDelta 模式 - 重写 PreConsumeBilling 为 BillingSession 工厂入口 - controller/relay.go 退款守卫改用 BillingSession.Refund() 修复的 Bug: - 令牌额度泄漏:PreConsumeTokenQuota 成功但 DecreaseUserQuota 失败时未回滚 - 订阅退款遗漏:FinalPreConsumedQuota=0 但 SubscriptionPreConsumed>0 时跳过退款 - 订阅多扣费:subConsume 强制为 1 但 FinalPreConsumedQuota 不同步 - 退款路径不统一:钱包/订阅退款逻辑现统一由 FundingSource.Refund 分派
- Settle 部分失败保护:新增 fundingSettled 标记,资金来源提交后 令牌调整失败不再导致 Refund 误退已结算的资金 - 订阅多扣费修复:trySubscription 传 subConsume 而非 preConsumedQuota 给 preConsume,保证三者(amount/preConsume/FinalPreConsumedQuota)一致 - 令牌回滚错误记录:preConsume 中 funding 失败时令牌回滚错误不再丢弃 - 移除钱包路径死代码:用户额度不足的 strings.Contains 匹配不可能命中 - WalletFunding.Refund 不重试:IncreaseUserQuota 非幂等,重试会多退
…e recharge card tabs - Defaulting to subscriptions when available and avoiding initial flash when no plans exist. - Adjust the wide-screen layout to place wallet and invite sections side by side, simplify the subscription header and controls, and add padding to prevent card borders from clipping. - Update related i18n strings by adding the new tab label and removing the obsolete subscription blurb.
…iption-card-when-no-plans ✨ refactor(wallet): Top-up layout to embed subscription plans into the recharge card tabs
…-session refactor: 抽象统一计费会话 BillingSession
Add a lightweight active-subscription check to skip subscription pre-consume when none exist, reducing unnecessary transactions and locks. In the subscription UI, disable subscription-first options when no active plan is available, show the effective fallback to wallet with a clear notice, and distinguish “invalidated” from “expired” states. Update i18n strings across supported locales to reflect the new messages and status labels.
Aligns the error variable types in the subscription-first path so that quota fallback checks use the correct NewAPIError. This prevents build failures and preserves the intended wallet fallback when subscription pre-consume returns an insufficient quota error.
Routes quota alerts through a subscription-specific check when billing from subscriptions, preventing wallet-based thresholds from triggering false warnings. Updates the notification settings description and localization keys to clarify that both wallet and subscription balances are monitored.
…n-quota-notify 🔔 feat: Add subscription-aware quota notifications and update UI copy
…-preference-fallback ✨ chore: Improve subscription billing fallback and UI states
…tumNous#2881) 当上游为 AWS Bedrock 时,message_delta 的 usage 可能缺少 input_tokens、 cache_creation_input_tokens、cache_read_input_tokens 等字段,导致与原生 Anthropic 格式不一致。从 message_start 积累的 claudeInfo 中补全这些字段后 重新序列化,确保客户端收到一致的 usage 格式。
Modified the formatUserLogs function to include a startIdx parameter, allowing for more flexible log ID assignment. Updated calls to this function in GetLogByTokenId and GetUserLogs to pass the appropriate starting index.
feat: add Codex channel disclaimer (i18n, OpenAI terms)
feat: Force beta=true parameter for Anthropic channel
feat(oauth): implement custom OAuth provider
fix: Claude stream block index/type transitions
fix: add paragraph breaks between reasoning summary chunks
# Conflicts: # service/openaicompat/chat_to_responses.go
…t-stream feat: channel test with stream=true
…fo-input-token fix: 使用openai兼容接口调用部分渠道在最终端点为claude原生端点下还是走了openai扣减input_token的逻辑
…amOverrideEditorModal
为渠道参数覆盖可视化规则提供拖拽排序支持
…4f8a4248b0ab3b03ba703796ea3 fix: kling risk fail return openAIVideo error
fix: add explicit docker-compose networks
…ride-beta-header-append feat:support $keep_only_declared and deduped $append for header override
chore: update model lists for frequently used channels
…tion links and submission checks
Round remaining balance
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Pro Run ID: 📒 Files selected for processing (1)
🚧 Files skipped from review as they are similar to previous changes (1)
WalkthroughThe PR modifies quota consumption in Changes
Sequence Diagram(s)(Skipped — change is a small internal control-flow tweak, not a multi-component feature requiring a sequence diagram.) Estimated code review effort🎯 2 (Simple) | ⏱️ ~10 minutes Possibly related PRs
Suggested reviewers
Poem
🚥 Pre-merge checks | ✅ 2 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (2 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches🧪 Generate unit tests (beta)
📝 Coding Plan
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
48ef46b to
ecd7203
Compare
…non-Claude providers PostClaudeConsumeQuota assumed Claude usage semantics (input_tokens excludes cached tokens) for all providers. When a Claude-format request (/v1/messages) is relayed to a non-Claude provider (e.g. Gemini, OpenAI), the upstream returns promptTokens that already include cached tokens. Without subtracting them, cache tokens were billed twice: once as part of promptTokens and again as cacheTokens with cacheRatio. The fix aligns with the existing logic in compatible_handler.go's postConsumeQuota: check GetFinalRequestRelayFormat() and only subtract cached tokens from promptTokens when the usage follows non-Claude semantics. The OpenRouter-specific cache creation estimation logic is preserved as-is. --- 修复了 PostClaudeConsumeQuota 中缓存 token 重复计费的问题。 当 Claude 格式请求(/v1/messages)被转发到非 Claude 渠道(如 Gemini、OpenAI 等)时, 上游返回的 promptTokens 已包含缓存 token。由于未将缓存 token 从 promptTokens 中减去, 缓存 token 被计费两次:一次作为 promptTokens 的一部分,另一次按 cacheRatio 单独计费。 修复方式与 compatible_handler.go 中 postConsumeQuota 的现有逻辑保持一致: 通过 GetFinalRequestRelayFormat() 判断 usage 语义,仅在非 Claude 语义时 才从 promptTokens 中减去缓存 token。OpenRouter 特有的缓存创建 token 估算逻辑保持不变。
ecd7203 to
2370be6
Compare
|
@seefs001 @Calcium-Ion Could you please take a look at this fix when you have time? Thank you! |
|
问题场景是:claude code中调用gemini时出现的,缓存token与输入token重复计费 |
Problem
When a Claude-format request (
/v1/messages) is relayed to a non-Claude provider (e.g. Gemini, OpenAI),PostClaudeConsumeQuotaincorrectly assumes Claude usage semantics for the returned usage data.Claude API returns
input_tokensthat already excludes cached tokens — so no subtraction is needed.Gemini / OpenAI / etc. return
prompt_tokensthat includes cached tokens — socacheTokensmust be subtracted frompromptTokensbefore addingcacheTokens * cacheRatio, otherwise cached tokens get billed twice.Billing breakdown (before fix)
promptTokenscacheTokensThe 15802 cached tokens are charged twice: once at full input price, once at cache price.
Billing breakdown (after fix)
promptTokenscacheTokensRoot Cause
PostClaudeConsumeQuotaonly subtracted cache tokens forChannelTypeOpenRouter. The existingpostConsumeQuotainrelay/compatible_handler.goalready handles this correctly usingGetFinalRequestRelayFormat() == RelayFormatClaudeto distinguish usage semantics —PostClaudeConsumeQuotawas missing the same guard.Fix
Lift the cache-token subtraction out of the OpenRouter-only block and apply it to all non-Claude-semantic usage, consistent with
compatible_handler.go:The OpenRouter-specific cache-creation estimation logic is preserved as-is (it runs before the subtraction).
Affected Scenarios
Any request that:
/v1/messages)Made with Cursor
Summary by CodeRabbit