Skip to content

feat(openai): add responses reasoning input compatibility flag - #3237

Closed
bkidy wants to merge 5393 commits into
QuantumNous:mainfrom
bkidy:fix/responses-reasoning-compat
Closed

feat(openai): add responses reasoning input compatibility flag#3237
bkidy wants to merge 5393 commits into
QuantumNous:mainfrom
bkidy:fix/responses-reasoning-compat

Conversation

@bkidy

@bkidy bkidy commented Mar 12, 2026

Copy link
Copy Markdown

Summary

  • add a default-off REMOVE_RESPONSES_REASONING_INPUT flag for OpenAI Responses requests
  • strip reasoning* items from input when the flag is enabled before forwarding upstream
  • help partially compatible relays avoid account/session-bound errors such as Item with id 'rs_...' not found

Why

Some OpenAI-compatible relays can serve /v1/responses, but do not fully support account/session-bound reasoning state across turns. In those cases, replaying reasoning items from prior responses may fail with errors like:

  • Item with id 'rs_...' not found
  • invalid_encrypted_content

This PR adds a small compatibility escape hatch without changing the default behavior.

Behavior

  • default remains unchanged
  • when REMOVE_RESPONSES_REASONING_INPUT=true, reasoning* items are removed from the Responses input array before the request is sent upstream
  • malformed input items are preserved as-is

Trade-offs

This compatibility mode may reduce some Responses-native continuity/cost benefits because account-bound reasoning items are no longer replayed.

Summary by CodeRabbit

Release Notes

  • New Features
    • Added configurable option to remove reasoning data from responses via environment variable.
    • Integrated filtering mechanism with error handling for response processing.

Calcium-Ion and others added 30 commits February 6, 2026 16:16
🔒 fix(security): sanitize AI-generated HTML to prevent XSS in playground
…idation

- Add configurable per-user token creation limit (max_user_tokens)
- Sanitize search input patterns to prevent expensive queries
- Add per-user search rate limiting (by user ID)
- Add pagination to search endpoint with strict page size cap
- Skip empty search fields instead of matching nothing
- Hide internal errors from API responses
- Fix Interface2String float64 formatting causing config parse failures
- Add float-string fallback in config system for int/uint fields
fix: harden token search with pagination, rate limiting and input validation
- Change ESCAPE character from '\' to '!' for compatibility with MySQL/PostgreSQL/SQLite
- Adjust sanitization logic to escape '!' and '_' correctly, improving input validation for search queries
fix: /v1/chat/completions -> /v1/responses json_schema
将散落在多个文件中的预扣费/结算/退款逻辑抽象为统一的 BillingSession 生命周期管理:

- 新增 BillingSettler 接口 (relay/common/billing.go) 避免循环引用
- 新增 FundingSource 接口 + WalletFunding / SubscriptionFunding 实现 (service/funding_source.go)
- 新增 BillingSession 封装预扣/结算/退款原子操作 (service/billing_session.go)
- 新增 SettleBilling 统一结算辅助函数,替换各 handler 中的 quotaDelta 模式
- 重写 PreConsumeBilling 为 BillingSession 工厂入口
- controller/relay.go 退款守卫改用 BillingSession.Refund()

修复的 Bug:
- 令牌额度泄漏:PreConsumeTokenQuota 成功但 DecreaseUserQuota 失败时未回滚
- 订阅退款遗漏:FinalPreConsumedQuota=0 但 SubscriptionPreConsumed>0 时跳过退款
- 订阅多扣费:subConsume 强制为 1 但 FinalPreConsumedQuota 不同步
- 退款路径不统一:钱包/订阅退款逻辑现统一由 FundingSource.Refund 分派
- Settle 部分失败保护:新增 fundingSettled 标记,资金来源提交后
  令牌调整失败不再导致 Refund 误退已结算的资金
- 订阅多扣费修复:trySubscription 传 subConsume 而非 preConsumedQuota
  给 preConsume,保证三者(amount/preConsume/FinalPreConsumedQuota)一致
- 令牌回滚错误记录:preConsume 中 funding 失败时令牌回滚错误不再丢弃
- 移除钱包路径死代码:用户额度不足的 strings.Contains 匹配不可能命中
- WalletFunding.Refund 不重试:IncreaseUserQuota 非幂等,重试会多退
…e recharge card tabs

- Defaulting to subscriptions when available and avoiding initial flash when no plans exist.
- Adjust the wide-screen layout to place wallet and invite sections side by side, simplify the subscription header and controls, and add padding to prevent card borders from clipping.
- Update related i18n strings by adding the new tab label and removing the obsolete subscription blurb.
…iption-card-when-no-plans

✨ refactor(wallet): Top-up layout to embed subscription plans into the recharge card tabs
…-session

refactor: 抽象统一计费会话 BillingSession
Add a lightweight active-subscription check to skip subscription pre-consume when none exist, reducing unnecessary transactions and locks. In the subscription UI, disable subscription-first options when no active plan is available, show the effective fallback to wallet with a clear notice, and distinguish “invalidated” from “expired” states. Update i18n strings across supported locales to reflect the new messages and status labels.
Aligns the error variable types in the subscription-first path so that quota fallback checks use the correct NewAPIError.
This prevents build failures and preserves the intended wallet fallback when subscription pre-consume returns an insufficient quota error.
Routes quota alerts through a subscription-specific check when billing from subscriptions, preventing wallet-based thresholds from triggering false warnings.
Updates the notification settings description and localization keys to clarify that both wallet and subscription balances are monitored.
…n-quota-notify

🔔 feat: Add subscription-aware quota notifications and update UI copy
…-preference-fallback

✨ chore: Improve subscription billing fallback and UI states
…tumNous#2881)

当上游为 AWS Bedrock 时,message_delta 的 usage 可能缺少 input_tokens、
cache_creation_input_tokens、cache_read_input_tokens 等字段,导致与原生
Anthropic 格式不一致。从 message_start 积累的 claudeInfo 中补全这些字段后
重新序列化,确保客户端收到一致的 usage 格式。
Modified the formatUserLogs function to include a startIdx parameter, allowing for more flexible log ID assignment. Updated calls to this function in GetLogByTokenId and GetUserLogs to pass the appropriate starting index.
feitianbubu and others added 25 commits March 6, 2026 16:32
…es-body-no-retry

fix(relay): skip retries for bad response body errors
Keep the model pricing editor wording aligned with the new price-based UI while exposing cache, image, and audio pricing in the marketplace so users can see the full configured pricing model.
Introduce a billing display mode feature allowing users to toggle between price and ratio views. Update relevant components and hooks to support this new functionality, ensuring consistent pricing information is displayed across the application.
Add siteDisplayType prop across various pricing components to conditionally render pricing information based on the selected display type. This update enhances the user experience by ensuring that pricing details are accurately represented according to the chosen display mode, particularly for token-based views.
为渠道参数覆盖可视化规则提供拖拽排序支持
…4f8a4248b0ab3b03ba703796ea3

fix: kling risk fail return openAIVideo error
…ride-beta-header-append

feat:support $keep_only_declared and deduped $append for header override
chore: update model lists for frequently used channels
@bkidy

bkidy commented Mar 12, 2026

Copy link
Copy Markdown
Author

I hit this in production with OpenAI Responses routed through a partially compatible relay.
The relay can serve /v1/responses, but replaying account/session-bound reasoning items across turns may fail with errors like:

  • Item with id 'rs_...' not found
  • invalid_encrypted_content
    This PR keeps default behavior unchanged and adds a small compatibility escape hatch:
  • disabled by default
  • enabled only via REMOVE_RESPONSES_REASONING_INPUT=true
  • strips reasoning* items from the Responses input array before forwarding upstream
    I kept the change intentionally small to avoid affecting the normal OpenAI Responses path.

@coderabbitai

coderabbitai Bot commented Mar 12, 2026

Copy link
Copy Markdown
Contributor

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 7acb97f1-4389-4439-8db3-99b22a806da9

📥 Commits

Reviewing files that changed from the base of the PR and between 4e1b05e and 86494be.

📒 Files selected for processing (4)
  • common/init.go
  • constant/env.go
  • dto/openai_request.go
  • relay/channel/openai/adaptor.go

Walkthrough

The pull request introduces an optional feature to remove reasoning-related data from OpenAI responses requests. It adds an environment variable configuration flag, a JSON filtering method, and integrates the filtering into the request processing pipeline.

Changes

Cohort / File(s) Summary
Configuration & Feature Flag
common/init.go, constant/env.go
Added environment variable REMOVE_RESPONSES_REASONING_INPUT (defaults to false) and initialized it as a package-level boolean flag.
Request Processing
dto/openai_request.go, relay/channel/openai/adaptor.go
Implemented RemoveReasoningFromInput() method to filter JSON array items by type and integrated conditional call in adaptor with error handling when flag is enabled.

Sequence Diagram

sequenceDiagram
    participant Env as Environment
    participant Init as Initialization
    participant Flag as constant.RemoveResponsesReasoningInput
    participant Adaptor as OpenAI Adaptor
    participant Request as OpenAIResponsesRequest
    participant DTO as Request DTO

    Env->>Init: Read REMOVE_RESPONSES_REASONING_INPUT
    Init->>Flag: Set boolean value
    Adaptor->>Flag: Check if enabled
    alt Flag Enabled
        Adaptor->>Request: Call RemoveReasoningFromInput()
        Request->>DTO: Parse Input as JSON array
        DTO->>DTO: Filter items (remove reasoning type)
        DTO->>Request: Assign filtered array back to Input
        Request->>Adaptor: Return result
    else Flag Disabled
        Adaptor->>Adaptor: Continue without filtering
    end
Loading

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~15 minutes

Possibly related PRs

  • Fix reasoning adaptor for openrouter #1577: Modifies reasoning-related field handling in the same adaptor file (relay/channel/openai/adaptor.go), involving reasoning data transformation for different API providers.

Poem

🐰 A filter whispers through the streams,
Removing reasoning from dreams,
With flags and JSON carefully dressed,
The responses flow, now cleaned and blessed! ✨

🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately describes the main change: adding an environment-driven compatibility flag (REMOVE_RESPONSES_REASONING_INPUT) for OpenAI Responses request handling to filter reasoning items.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
📝 Coding Plan
  • Generate coding plan for human review comments

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

Tip

You can validate your CodeRabbit configuration file in your editor.

If your editor has YAML language server, you can enable auto-completion and validation by adding # yaml-language-server: $schema=https://coderabbit.ai/integrations/schema.v2.json at the top of your CodeRabbit configuration file.

@RedwindA

Copy link
Copy Markdown
Contributor

建议研究并使用渠道亲和性功能

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.