Conversation
calculateCostFromTokens charged reasoning_tokens at the full reasoning rate ON TOP of completion_tokens: cost += completion_tokens * output_rate cost += reasoning_tokens * (reasoning_rate || output_rate) But reasoning tokens are a subset of completion tokens everywhere we consume them: OpenAI counts reasoning_tokens inside completion_tokens, and toOpenAIUsage's gemini extractor folds thoughtsTokenCount into completionTokens. Every reasoning request was therefore billed for its thinking twice — and since MODEL_PRICING sets reasoning == output for most entries, that is a straight 2x on the thinking portion. This is the same contract the function already applies one block up, where cached/cache_creation are subtracted because prompt_tokens is cache-inclusive. Reasoning now bills only the differential, and only when a model prices it apart from output. Two call sites reported gemini usage with thoughts OUTSIDE the completion count, which would have turned the fix into an under-charge for gemini. Both now fold thoughts in, matching what toOpenAIUsage already did. Total gemini cost is unchanged by the fold — candidates*output + thoughts*reasoning either way — only the field split changes; there is a test pinning that equivalence. Reported cost for reasoning models goes DOWN after this change. That is the point: the previous figures over-counted. 由 Claude Code 辅助生成
korvin2000
pushed a commit
to korvin2000/OmniRoute
that referenced
this pull request
Aug 2, 2026
Co-authored-by: luoyide <ydhome.code@gmail.com> Inspired-by: decolua/9router#2762
afandiaziz
pushed a commit
to afandiaziz/9router
that referenced
this pull request
Aug 9, 2026
29 PR upstream di-cherry-pick (semua masih open upstream per 2026-08-09). Rincian lengkap + link per PR ada di FORK-CHANGES.md. P1 skala 2475 koneksi : decolua#2798 decolua#410 decolua#2879 decolua#879 decolua#2997 P2 akurasi token/usage: decolua#2422 decolua#2658 decolua#2762 decolua#2453 decolua#2668 decolua#2361 P3 provider & combo : decolua#2526 decolua#3125 decolua#1434 decolua#2689 decolua#2439 decolua#2724 decolua#2647 decolua#1805 decolua#2909 decolua#2853 decolua#2508 decolua#2928 decolua#2345 decolua#2112 decolua#2786 P4 keamanan : decolua#1666 decolua#2776 Revert decolua#664: menambah transformRequest kedua di DefaultExecutor sehingga menimpa yang pertama dan mematikan stream_options/text.format/ injectReasoningContent/stripUnsupportedParams — termasuk PR decolua#3081 yang sudah dipakai produksi. Test: 88 gagal / 1783 lulus — nol regresi vs baseline v0.5.50 (88/1656).
golamrabbi696
added a commit
to golamrabbi696/EzRouter
that referenced
this pull request
Aug 12, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
calculateCostFromTokenscharges reasoning tokens on top of the completion tokens that already contain them.Reasoning tokens are a subset of completion tokens everywhere they are consumed:
reasoning_tokensinsidecompletion_tokens(spec), andextractUsage's OpenAI branch reads it straight out ofcompletion_tokens_details.toOpenAIUsage's gemini extractor returnscompletionTokens: candidates + thoughtsandreasoningTokens: thoughts— already folded.So every reasoning request was billed for its thinking twice. Because
MODEL_PRICINGsetsreasoningequal tooutputfor most entries (claude-sonnet-4-6: output 15.00, reasoning 15.00,gpt-5: output 10.00, reasoning 10.00, …),pricing.reasoning || pricing.outputresolves to the output rate and the second line is a straight 2x on the thinking portion. On a thinking-heavy request that is a ~30% overstatement of the whole call.The fix
Bill only the differential, and only when a model actually prices reasoning apart from output:
This is the same contract the function already applies one block up, where
cachedandcache_creationare subtracted fromprompt_tokensprecisely because that field is cache-inclusive. The comment there already spells it out: "prompt_tokens is cache-inclusive (see canonicalizeUsage): cached + cache_creation are subsets, so subtract both to avoid charging them at the full input rate." Completion tokens deserve the same treatment.Why two gemini call sites change too
extractUsage(usageTracking.js) andextractUsageFromResponse(requestDetail.js) both reported gemini usage withthoughtsTokenCountoutsidecandidatesTokenCount. Left alone, the pricing change would have turned into an under-charge for gemini. Both now fold thoughts in, matching whattoOpenAIUsagehas always done — so the reasoning-inclusive invariant holds on every path that reaches the cost function.Total gemini cost is unchanged by the fold:
candidates*output + thoughts*reasoningeither way. Only the field split changes. There is a test pinning that equivalence explicitly.Heads-up for reviewers
Reported cost for reasoning models goes down after this change. That is the intended outcome — the previous figures over-counted — but it will be visible in the dashboard, so it should not come as a surprise.
Verification
toOpenAIUsage), and the cost math for reasoning priced at / above / absent from the output rate, plus the gemini cost-equivalence guard.masterbaseline, zero regression.本 PR 由 Claude Code 辅助完成