Skip to content

fix(pricing): stop billing reasoning tokens twice - #2762

Open
yidecode wants to merge 1 commit into
decolua:masterfrom
yidecode:pr/reasoning-token-double-billing
Open

yidecode wants to merge 1 commit into
decolua:masterfrom
yidecode:pr/reasoning-token-double-billing

Conversation

@yidecode

Copy link
Copy Markdown
Contributor

calculateCostFromTokens charges reasoning tokens on top of the completion tokens that already contain them.

const outputTokens = tokens.completion_tokens || tokens.output_tokens || 0;
cost += outputTokens * (pricing.output / 1000000);

const reasoningTokens = tokens.reasoning_tokens || 0;
if (reasoningTokens > 0) {
  cost += reasoningTokens * ((pricing.reasoning || pricing.output) / 1000000);   // added, not substituted
}

Reasoning tokens are a subset of completion tokens everywhere they are consumed:

  • OpenAI counts reasoning_tokens inside completion_tokens (spec), and extractUsage's OpenAI branch reads it straight out of completion_tokens_details.
  • toOpenAIUsage's gemini extractor returns completionTokens: candidates + thoughts and reasoningTokens: thoughts — already folded.

So every reasoning request was billed for its thinking twice. Because MODEL_PRICING sets reasoning equal to output for most entries (claude-sonnet-4-6: output 15.00, reasoning 15.00, gpt-5: output 10.00, reasoning 10.00, …), pricing.reasoning || pricing.output resolves to the output rate and the second line is a straight 2x on the thinking portion. On a thinking-heavy request that is a ~30% overstatement of the whole call.

The fix

Bill only the differential, and only when a model actually prices reasoning apart from output:

if (reasoningTokens > 0 && pricing.reasoning) {
  cost += reasoningTokens * ((pricing.reasoning - pricing.output) / 1000000);
}

This is the same contract the function already applies one block up, where cached and cache_creation are subtracted from prompt_tokens precisely because that field is cache-inclusive. The comment there already spells it out: "prompt_tokens is cache-inclusive (see canonicalizeUsage): cached + cache_creation are subsets, so subtract both to avoid charging them at the full input rate." Completion tokens deserve the same treatment.

Why two gemini call sites change too

extractUsage (usageTracking.js) and extractUsageFromResponse (requestDetail.js) both reported gemini usage with thoughtsTokenCount outside candidatesTokenCount. Left alone, the pricing change would have turned into an under-charge for gemini. Both now fold thoughts in, matching what toOpenAIUsage has always done — so the reasoning-inclusive invariant holds on every path that reaches the cost function.

Total gemini cost is unchanged by the fold: candidates*output + thoughts*reasoning either way. Only the field split changes. There is a test pinning that equivalence explicitly.

Heads-up for reviewers

Reported cost for reasoning models goes down after this change. That is the intended outcome — the previous figures over-counted — but it will be visible in the dashboard, so it should not come as a surprise.

Verification

  • 9 new tests: the reasoning-inclusive convention (gemini fold, OpenAI passthrough, agreement with toOpenAIUsage), and the cost math for reasoning priced at / above / absent from the output rate, plus the gemini cost-equivalence guard.
  • Full suite: failure set identical to the master baseline, zero regression.

本 PR 由 Claude Code 辅助完成

calculateCostFromTokens charged reasoning_tokens at the full reasoning rate
ON TOP of completion_tokens:

  cost += completion_tokens * output_rate
  cost += reasoning_tokens  * (reasoning_rate || output_rate)

But reasoning tokens are a subset of completion tokens everywhere we consume
them: OpenAI counts reasoning_tokens inside completion_tokens, and
toOpenAIUsage's gemini extractor folds thoughtsTokenCount into
completionTokens. Every reasoning request was therefore billed for its
thinking twice — and since MODEL_PRICING sets reasoning == output for most
entries, that is a straight 2x on the thinking portion.

This is the same contract the function already applies one block up, where
cached/cache_creation are subtracted because prompt_tokens is cache-inclusive.
Reasoning now bills only the differential, and only when a model prices it
apart from output.

Two call sites reported gemini usage with thoughts OUTSIDE the completion
count, which would have turned the fix into an under-charge for gemini. Both
now fold thoughts in, matching what toOpenAIUsage already did. Total gemini
cost is unchanged by the fold — candidates*output + thoughts*reasoning either
way — only the field split changes; there is a test pinning that equivalence.

Reported cost for reasoning models goes DOWN after this change. That is the
point: the previous figures over-counted.

由 Claude Code 辅助生成
korvin2000 pushed a commit to korvin2000/OmniRoute that referenced this pull request Aug 2, 2026
Co-authored-by: luoyide <ydhome.code@gmail.com>
Inspired-by: decolua/9router#2762
afandiaziz pushed a commit to afandiaziz/9router that referenced this pull request Aug 9, 2026
29 PR upstream di-cherry-pick (semua masih open upstream per 2026-08-09).
Rincian lengkap + link per PR ada di FORK-CHANGES.md.

P1 skala 2475 koneksi : decolua#2798 decolua#410 decolua#2879 decolua#879 decolua#2997
P2 akurasi token/usage: decolua#2422 decolua#2658 decolua#2762 decolua#2453 decolua#2668 decolua#2361
P3 provider & combo   : decolua#2526 decolua#3125 decolua#1434 decolua#2689 decolua#2439 decolua#2724 decolua#2647 decolua#1805
                        decolua#2909 decolua#2853 decolua#2508 decolua#2928 decolua#2345 decolua#2112 decolua#2786
P4 keamanan           : decolua#1666 decolua#2776

Revert decolua#664: menambah transformRequest kedua di DefaultExecutor sehingga
menimpa yang pertama dan mematikan stream_options/text.format/
injectReasoningContent/stripUnsupportedParams — termasuk PR decolua#3081 yang
sudah dipakai produksi.

Test: 88 gagal / 1783 lulus — nol regresi vs baseline v0.5.50 (88/1656).
golamrabbi696 added a commit to golamrabbi696/EzRouter that referenced this pull request Aug 12, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant