Skip to content

feat(usage): surface Claude thinking token counts to clients - #2770

Open
yidecode wants to merge 2 commits into
decolua:masterfrom
yidecode:pr/claude-thinking-token-visibility
Open

yidecode wants to merge 2 commits into
decolua:masterfrom
yidecode:pr/claude-thinking-token-visibility

Conversation

@yidecode

Copy link
Copy Markdown
Contributor

Stacked on #2762 — that one has to land first. Adding reasoning_tokens for Claude while calculateCostFromTokens still adds reasoning on top of completion would make Claude join the double-billing #2762 removes. This branch contains #2762's commit; it will rebase to a single commit once that merges.

Anthropic reports thinking tokens on message_delta as usage.output_tokens_details.thinking_tokens. Nothing consumed it — the Claude branches of extractUsage, extractUsageFromResponse and toOpenAIUsage's claude extractor all stopped at input/output/cache — so the count never reached cost tracking, the request-detail record, or the client.

The non-obvious part

Extracting the count is not enough. claude-to-openai builds its own usage object and assigns it to the shared stream state, mixing Claude field names into it:

state.usage = { prompt_tokens, completion_tokens, total_tokens,
                input_tokens, output_tokens };
state.usage.output_tokens_details = { thinking_tokens };

stream.js then hands that object to filterUsageForFormat before emitting the final chunk — and the OpenAI allow-list has no output_tokens_details, because that is a Claude name. So a naive fix extracts the count correctly and still drops it on the way out. The stored request-detail record shows the leak exactly: a hybrid {prompt_tokens, …, input_tokens, output_tokens, output_tokens_details} object that no single format's allow-list fully matches.

The count is therefore mirrored under both OpenAI spellings (reasoning_tokens and completion_tokens_details.reasoning_tokens), and output_tokens_details is added to the Claude allow-list so Claude-format clients keep it too.

completion_tokens is deliberately left alone — thinking tokens are already inside output_tokens, so the reasoning-inclusive convention #2762 relies on still holds.

Why it matters

GitHub Copilot returns a signed but empty thinking block for its 4.7+ Claude shims (claude-sonnet-5, claude-opus-4.7, claude-opus-4.8, claude-fable-5) — the reasoning happens, the text is withheld server-side. Verified end-to-end against a live gateway:

sonnet-4.6   reasoning_content=1012ch   reasoning_tokens=426
sonnet-5     reasoning_content=   0ch   reasoning_tokens=886

Without the count, that second row is indistinguishable from a broken pipeline — which is exactly how it gets reported. With it, a client can tell "withheld upstream" from "the proxy ate my thinking".

Verification

  • 4 new tests driving the real createSSETransformStreamWithLogger(CLAUDE → OPENAI) transform, including the withheld-text case and the Claude-format allow-list.
  • Full suite: failure set identical to the master baseline, zero regression.

本 PR 由 Claude Code 辅助完成

yidecode added 2 commits July 22, 2026 12:20
calculateCostFromTokens charged reasoning_tokens at the full reasoning rate
ON TOP of completion_tokens:

  cost += completion_tokens * output_rate
  cost += reasoning_tokens  * (reasoning_rate || output_rate)

But reasoning tokens are a subset of completion tokens everywhere we consume
them: OpenAI counts reasoning_tokens inside completion_tokens, and
toOpenAIUsage's gemini extractor folds thoughtsTokenCount into
completionTokens. Every reasoning request was therefore billed for its
thinking twice — and since MODEL_PRICING sets reasoning == output for most
entries, that is a straight 2x on the thinking portion.

This is the same contract the function already applies one block up, where
cached/cache_creation are subtracted because prompt_tokens is cache-inclusive.
Reasoning now bills only the differential, and only when a model prices it
apart from output.

Two call sites reported gemini usage with thoughts OUTSIDE the completion
count, which would have turned the fix into an under-charge for gemini. Both
now fold thoughts in, matching what toOpenAIUsage already did. Total gemini
cost is unchanged by the fold — candidates*output + thoughts*reasoning either
way — only the field split changes; there is a test pinning that equivalence.

Reported cost for reasoning models goes DOWN after this change. That is the
point: the previous figures over-counted.

由 Claude Code 辅助生成
Stacked on the reasoning-token double-billing fix — that one has to land first,
otherwise adding reasoning_tokens for Claude makes Claude join the double-billing
it removes.

Anthropic reports thinking tokens on message_delta as
usage.output_tokens_details.thinking_tokens. Nothing consumed it: the Claude
branches of extractUsage / extractUsageFromResponse and toOpenAIUsage's claude
extractor all stopped at input/output/cache, so the count never reached cost
tracking, the request-detail record, or the client.

Wiring it up needed one non-obvious step. claude-to-openai builds its own usage
object and assigns it to the shared stream state, mixing Claude field names in:

  state.usage = { prompt_tokens, completion_tokens, total_tokens,
                  input_tokens, output_tokens }
  state.usage.output_tokens_details = { thinking_tokens }

stream.js then passes that object through filterUsageForFormat before emitting
the final chunk, and the OpenAI allow-list has no `output_tokens_details` — so a
naive fix extracts the count correctly and still drops it on the way out. The
count is therefore mirrored under both OpenAI spellings (`reasoning_tokens` and
`completion_tokens_details.reasoning_tokens`), and `output_tokens_details` is
added to the Claude allow-list so Claude-format clients keep it too.

Note that thinking tokens are already inside output_tokens, so completion_tokens
is left alone — the reasoning-inclusive convention the pricing fix relies on
still holds.

This matters most where the thinking TEXT is unavailable: GitHub Copilot returns
a signed but EMPTY thinking block for its 4.7+ Claude shims (sonnet-5, opus-4.7,
opus-4.8, fable-5). Verified end-to-end against a live gateway:

  sonnet-4.6  reasoning_content=1012ch  reasoning_tokens=426
  sonnet-5    reasoning_content=   0ch  reasoning_tokens=886

Without the count, that second row is indistinguishable from a broken pipeline.

由 Claude Code 辅助生成
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant