Repository navigation
Conversation
…eam_options (Twigpine#448) Root cause: MiMo and Gitlawb OpenGateway explicitly strip `stream_options` from request bodies (it causes a "Param Incorrect" error on MiMo's API). Without `stream_options: { include_usage: true }`, the provider never includes usage in its streaming response. The shim seeds each message with { input_tokens: 0, output_tokens: 0 } and those values never get updated, so the status-line always shows 0% context used. Fix 1 — tokens.ts: `getTokenUsage` now returns `undefined` when both `input_tokens` and `output_tokens` are zero. A real API response always has ≥ 1 output token, so all-zero usage unambiguously means "provider did not report usage". Returning `undefined` causes `getCurrentUsage()` to return `null` → `calculateContextPercentages(null, ...)` returns `{ used: null }` → status-line shows "N/A" (empty) instead of misleading "0% used". Fix 2 — openaiShim.ts: Accumulate `lastSeenUsage` across all SSE chunks. Some providers emit usage in an early chunk (before the finish_reason chunk); without this accumulator the `message_delta` is emitted without usage even when the stream contained real usage data earlier. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
jatmn
left a comment
There was a problem hiding this comment.
Thanks for the contribution. I took a pass through the linked issue, the status-line path, and the OpenAI shim changes, and found one issue that needs to be addressed.
Findings
- [P2] Do not turn the linked status-line bug into a silent no-usage state
src/utils/tokens.ts:39
For the MiMo/OpenGateway repro in #448, this still leaves the status-line input without token usage after completed turns: the stream path reports zero usage, the session totals incontext_window.total_input_tokens/total_output_tokensremain zero, and this new guard changescurrent_usageplusused_percentagetonull, which is the same shape documented for "no messages yet". That hides the misleading0%display, but it does not make the status line report context usage for the provider/path-specific bug the issue is tracking, so commonjq '.context_window.used_percentage // empty'configurations now print nothing after real model turns. Please either route a real/estimated usage value into this path, or narrow the PR so it no longer presents this as fixing #448 and make the intended "unsupported usage" state explicit in the status-line contract and tests.
…P turn When a user switches providers mid-session (Anthropic → MiMo), the most recent assistant messages have all-zero usage (provider doesn't report it). Previously, getCurrentUsage() would skip those zero-usage messages and return the last Anthropic message's numbers — surfacing stale context % from a previous API call instead of N/A. Fix: in the getCurrentUsage loop, when we encounter a real (non-synthetic) assistant message that getTokenUsage() skipped because usage is all-zero, stop iteration immediately and return null rather than walking further back through the message history. Also add: - Tests for getCurrentUsage mixed-session and empty-session behavior - Shim test asserting hasEmittedFinalUsage prevents a duplicate message_delta when a trailing empty-choices chunk arrives after lastSeenUsage already triggered the first message_delta Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
jatmn
left a comment
There was a problem hiding this comment.
Thanks for the update. I took another pass through the changed paths and found an issue that still needs to be addressed before this is ready.
Findings
- [P2] Stop all-zero usage from falling back in every token helper
src/utils/tokens.ts:39
This guard makesgetTokenUsage()hide all-zero assistant usage for every caller, but onlygetCurrentUsage()gets the new "stop at unsupported usage" behavior. The other shared helpers still keep walking backward, so a mixed session can still reuse stale Anthropic usage after a MiMo/OpenGateway turn. For example, with an older assistant message at{ input_tokens: 1000, output_tokens: 200 }followed by a real MiMo assistant message at{ input_tokens: 0, output_tokens: 0 },tokenCountFromLastAPIResponse()andfinalContextTokensFromLastResponse()now return1200,messageTokenCountFromLastAPIResponse()returns200, andtokenCountWithEstimation()anchors on the old response plus later-message estimates. Those helpers feed prompt warnings, session-memory/compact metadata, task-budget accounting, and the auto-compact/blocking-limit path, so the stale-usage problem this PR fixes for the status line remains in adjacent runtime decisions. Please either keep all-zero usage visible to callers that need to distinguish "unsupported latest response" from "no usage found", or give the shared search helpers the same stop-at-current-unsupported-turn behavior and add mixed-provider coverage for the non-status-line callers too.
getTokenUsage returning undefined for all-zero usage broke every shared helper that walks messages backward: tokenCountFromLastAPIResponse, finalContextTokensFromLastResponse, messageTokenCountFromLastAPIResponse, and tokenCountWithEstimation would skip past a MiMo/OpenGateway turn (zero usage) and find stale Anthropic data from the previous provider. This restores getTokenUsage to its original form — it always returns the raw usage object. The all-zero sentinel check is now inlined into getCurrentUsage alone (via getTokenCountFromUsage(usage) === 0), which is the only caller where "provider didn't report usage → show N/A" is the correct behavior. All other callers see the same zero values they saw before this PR, which is harmless and not a regression. The removed case-b fallback block in getCurrentUsage was dead code after this change: getTokenUsage no longer returns undefined for real assistant messages, so the inline message-type guard was never reached. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
|
Addressed in the latest commit ( What changed: The The all-zero check is now inlined exclusively in Side effect cleanup: the old "case b" fallback block in Tests updated:
|
jatmn
left a comment
There was a problem hiding this comment.
Thanks for the update. I took another pass through the changed paths and found an issue that still needs to be addressed before this is ready.
Findings
- [P2] Do not let early usage suppress the later final usage chunk
src/services/api/openaiShim.ts:1770
lastSeenUsageis now emitted on thefinish_reasonchunk even when that chunk has nousage, and that setshasEmittedFinalUsage = true. If a provider sends an earlier usage-bearing chunk, then the stop chunk, then the existing trailing empty-choices usage chunk with the final cumulative token counts, the final chunk is ignored becausehasEmittedFinalUsageis already true. That regresses the empty-choices fallback path this file already had for providers that report final stream usage after the stop chunk, and the new duplicate test only covers the case where the trailing usage is identical to the early usage. Please keep the later empty-choices usage able to replace or emit the final counts, or avoid marking final usage emitted when the stop event only used carried-forwardlastSeenUsage.
…real usage When a stop chunk arrives without chunkUsage and falls back to lastSeenUsage, hasEmittedFinalUsage must NOT be set — a subsequent trailing empty-choices chunk carrying the definitive final token counts should still be allowed to fire. Previously `if (finalUsage)` used the fallback value, preventing the trailing chunk from emitting. Changed to `if (chunkUsage)` so only a stop chunk with real provider-supplied usage marks the flag done. Also updates the existing duplicate-emission test to reflect the corrected behavior and adds two new tests covering the regression scenario. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
|
Fixed in latest commit. Changed Three tests replaced/added in
|
jatmn
left a comment
There was a problem hiding this comment.
Thanks for the update. I rechecked the changed paths and found an issue that still needs to be addressed.
Findings
- [P2] Avoid double-counting usage when a trailing final chunk arrives
src/services/api/openaiShim.ts:1770
The stop chunk now emitslastSeenUsageas amessage_deltawhen it has nochunkUsage, and the trailing empty-choices path can then emit another usage-bearingmessage_deltafor the same API response. The new regression tests even cover this as two deltas, butclaude.tstreats every usage-bearingmessage_deltaas additive by callingaddToTotalSessionCost, which incrementsSTATE.modelUsagetotals incost-tracker.ts. For providers that send early/provisional usage followed by a final empty-choices usage chunk, this will add both the provisional counts and the final counts to the status-line totals/cost instead of replacing the first value with the final one. The same condition also replays stalelastSeenUsageon a trailing empty-choices chunk that has nousageof its own, so an otherwise harmless empty trailer can duplicate accounting too. Please make the carried-forward stop usage non-additive, or otherwise ensure the later final usage supersedes it without charging/updating the global totals twice.
…arly provisional + trailing final chunks Providers like MiMo/OpenGateway emit an early chunk with provisional usage data, then a trailing empty-choices chunk with the real final totals. Previously the stop-chunk handler fell back to `lastSeenUsage` when `chunkUsage` was absent, which emitted the provisional counts at stop time; the in-loop post-check then fired again on the trailing chunk and emitted the real final counts -- both flowing into addToTotalSessionCost and doubling the session totals. Fix: - At stop time, only attach usage when the stop chunk itself carries it. - Add a post-stream fallback (after the reader is released) that emits lastSeenUsage once for providers that send usage only in an early chunk with no trailing empty-choices chunk. The existing in-loop trailing-chunk check already handles the trailing-chunk case; the post-stream path closes the gap for providers that omit it. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
|
Thanks for the precise diagnosis, @jatmn. Root cause confirmed: when a provider emits usage in an early/provisional chunk and the stop chunk carries no usage, the stop-chunk handler fell back to Fix (latest commit
The three cases now behave correctly:
|
jatmn
left a comment
There was a problem hiding this comment.
Thanks for the update. I rechecked the changed paths and found one remaining issue that needs to be addressed.
Findings
- [P2] Update the status-line contract for unsupported usage
src/tools/AgentTool/built-in/statuslineSetup.ts:58
getCurrentUsage()now returnsnullafter a completed assistant turn when the latest provider returned the shimmed all-zero usage record, andStatusLine.tsxpasses that through tocontext_window.current_usage,used_percentage, andremaining_percentage. The built-in status-line setup prompt still documents thosenullvalues as meaning only "no messages yet", so/statusline setupwill teach users and generated status-line commands the wrong contract for the new MiMo/OpenGateway unsupported-usage state. Please update this status-line schema/help text, or expose a distinguishable unsupported-usage state, so the user-facing contract matches the behavior this PR introduces.
…tate getCurrentUsage() returns null in two distinct cases: 1. No assistant messages yet (existing behavior, already documented) 2. Active provider does not report token usage — e.g. providers that strip stream_options (MiMo, OpenGateway) return all-zero usage records, which getCurrentUsage() treats as unsupported and returns null for. The /statusline setup prompt previously documented only case 1, so generated status-line commands would mislead users into thinking null only means "session not started". Update current_usage, used_percentage, and remaining_percentage inline docs to call out both null triggers. The existing // empty jq fallback in the examples already handles both cases correctly — no example changes needed. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
|
Fixed in commit Updated the three field comments in The |
jatmn
left a comment
There was a problem hiding this comment.
Thanks for the update. I rechecked the changed paths and found one issue that still needs to be addressed.
Findings
- [P2] Do not finalize usage from empty chunks that carry no usage
src/services/api/openaiShim.ts:1788
The trailing empty-choices branch now falls back tolastSeenUsageeven when the empty chunk itself has nousage, and then setshasEmittedFinalUsage = true. If a provider sends early/provisional usage, the stop chunk, an empty keepalive/trailer with no usage, and then the real final empty-choices usage chunk, the no-usage chunk will replay the provisional counts and suppress the later final totals.claude.tstreats every usage-bearingmessage_deltaas additive session accounting, so this leaves status-line totals and costs based on the stale provisional usage instead of the definitive final chunk. Please only let this in-loop empty-choices path finalize when the current chunk has real usage, and leave thelastSeenUsagefallback to the post-stream path where no later final chunk can still arrive.
|
closing as abandoned |
Root cause
Issue #448 (status line always shows 0% context used) still reproduces on current main for MiMo V2.5 Pro via Gitlawb OpenGateway, as confirmed by @TheGoddessInari.
The root cause is a two-step silent failure:
stream_optionsis stripped — bothgitlawb-opengateway.tsandxiaomi-mimo.ts(andruntimeMetadata.tsfor allmimo-v2*models) include'stream_options'inremoveBodyFields. This was intentionally added in fix: MiMo remove unsupported body fields and preserve reasoning content #1253 because MiMo's API returns a "Param Incorrect" error whenstream_optionsis present. The stripping is correct; the missing fallback is the bug.No usage → all-zero message — without
stream_options: { include_usage: true }, the provider never includes usage in its streaming response. The shim always seedsmessage_start.message.usagewith{ input_tokens: 0, output_tokens: 0 }. Since no real usage ever arrives viamessage_delta, those zeros persist.getCurrentUsage(messages)reads the last assistant message's usage, gets{ 0, 0, 0, 0 }, andcalculateContextPercentagescomputes0%.Changes
Fix 1 —
src/utils/tokens.ts:getTokenUsagezero guardReturns
undefinedwhen bothinput_tokensandoutput_tokensare zero. A real API response always has ≥ 1 output token, so an all-zero record is an unambiguous signal that the provider did not report usage.Fix 2 —
src/utils/tokens.ts:getCurrentUsagemixed-session guardPreviously, when
getTokenUsageskipped a zero-usage message, the loop fell back to older messages. In a mixed session (started on Anthropic, then switched to MiMo), this would surface stale context numbers from the previous Anthropic turn instead of N/A.Now: when the loop finds a real (non-synthetic) assistant message that
getTokenUsageskipped because usage is all-zero, it stops immediately and returnsnullrather than walking backwards into stale history. Non-assistant messages (user, tool_result) and synthetic messages still trigger the normal fallback walk.Together, Fix 1 + Fix 2 mean:
getCurrentUsage= null →used_percentage= null → status bar: N/AFix 3 —
src/services/api/openaiShim.ts:lastSeenUsageaccumulatorAdds a
lastSeenUsageaccumulator so usage arriving in an early SSE chunk carries forward to themessage_deltaat stop time. The existing empty-choices fallback path is also updated to uselastSeenUsage.Tests
tokens.test.ts—getTokenUsage: non-zero, all-zero, partialtokens.test.ts—getCurrentUsage: pure session, zero session, mixed-session, no messagesopenaiShim.test.ts— early-chunk usage forwarded tomessage_delta; trailing empty-choices chunk does not emit a duplicatemessage_deltaVerification
Known remaining limitation
context_window.total_input_tokens/total_output_tokensin the status-line JSON (sourced fromSTATE.modelUsageviaaddToTotalSessionCost) still show0for MiMo sessions — the write path inclaude.tsis not changed here. Getting those fields accurate would require either client-side token estimation or MiMo adding support forstream_options. Theused_percentagefield (the one visible in a typical status bar) is correctly N/A.🤖 Generated with Claude Code