Skip to content

fix(status-line): show N/A instead of 0% for providers that strip stream_options (#448) - #1444

Closed
ajsai47 wants to merge 6 commits into
Twigpine:mainfrom
ajsai47:fix/status-line-zero-usage-3p-providers
Closed

ajsai47 wants to merge 6 commits into
Twigpine:mainfrom
ajsai47:fix/status-line-zero-usage-3p-providers

Conversation

@ajsai47

@ajsai47 ajsai47 commented May 30, 2026 •

Copy link
Copy Markdown

Root cause

Issue #448 (status line always shows 0% context used) still reproduces on current main for MiMo V2.5 Pro via Gitlawb OpenGateway, as confirmed by @TheGoddessInari.

The root cause is a two-step silent failure:

  1. stream_options is stripped — both gitlawb-opengateway.ts and xiaomi-mimo.ts (and runtimeMetadata.ts for all mimo-v2* models) include 'stream_options' in removeBodyFields. This was intentionally added in fix: MiMo remove unsupported body fields and preserve reasoning content #1253 because MiMo's API returns a "Param Incorrect" error when stream_options is present. The stripping is correct; the missing fallback is the bug.

  2. No usage → all-zero message — without stream_options: { include_usage: true }, the provider never includes usage in its streaming response. The shim always seeds message_start.message.usage with { input_tokens: 0, output_tokens: 0 }. Since no real usage ever arrives via message_delta, those zeros persist. getCurrentUsage(messages) reads the last assistant message's usage, gets { 0, 0, 0, 0 }, and calculateContextPercentages computes 0%.

Changes

Fix 1 — src/utils/tokens.ts: getTokenUsage zero guard

Returns undefined when both input_tokens and output_tokens are zero. A real API response always has ≥ 1 output token, so an all-zero record is an unambiguous signal that the provider did not report usage.

Fix 2 — src/utils/tokens.ts: getCurrentUsage mixed-session guard

Previously, when getTokenUsage skipped a zero-usage message, the loop fell back to older messages. In a mixed session (started on Anthropic, then switched to MiMo), this would surface stale context numbers from the previous Anthropic turn instead of N/A.

Now: when the loop finds a real (non-synthetic) assistant message that getTokenUsage skipped because usage is all-zero, it stops immediately and returns null rather than walking backwards into stale history. Non-assistant messages (user, tool_result) and synthetic messages still trigger the normal fallback walk.

Together, Fix 1 + Fix 2 mean:

  • Pure MiMo session → getCurrentUsage = null → used_percentage = null → status bar: N/A
  • Mixed session → same (stops at the most recent all-zero MiMo message)
  • Anthropic session → unchanged

Fix 3 — src/services/api/openaiShim.ts: lastSeenUsage accumulator

Adds a lastSeenUsage accumulator so usage arriving in an early SSE chunk carries forward to the message_delta at stop time. The existing empty-choices fallback path is also updated to use lastSeenUsage.

Tests

  • 4 tests in tokens.test.ts — getTokenUsage: non-zero, all-zero, partial
  • 4 tests in tokens.test.ts — getCurrentUsage: pure session, zero session, mixed-session, no messages
  • 2 tests in openaiShim.test.ts — early-chunk usage forwarded to message_delta; trailing empty-choices chunk does not emit a duplicate message_delta

Verification

bun test src/utils/tokens.test.ts src/services/api/openaiShim.test.ts
# → 116 pass, 0 fail

bun run build
# → ✓ Built openclaude v0.15.0

Known remaining limitation

context_window.total_input_tokens / total_output_tokens in the status-line JSON (sourced from STATE.modelUsage via addToTotalSessionCost) still show 0 for MiMo sessions — the write path in claude.ts is not changed here. Getting those fields accurate would require either client-side token estimation or MiMo adding support for stream_options. The used_percentage field (the one visible in a typical status bar) is correctly N/A.

🤖 Generated with Claude Code

…eam_options (Twigpine#448)

Root cause: MiMo and Gitlawb OpenGateway explicitly strip `stream_options` from
request bodies (it causes a "Param Incorrect" error on MiMo's API). Without
`stream_options: { include_usage: true }`, the provider never includes usage in
its streaming response. The shim seeds each message with { input_tokens: 0,
output_tokens: 0 } and those values never get updated, so the status-line
always shows 0% context used.

Fix 1 — tokens.ts: `getTokenUsage` now returns `undefined` when both
`input_tokens` and `output_tokens` are zero. A real API response always has
≥ 1 output token, so all-zero usage unambiguously means "provider did not
report usage". Returning `undefined` causes `getCurrentUsage()` to return
`null` → `calculateContextPercentages(null, ...)` returns `{ used: null }` →
status-line shows "N/A" (empty) instead of misleading "0% used".

Fix 2 — openaiShim.ts: Accumulate `lastSeenUsage` across all SSE chunks.
Some providers emit usage in an early chunk (before the finish_reason chunk);
without this accumulator the `message_delta` is emitted without usage even
when the stream contained real usage data earlier.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
@jatmn jatmn closed this May 30, 2026
@jatmn jatmn reopened this May 30, 2026

@jatmn jatmn left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the contribution. I took a pass through the linked issue, the status-line path, and the OpenAI shim changes, and found one issue that needs to be addressed.

Findings

  • [P2] Do not turn the linked status-line bug into a silent no-usage state
    src/utils/tokens.ts:39
    For the MiMo/OpenGateway repro in #448, this still leaves the status-line input without token usage after completed turns: the stream path reports zero usage, the session totals in context_window.total_input_tokens / total_output_tokens remain zero, and this new guard changes current_usage plus used_percentage to null, which is the same shape documented for "no messages yet". That hides the misleading 0% display, but it does not make the status line report context usage for the provider/path-specific bug the issue is tracking, so common jq '.context_window.used_percentage // empty' configurations now print nothing after real model turns. Please either route a real/estimated usage value into this path, or narrow the PR so it no longer presents this as fixing #448 and make the intended "unsupported usage" state explicit in the status-line contract and tests.

…P turn

When a user switches providers mid-session (Anthropic → MiMo), the most
recent assistant messages have all-zero usage (provider doesn't report it).
Previously, getCurrentUsage() would skip those zero-usage messages and
return the last Anthropic message's numbers — surfacing stale context %
from a previous API call instead of N/A.

Fix: in the getCurrentUsage loop, when we encounter a real (non-synthetic)
assistant message that getTokenUsage() skipped because usage is all-zero,
stop iteration immediately and return null rather than walking further back
through the message history.

Also add:
- Tests for getCurrentUsage mixed-session and empty-session behavior
- Shim test asserting hasEmittedFinalUsage prevents a duplicate message_delta
  when a trailing empty-choices chunk arrives after lastSeenUsage already
  triggered the first message_delta

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

@jatmn jatmn left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the update. I took another pass through the changed paths and found an issue that still needs to be addressed before this is ready.

Findings

  • [P2] Stop all-zero usage from falling back in every token helper
    src/utils/tokens.ts:39
    This guard makes getTokenUsage() hide all-zero assistant usage for every caller, but only getCurrentUsage() gets the new "stop at unsupported usage" behavior. The other shared helpers still keep walking backward, so a mixed session can still reuse stale Anthropic usage after a MiMo/OpenGateway turn. For example, with an older assistant message at { input_tokens: 1000, output_tokens: 200 } followed by a real MiMo assistant message at { input_tokens: 0, output_tokens: 0 }, tokenCountFromLastAPIResponse() and finalContextTokensFromLastResponse() now return 1200, messageTokenCountFromLastAPIResponse() returns 200, and tokenCountWithEstimation() anchors on the old response plus later-message estimates. Those helpers feed prompt warnings, session-memory/compact metadata, task-budget accounting, and the auto-compact/blocking-limit path, so the stale-usage problem this PR fixes for the status line remains in adjacent runtime decisions. Please either keep all-zero usage visible to callers that need to distinguish "unsupported latest response" from "no usage found", or give the shared search helpers the same stop-at-current-unsupported-turn behavior and add mixed-provider coverage for the non-status-line callers too.

getTokenUsage returning undefined for all-zero usage broke every shared
helper that walks messages backward: tokenCountFromLastAPIResponse,
finalContextTokensFromLastResponse, messageTokenCountFromLastAPIResponse,
and tokenCountWithEstimation would skip past a MiMo/OpenGateway turn
(zero usage) and find stale Anthropic data from the previous provider.

This restores getTokenUsage to its original form — it always returns the
raw usage object. The all-zero sentinel check is now inlined into
getCurrentUsage alone (via getTokenCountFromUsage(usage) === 0), which
is the only caller where "provider didn't report usage → show N/A" is
the correct behavior. All other callers see the same zero values they
saw before this PR, which is harmless and not a regression.

The removed case-b fallback block in getCurrentUsage was dead code after
this change: getTokenUsage no longer returns undefined for real assistant
messages, so the inline message-type guard was never reached.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
@ajsai47

ajsai47 commented May 31, 2026

Copy link
Copy Markdown
Author

Addressed in the latest commit (bc5f2fe).

What changed:

The getTokenUsage zero-usage guard has been reverted. The function now returns the raw usage object for all real assistant messages, including all-zero ones. This restores pre-PR behaviour for every shared helper (tokenCountFromLastAPIResponse, finalContextTokensFromLastResponse, messageTokenCountFromLastAPIResponse, tokenCountWithEstimation) — they see zero values, not stale Anthropic data.

The all-zero check is now inlined exclusively in getCurrentUsage, using getTokenCountFromUsage(usage) === 0. This covers both the pure-3P case (status line shows N/A) and the mixed-session case (stops at the MiMo turn, does not fall back to the older Anthropic record).

Side effect cleanup: the old "case b" fallback block in getCurrentUsage is now dead code — getTokenUsage no longer returns undefined for real assistant messages, so the message-type guard after the if (usage) branch was never reachable. Removed it to avoid confusion.

Tests updated:

  • getTokenUsage zero-usage test: now asserts the zero-object is returned (not undefined)
  • getCurrentUsage tests: unchanged in assertion — null for pure-3P, null for mixed sessions

@jatmn jatmn left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the update. I took another pass through the changed paths and found an issue that still needs to be addressed before this is ready.

Findings

  • [P2] Do not let early usage suppress the later final usage chunk
    src/services/api/openaiShim.ts:1770
    lastSeenUsage is now emitted on the finish_reason chunk even when that chunk has no usage, and that sets hasEmittedFinalUsage = true. If a provider sends an earlier usage-bearing chunk, then the stop chunk, then the existing trailing empty-choices usage chunk with the final cumulative token counts, the final chunk is ignored because hasEmittedFinalUsage is already true. That regresses the empty-choices fallback path this file already had for providers that report final stream usage after the stop chunk, and the new duplicate test only covers the case where the trailing usage is identical to the early usage. Please keep the later empty-choices usage able to replace or emit the final counts, or avoid marking final usage emitted when the stop event only used carried-forward lastSeenUsage.

…real usage

When a stop chunk arrives without chunkUsage and falls back to lastSeenUsage,
hasEmittedFinalUsage must NOT be set — a subsequent trailing empty-choices chunk
carrying the definitive final token counts should still be allowed to fire.

Previously `if (finalUsage)` used the fallback value, preventing the trailing
chunk from emitting. Changed to `if (chunkUsage)` so only a stop chunk with
real provider-supplied usage marks the flag done.

Also updates the existing duplicate-emission test to reflect the corrected
behavior and adds two new tests covering the regression scenario.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
@ajsai47

ajsai47 commented May 31, 2026

Copy link
Copy Markdown
Author

Fixed in latest commit. Changed if (finalUsage) to if (chunkUsage) on the stop-chunk path — hasEmittedFinalUsage is now only set when the stop chunk itself carries real usage, not when we're using the carried-forward lastSeenUsage fallback. This lets a subsequent trailing empty-choices chunk with the definitive cumulative counts still fire.

Three tests replaced/added in openaiShim.test.ts:

  1. Regression: trailing chunk fires when stop chunk used lastSeenUsage fallback (2× message_delta emitted)
  2. Trailing chunk with higher real counts supersedes the provisional fallback
  3. Guard still suppresses trailing chunk when stop chunk had real chunkUsage (1× message_delta)

@jatmn jatmn left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the update. I rechecked the changed paths and found an issue that still needs to be addressed.

Findings

  • [P2] Avoid double-counting usage when a trailing final chunk arrives
    src/services/api/openaiShim.ts:1770
    The stop chunk now emits lastSeenUsage as a message_delta when it has no chunkUsage, and the trailing empty-choices path can then emit another usage-bearing message_delta for the same API response. The new regression tests even cover this as two deltas, but claude.ts treats every usage-bearing message_delta as additive by calling addToTotalSessionCost, which increments STATE.modelUsage totals in cost-tracker.ts. For providers that send early/provisional usage followed by a final empty-choices usage chunk, this will add both the provisional counts and the final counts to the status-line totals/cost instead of replacing the first value with the final one. The same condition also replays stale lastSeenUsage on a trailing empty-choices chunk that has no usage of its own, so an otherwise harmless empty trailer can duplicate accounting too. Please make the carried-forward stop usage non-additive, or otherwise ensure the later final usage supersedes it without charging/updating the global totals twice.

…arly provisional + trailing final chunks

Providers like MiMo/OpenGateway emit an early chunk with provisional usage
data, then a trailing empty-choices chunk with the real final totals.
Previously the stop-chunk handler fell back to `lastSeenUsage` when
`chunkUsage` was absent, which emitted the provisional counts at stop time;
the in-loop post-check then fired again on the trailing chunk and emitted the
real final counts -- both flowing into addToTotalSessionCost and doubling the
session totals.

Fix:
- At stop time, only attach usage when the stop chunk itself carries it.
- Add a post-stream fallback (after the reader is released) that emits
  lastSeenUsage once for providers that send usage only in an early chunk
  with no trailing empty-choices chunk.

The existing in-loop trailing-chunk check already handles the trailing-chunk
case; the post-stream path closes the gap for providers that omit it.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
@ajsai47

ajsai47 commented May 31, 2026

Copy link
Copy Markdown
Author

Thanks for the precise diagnosis, @jatmn.

Root cause confirmed: when a provider emits usage in an early/provisional chunk and the stop chunk carries no usage, the stop-chunk handler fell back to lastSeenUsage via chunkUsage ?? lastSeenUsage — emitting provisional counts at stop time. Then if a trailing empty-choices chunk arrived, the in-loop post-check fired again with the real final counts, causing both to flow into addToTotalSessionCost.

Fix (latest commit ff12bdd):

  1. Stop chunk — now only attaches usage when the chunk itself carries it (chunkUsage). No lastSeenUsage fallback here. The comment explains why.

  2. Trailing-chunk in-loop check — unchanged; handles providers that send real final usage in a trailing empty-choices chunk correctly.

  3. Post-stream fallback (new) — after reader.releaseLock(), emits lastSeenUsage once for providers that send usage only in an early chunk with no trailing chunk at all. This closes the "provider sends [early_usage] [stop] with no trailing chunk" gap that removing the fallback from the stop chunk would otherwise create.

The three cases now behave correctly:

  • [stop_with_usage] → emits at stop, hasEmittedFinalUsage = true, nothing else fires ✓
  • [early_usage] [stop] [trailing_with_final_usage] → stop emits nothing; trailing in-loop check emits real final once ✓
  • [early_usage] [stop] (no trailing chunk) → stop emits nothing; post-stream fallback emits lastSeenUsage once ✓

@jatmn jatmn left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the update. I rechecked the changed paths and found one remaining issue that needs to be addressed.

Findings

  • [P2] Update the status-line contract for unsupported usage
    src/tools/AgentTool/built-in/statuslineSetup.ts:58
    getCurrentUsage() now returns null after a completed assistant turn when the latest provider returned the shimmed all-zero usage record, and StatusLine.tsx passes that through to context_window.current_usage, used_percentage, and remaining_percentage. The built-in status-line setup prompt still documents those null values as meaning only "no messages yet", so /statusline setup will teach users and generated status-line commands the wrong contract for the new MiMo/OpenGateway unsupported-usage state. Please update this status-line schema/help text, or expose a distinguishable unsupported-usage state, so the user-facing contract matches the behavior this PR introduces.

…tate

getCurrentUsage() returns null in two distinct cases:
1. No assistant messages yet (existing behavior, already documented)
2. Active provider does not report token usage — e.g. providers that strip
   stream_options (MiMo, OpenGateway) return all-zero usage records, which
   getCurrentUsage() treats as unsupported and returns null for.

The /statusline setup prompt previously documented only case 1, so generated
status-line commands would mislead users into thinking null only means
"session not started". Update current_usage, used_percentage, and
remaining_percentage inline docs to call out both null triggers.

The existing // empty jq fallback in the examples already handles both cases
correctly — no example changes needed.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
@ajsai47

ajsai47 commented Jun 1, 2026

Copy link
Copy Markdown
Author

Fixed in commit 6dfc922 (docs(status-line): update schema to document unsupported-usage null state).

Updated the three field comments in statuslineSetup.ts to call out both null cases:

"current_usage": {  // Token usage from last API call.
                    // null when: no messages yet, OR the active provider does not
                    // report token usage (e.g. providers that strip stream_options
                    // such as MiMo/OpenGateway).
  ...
} | null,
"used_percentage": number | null,    // null when: no messages yet, OR provider does not report usage.
"remaining_percentage": number | null // null when: no messages yet, OR provider does not report usage.

The // empty jq fallback in the examples already handles both null cases correctly — no example changes needed.

@jatmn jatmn left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the update. I rechecked the changed paths and found one issue that still needs to be addressed.

Findings

  • [P2] Do not finalize usage from empty chunks that carry no usage
    src/services/api/openaiShim.ts:1788
    The trailing empty-choices branch now falls back to lastSeenUsage even when the empty chunk itself has no usage, and then sets hasEmittedFinalUsage = true. If a provider sends early/provisional usage, the stop chunk, an empty keepalive/trailer with no usage, and then the real final empty-choices usage chunk, the no-usage chunk will replay the provisional counts and suppress the later final totals. claude.ts treats every usage-bearing message_delta as additive session accounting, so this leaves status-line totals and costs based on the stale provisional usage instead of the definitive final chunk. Please only let this in-loop empty-choices path finalize when the current chunk has real usage, and leave the lastSeenUsage fallback to the post-stream path where no later final chunk can still arrive.

@jatmn

jatmn commented Jun 27, 2026

Copy link
Copy Markdown
Collaborator

closing as abandoned

@jatmn jatmn closed this Jun 27, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Status line usage always reports 0% (Global token state not updated in QueryEngine)

2 participants