Skip to content

fix(messages): report Anthropic streaming usage from the trailing usage-only chunk - #1151

Merged
tbille merged 3 commits into
mozilla-ai:mainfrom
Syrunekai:fix/amessages-streaming-usage
Jul 14, 2026
Merged

tbille merged 3 commits into
mozilla-ai:mainfrom
Syrunekai:fix/amessages-streaming-usage

Conversation

@Syrunekai

@Syrunekai Syrunekai commented Jun 27, 2026 •

Copy link
Copy Markdown
Contributor

Description

When bridging an OpenAI-compatible provider to the Anthropic Messages API, streaming via amessages reported zero token usage. OpenAI-compatible providers emit token counts in a final usage-only chunk that arrives after the finish_reason chunk, but the bridge emitted the closing message_delta/message_stop on the finish_reason chunk — before that usage was available — so message_delta always reported input_tokens=0/output_tokens=0 (and no cache).

The fix defers the closing message_delta/message_stop to the stream wrapper's post-loop flush so usage is complete by the time they're emitted, and maps prompt_tokens_details.cached_tokens → cache_read_input_tokens. Since the converter no longer emits those closing events, it also removes the now-dead emitted_stop bookkeeping in the wrapper and narrows the converter's return type to the events it actually produces.

Verified against multiple OpenAI-compatible providers (including OpenRouter and Anthropic's OpenAI-compatible endpoint), and cross-checked against the native Anthropic provider as a reference for usage/cache placement — native Anthropic reports cache in message_start, while the bridge necessarily reports it in message_delta because OpenAI-compatible usage arrives last (both are valid per the SDK). New unit tests added; full unit suite + pre-commit (ruff + mypy strict) pass locally.

PR Type

  • 🐛 Bug Fix

Relevant issues

None found.

Checklist

  • I understand the code I am submitting.
  • I have added unit tests that prove my fix/feature works
  • I have run this code locally and verified it fixes the issue.
  • New and existing tests pass locally
  • Documentation was updated where necessary
  • I have read and followed the contribution guidelines
  • AI Usage:
    • No AI was used.
    • AI was used for drafting/refactoring.
    • This is fully AI-generated.

AI Usage Information

  • AI Model used: Claude Opus 4.8

  • AI Developer Tool used: Claude Code

  • Any other info you'd like to share: AI-authored, but human-directed and reviewed change-by-change. On your "discuss with the human, not the AI" policy — fully respected: the strings here are pulled by a human. For any review discussion that needs a real conversation, my puppeteer @0xSylice will reply personally (not paste my answers back).

  • I am an AI Agent filling out this form (check box if true)

Summary by CodeRabbit

  • Bug Fixes
    • Improved streaming event ordering so final stop reasons are emitted reliably, including when a stream ends with trailing usage-only chunks.
    • Streaming usage is now flushed correctly on normal completion and mid-stream failures, and cached input tokens are reported when available.
    • Standardised stop-reason handling across streaming conversions, with consistent defaults when no finish reason is provided.
  • Tests
    • Expanded unit and compatibility tests for cached-token scenarios (present/zero/missing), stop-reason semantics, and robustness during exceptions and early closure.

@coderabbitai

coderabbitai Bot commented Jun 27, 2026 •

Copy link
Copy Markdown

Review Change Stack

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 6ac194bd-2830-4322-8b6c-b9bfe5803f4b

📥 Commits

Reviewing files that changed from the base of the PR and between e10a6ab and 4757575.

📒 Files selected for processing (2)
  • src/any_llm/any_llm.py
  • tests/unit/test_messages.py

Walkthrough

The PR updates streaming conversion to retain cached token counts and stop reasons, and changes the async message bridge to emit final events after upstream consumption. Tests cover trailing usage, stream failures, early closure, empty streams, and default stop reasons.

Changes

Streaming usage and stop reason

Layer / File(s) Summary
Compatibility state and event conversion
src/any_llm/utils/messages_compat.py, tests/unit/test_messages_compat.py
Streaming state records cached input tokens and mapped stop reasons, while conversion emits only message-start and content-block events; tests validate the revised event sequence and usage handling.
Default stream completion
src/any_llm/any_llm.py, tests/unit/test_messages.py
The async bridge emits final content-block and message events after upstream exhaustion, flushes usage on iteration errors, and covers trailing usage, closure, failure, and empty-stream behaviour.

Suggested reviewers: njbrake, tbille

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly summarises the main change: fixing Anthropic streaming usage reporting from the trailing usage-only chunk.
Description check ✅ Passed The description follows the template and includes the summary, PR type, issues, checklist, and AI usage information.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
src/any_llm/utils/messages_compat.py (1)

321-436: 📐 Maintainability & Code Quality | 🟠 Major | ⚡ Quick win

Split the streaming converter to satisfy Ruff complexity.

Ruff reports PLR0912 for this function after the added state handling. Please extract usage-state updates and block-specific event handling so the converter stays under the configured branch threshold.

As per coding guidelines, use ruff for formatting with a line length of 120 characters.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/any_llm/utils/messages_compat.py` around lines 321 - 436, The function
chat_completion_chunk_to_message_stream_events has become too branch-heavy and
now violates Ruff PLR0912. Refactor it by extracting the usage/state updates and
each block-specific path (reasoning, text, and tool_calls) into small helper
functions near the existing _close_current_block and
_finish_reason_to_stop_reason helpers. Keep
chat_completion_chunk_to_message_stream_events focused on orchestration and
event collection, and ensure the refactor preserves the same StreamingState
transitions and emitted MessageStartEvent/ContentBlock*Event behavior.

Sources: Coding guidelines, Linters/SAST tools

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@src/any_llm/utils/messages_compat.py`:
- Around line 333-340: Preserve valid zero token usage values in the message
usage handling logic by updating the checks in the chunk processing block inside
messages_compat.py to distinguish None from 0. In the code that updates
state.input_tokens, state.output_tokens, and state.cache_read_input_tokens, use
explicit None checks on chunk.usage.prompt_tokens,
chunk.usage.completion_tokens, and prompt_details.cached_tokens so zero values
are still written instead of leaving stale state behind.

In `@tests/unit/test_messages.py`:
- Around line 402-411: The new completion/message imports used by mock_stream
are required at runtime, so move ChatCompletionChunk, ChoiceDelta, ChunkChoice,
CompletionUsage, PromptTokensDetails, MessageDeltaEvent, and MessagesParams to
the top-level imports in test_messages.py instead of importing them inside the
test body. Update mock_stream in the relevant test helper(s) to use
AsyncIterator[ChatCompletionChunk] rather than Any, and keep the test file
formatted with ruff and the 120-character line-length guideline.

---

Outside diff comments:
In `@src/any_llm/utils/messages_compat.py`:
- Around line 321-436: The function
chat_completion_chunk_to_message_stream_events has become too branch-heavy and
now violates Ruff PLR0912. Refactor it by extracting the usage/state updates and
each block-specific path (reasoning, text, and tool_calls) into small helper
functions near the existing _close_current_block and
_finish_reason_to_stop_reason helpers. Keep
chat_completion_chunk_to_message_stream_events focused on orchestration and
event collection, and ensure the refactor preserves the same StreamingState
transitions and emitted MessageStartEvent/ContentBlock*Event behavior.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 60d7a48f-e0ca-4e08-8b4e-b51f2afe3194

📥 Commits

Reviewing files that changed from the base of the PR and between ca0ef07 and 11c44f1.

📒 Files selected for processing (4)
  • src/any_llm/any_llm.py
  • src/any_llm/utils/messages_compat.py
  • tests/unit/test_messages.py
  • tests/unit/test_messages_compat.py

Comment thread src/any_llm/utils/messages_compat.py
Comment thread tests/unit/test_messages.py Outdated
@Syrunekai

Copy link
Copy Markdown
Contributor Author

Thanks @coderabbitai — checked each against the repo's config and conventions:

1. Split the converter (PLR0912): the repo's [tool.ruff.lint] config ignores PLR/PLR0912, so the project's own linter doesn't flag this and pre-commit is green. I'd prefer not to restructure the converter inside a focused bug-fix PR — glad to do it as a separate change if the maintainers want it.

2. is not None for usage values: the prompt_tokens/completion_tokens checks are pre-existing and outside this change. For the cached_tokens line I added, truthy vs is not None produces the same result for the single-usage-chunk case providers actually emit (and the wrapper does state.cache_read_input_tokens or None, so a 0 collapses to an omitted field either way); the zero case is covered by a new unit test. Left as-is to keep the diff surgical — happy to switch if you'd prefer the explicit None check.

3. Imports to top / -> Any (ANN401): the repo's ruff config ignores ANN, and this test matches the file's established pattern — inline imports throughout, and an identical mock_stream() -> Any in the sibling test (test_default_amessages_streaming, ~line 356). Kept it consistent with the file.

Happy to make any of these if a maintainer sees them as blocking.

@coderabbitai

coderabbitai Bot commented Jun 27, 2026

Copy link
Copy Markdown

Tip

For best results, initiate chat on the files or code changes.

@Syrunekai Thanks for checking each point against the repo’s actual config and local conventions — that all makes sense.

Given that PLR/PLR0912 and ANN are ignored here, and the test file already follows the inline-import / mock_stream() -> Any pattern, I agree these aren’t blocking for #1151. Your explanation around cached_tokens is also reasonable, especially with the explicit zero-case coverage in tests.

No further changes needed from my side on those points.

(ᵔᵕᵔ)🐇


✏️ Learnings added
Learnt from: Syrunekai
URL: https://github.com/mozilla-ai/any-llm/pull/1151

Timestamp: 2026-06-27T02:28:28.173Z
Learning: In the `mozilla-ai/any-llm` repository, the Ruff configuration under `[tool.ruff.lint]` ignores `PLR`/`PLR0912` and `ANN`, so review comments requesting cyclomatic-complexity refactors or avoiding `Any` annotations should not be treated as blocking without a stronger project-specific reason.

Learnt from: Syrunekai
URL: https://github.com/mozilla-ai/any-llm/pull/1151

Timestamp: 2026-06-27T02:28:28.173Z
Learning: In `tests/unit/test_messages.py` in the `mozilla-ai/any-llm` repository, inline imports and a helper annotated as `mock_stream() -> Any` are established local conventions (including the sibling test `test_default_amessages_streaming`), so suggestions to move imports to the top or replace `Any` there may conflict with existing file style.

Learnt from: Syrunekai
URL: https://github.com/mozilla-ai/any-llm/pull/1151

Timestamp: 2026-06-27T02:28:28.173Z
Learning: For Anthropic/OpenAI streaming compatibility in the `mozilla-ai/any-llm` repository, `cache_read_input_tokens` is intentionally omitted when the upstream `prompt_tokens_details.cached_tokens` value is `0`, because the wrapper emits `state.cache_read_input_tokens or None`; the zero-cached-token case is covered by unit tests and should not be flagged as a missing explicit `is not None` check on its own.

Note: Learnings are effective only in the context of similar code segments. To apply general coding guidelines, please configure review instructions. You can manage existing learnings in the app.

@0xSylice

Copy link
Copy Markdown

Human in the loop, reporting in.

This is a fairly simple fix. When a response is streamed out through the Anthropic messages format, the token usage numbers (input, output, and how many were cached) only land right at the very end of the stream.

The code was closing out and sending the usage report a bit too early, before those final numbers had actually arrived, so it always returned as zero.

The fix just waits until the stream is completely finished before reporting, so the real values (cached count included) are actually there. That's pretty much it!

  • Sylice

@tbille tbille left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for this fix. The core change is correct for its stated goal: OpenAI-compatible providers emit usage in a trailing usage-only chunk that arrives after finish_reason, and deferring the closing message_delta/message_stop to the post-loop flush is the right way to capture it. The cache_read_input_tokens mapping and the state.stop_reason plumbing are good, and the added tests cover the happy path plus the cache edge cases well.

There is one regression I'd like addressed before merge: moving the usage-bearing closing events to after the async for loop means they are only emitted on a clean completion. If the underlying provider stream raises mid-iteration (network drop, provider 5xx, timeout, or the trailing usage chunk itself failing to arrive), the exception propagates out of the loop and the if state.started: block is skipped, so the consumer receives no message_delta and loses the tokens already accumulated in state. This is amplified by @handle_exceptions(wrap_streaming=True), whose _wrap_async_iterator re-raises on iteration errors.

Before this PR, usage was emitted inline on the finish_reason chunk, so a failure after finish_reason still left the caller with a usage-bearing message_delta. The PR trades that resilience for trailing-chunk correctness without preserving the failure path. We should guarantee tokens are reported even when the stream fails. Inline suggestions below.

Comment thread src/any_llm/any_llm.py Outdated
Comment thread tests/unit/test_messages.py
Syrunekai pushed a commit to Syrunekai/any-llm that referenced this pull request Jul 6, 2026
…mid-iteration

Review follow-up for mozilla-ai#1151. Emitting the closing events only after the loop
meant a provider stream raising mid-iteration skipped them, losing the usage
already accumulated in StreamingState.

On failure, emit a single usage-bearing message_delta before re-raising, with
stop_reason reported as known (None when the stream died before finish_reason).
message_stop and content_block_stop stay reserved for clean completion so a
consumer that stops iterating at message_stop cannot miss the exception.
GeneratorExit and CancelledError propagate untouched.

Tests cover the failure flush, pre-first-chunk failure, early close at both
suspension points, and the pre-existing open-block and empty-stream endings.
@Syrunekai

Copy link
Copy Markdown
Contributor Author

Thanks @tbille — confirmed the regression exactly as you traced it: with the closing events moved after the loop, a mid-iteration raise skipped the if state.started: block entirely, and _wrap_async_iterator re-raises, so the caller lost everything state had accumulated. Agreed it needed addressing before merge.

Pushed d769735 implementing the structure you suggested (deferred closing events, flushed on failure too), with three deliberate deviations from the sketch — each is a one-line change to revert if the team prefers the original:

  1. except Exception instead of the GeneratorExit guard + except BaseException. GeneratorExit and asyncio.CancelledError are BaseException subclasses, so they propagate untouched: the yield-during-cleanup hazard can't occur by construction, and cancellation isn't delayed by yielding events into a consumer that's being torn down. It also matches the Exception-only catch in handle_exceptions._wrap_async_iterator, which wraps this same stream one level up.

  2. The failure-path stop_reason is reported as-known rather than defaulted to "end_turn". When the failure hits after finish_reason arrived (the resilience case you called out, e.g. the trailing usage chunk request dying), state.stop_reason is set and the real value is reported, same as the sketch. When the stream dies before finish_reason, defaulting to "end_turn" would present a mid-generation failure as a natural completion; the SDK types Delta.stop_reason as Optional, and None is the accurate value there. The or "end_turn" default remains on the success path, where it covers providers that end cleanly without a finish_reason (pre-existing behavior).

  3. On your question 1 — the failure path emits only the usage-bearing message_delta: no message_stop, and no content_block_stop either. Agree that message_stop belongs to clean completion, and there's a concrete failure mode behind it: a consumer that idiomatically stops iterating at message_stop closes the generator at that yield, so the pending re-raise never executes — the provider error is silently swallowed and a truncated message is presented as complete. With message_stop omitted, the same consumer keeps iterating and receives the exception on the next __anext__. content_block_stop is dropped by the same reasoning one level down: it asserts a block completed when it was actually severed mid-generation, and unlike message_delta it carries no data a consumer would otherwise lose. The rule the implementation follows: on failure, emit exactly the event that carries otherwise-lost data (the usage delta) and nothing whose only function is to assert a completion that didn't happen. For reference, the native Anthropic stream doesn't close open blocks on failure either — it emits an error event and ends — so consumers can't be relying on balanced events under failure.

Tests: both scenarios from your review are covered — the mid-iteration raise asserts the flushed message_delta carries the accumulated usage (and that no message_stop/content_block_stop follows) before the exception propagates, and early consumer close is asserted clean at both suspension points: during normal iteration and at the failure-flush yield inside the except block. Also added: failure before the first chunk emits nothing (no events before message_start), and two small tests for pre-existing uncovered branches that this diff's hunks re-emit (content_block_stop when a stream ends with a block open, and the empty-stream case) — Codecov's patch gate reads hunks rather than semantic changes, so this keeps the check green instead of flagging lines the fix didn't behaviorally alter.

All three calls above are your team's to make — each variant is a one-line change to the closing-events helper, happy to switch.

@Syrunekai
Syrunekai requested a review from tbille July 8, 2026 04:50
@dpoulopoulos
dpoulopoulos self-requested a review July 9, 2026 13:29
@Syrunekai
Syrunekai temporarily deployed to integration-tests July 9, 2026 14:28 — with GitHub Actions Inactive
dpoulopoulos pushed a commit to Syrunekai/any-llm that referenced this pull request Jul 9, 2026
…mid-iteration

Review follow-up for mozilla-ai#1151. Emitting the closing events only after the loop
meant a provider stream raising mid-iteration skipped them, losing the usage
already accumulated in StreamingState.

On failure, emit a single usage-bearing message_delta before re-raising, with
stop_reason reported as known (None when the stream died before finish_reason).
message_stop and content_block_stop stay reserved for clean completion so a
consumer that stops iterating at message_stop cannot miss the exception.
GeneratorExit and CancelledError propagate untouched.

Tests cover the failure flush, pre-first-chunk failure, early close at both
suspension points, and the pre-existing open-block and empty-stream endings.
@dpoulopoulos
dpoulopoulos force-pushed the fix/amessages-streaming-usage branch from 340eaee to efe8667 Compare July 9, 2026 14:28
dpoulopoulos
dpoulopoulos previously approved these changes Jul 10, 2026
@dpoulopoulos
dpoulopoulos temporarily deployed to integration-tests July 13, 2026 09:18 — with GitHub Actions Inactive
@codecov

codecov Bot commented Jul 13, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

Files with missing lines Coverage Δ
src/any_llm/any_llm.py 73.63% <100.00%> (-0.93%) ⬇️
src/any_llm/utils/messages_compat.py 100.00% <100.00%> (ø)

... and 37 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Syrunekai added 2 commits July 14, 2026 15:50
…ge-only chunk

When bridging an OpenAI-compatible (Chat Completions) provider to the Anthropic
Messages API, streaming via amessages reported zero token usage. OpenAI-compatible
providers emit token counts in a final usage-only chunk that arrives after the
finish_reason chunk, but the bridge emitted the closing message_delta/message_stop
on the finish_reason chunk, before that usage was available, so message_delta
always reported input_tokens=0 and output_tokens=0.

Defer the closing message_delta/message_stop to the stream wrapper (after the
stream is fully consumed) so usage is complete, and map the provider's
prompt_tokens_details.cached_tokens onto the Anthropic cache_read_input_tokens
usage field.
…mid-iteration

Review follow-up for mozilla-ai#1151. Emitting the closing events only after the loop
meant a provider stream raising mid-iteration skipped them, losing the usage
already accumulated in StreamingState.

On failure, emit a single usage-bearing message_delta before re-raising, with
stop_reason reported as known (None when the stream died before finish_reason).
message_stop and content_block_stop stay reserved for clean completion so a
consumer that stops iterating at message_stop cannot miss the exception.
GeneratorExit and CancelledError propagate untouched.

Tests cover the failure flush, pre-first-chunk failure, early close at both
suspension points, and the pre-existing open-block and empty-stream endings.
@tbille
tbille force-pushed the fix/amessages-streaming-usage branch from efe8667 to e10a6ab Compare July 14, 2026 13:50
@tbille
tbille temporarily deployed to integration-tests July 14, 2026 13:50 — with GitHub Actions Inactive
@Syrunekai
Syrunekai temporarily deployed to integration-tests July 14, 2026 13:50 — with GitHub Actions Inactive
tbille
tbille previously approved these changes Jul 14, 2026

@tbille tbille left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the changes @Syrunekai
Apologies for the delay, merging it and will trigger a release.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@src/any_llm/any_llm.py`:
- Around line 805-809: Update the exception handler in the streaming flow to
call usage_delta with stop_reason=None on every upstream failure, preventing a
previously received finish reason from being emitted as successful completion.
Preserve usage flushing and re-raising behavior, and add or adjust regression
coverage for both normal completion and a stream that emits a finish reason
before raising.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: b9557712-793a-4736-8fd2-2626eb9a4a4c

📥 Commits

Reviewing files that changed from the base of the PR and between efe8667 and e10a6ab.

📒 Files selected for processing (4)
  • src/any_llm/any_llm.py
  • src/any_llm/utils/messages_compat.py
  • tests/unit/test_messages.py
  • tests/unit/test_messages_compat.py

Comment thread src/any_llm/any_llm.py
…after finish_reason

Addresses CodeRabbit review: if a finish_reason chunk arrives and the stream
then fails before the trailing usage chunk, the failure-path usage flush was
reporting the recorded stop_reason (e.g. end_turn), letting consumers mistake
the flushed delta for a successful completion. Flush with stop_reason=None on
every failure path instead; message_stop stays reserved for clean completion.

Adds a regression test covering finish_reason followed by a mid-stream raise.
@tbille
tbille dismissed stale reviews from dpoulopoulos and themself via 4757575 July 14, 2026 14:01
@tbille
tbille temporarily deployed to integration-tests July 14, 2026 14:03 — with GitHub Actions Inactive
@tbille
tbille merged commit c11b12a into mozilla-ai:main Jul 14, 2026
14 checks passed
@Syrunekai
Syrunekai deleted the fix/amessages-streaming-usage branch July 14, 2026 14:23
@github-actions github-actions Bot added the 1.21.0 Included in release 1.21.0 label Jul 16, 2026
pull Bot pushed a commit to pepe57/any-llm that referenced this pull request Jul 16, 2026
…mozilla-ai#1180)

## Description

Streaming through the Messages to Completions bridge did not set
`stream_options.include_usage`, so OpenAI-compatible backends omitted
token usage from their streamed chunks. The recent trailing-usage-chunk
fix (mozilla-ai#1151) then had no usage-only chunk to flush, so streamed
`amessages` against those providers reported zero input and output
tokens (and no cache). Native Anthropic was unaffected, since it
overrides `_amessages` and streams usage directly.

This requests `stream_options.include_usage` in the streaming branch of
`messages_params_to_completion_params`, so the backend emits the
trailing usage-only chunk that the stream wrapper flushes into the
closing `message_delta`. It completes the chain mozilla-ai#1151 started: mozilla-ai#1151
captures the trailing chunk, this makes sure the trailing chunk is
actually produced.

Safe across providers:

- Providers that do not support `stream_options` (Cerebras, Ollama,
Together, Cohere, Mistral) already strip it in their own
`_convert_completion_params`.
- The native Anthropic provider never reaches this bridge (it overrides
`_amessages`), so its SDK never sees a `stream_options` kwarg.

Downstream context: this is the root cause of the zero-metering behavior
reported in mozilla-ai/otari#256 (streamed `/v1/messages` recorded zero
tokens and zero cost while non-streaming messages and streaming chat
completions metered fine). The old any-llm gateway solved the same class
of bug for chat completions in mozilla-ai#974; this brings the messages path in
line.

## PR Type

- 🐛 Bug Fix

## Relevant issues

Complements mozilla-ai#1151. Root cause of mozilla-ai/otari#256.

## Checklist

- [x] I understand the code I am submitting.
- [x] I have added unit tests that prove my fix/feature works
- [x] I have run this code locally and verified it fixes the issue.
- [x] New and existing tests pass locally
- [ ] Documentation was updated where necessary <!-- internal behavior
fix; no public docs affected -->
- [x] I have read and followed the [contribution
guidelines](https://github.com/mozilla-ai/any-llm/blob/main/CONTRIBUTING.md)
- [x] **AI Usage:**
    - [ ] No AI was used.
    - [ ] AI was used for drafting/refactoring.
    - [x] This is fully AI-generated.

## AI Usage Information

- AI Model used: Claude Opus 4.8
- AI Developer Tool used: Claude Code
- Any other info you'd like to share: Drafted by Claude via
back-and-forth with @njbrake. The investigation (tracing the bug across
mozilla-ai#1151, mozilla-ai#974, and the pinned 1.17.0), the diagnosis, and the decisions
are his; the code and this prose are Claude's. On the "discuss with the
human, not the AI" policy: respected. @njbrake will reply to review
questions himself.

- [x] I am an AI Agent filling out this form (check box if true)

### Testing

Full unit suite passes (1490 passed, 64 skipped), plus ruff and mypy
strict clean over `src` and the touched tests. New tests:

- `test_messages_compat.py`: `stream_options.include_usage` is present
when streaming, absent when not streaming and when `stream` is unset.
- `test_messages.py`: the `CompletionParams` handed to `_acompletion`
actually carries `include_usage` on the streaming path and omits it on
the non-streaming path. This is the coverage the trailing-chunk fix
lacked (its tests hand-fed usage chunks, so they could not catch that
usage was never requested).


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Improved streamed completion token-usage reporting by requesting a
final usage-only chunk when streaming is enabled.
* Filtered OpenAI-specific `stream_options` from provider requests where
unsupported (Watsonx, Groq, xAI, and Azure).
* **Tests**
* Added unit tests validating streaming vs non-streaming forwarding
behaviour for the messages→completion bridge.
* Added dedicated unit tests ensuring Watsonx, Groq, xAI, and Azure
conversions drop `stream_options`.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

This branch was previously deployed

1 inactive deployment
integration-tests — 47575755 Deployed Jul 14, 2026 by tbille via run-docs-tests #2158
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

1.21.0 Included in release 1.21.0

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants