Skip to content

fix(messages): request usage on streaming so the bridge meters tokens - #1180

Merged
njbrake merged 5 commits into
mainfrom
fix/messages-stream-include-usage
Jul 16, 2026
Merged

njbrake merged 5 commits into
mainfrom
fix/messages-stream-include-usage

Conversation

@njbrake

@njbrake njbrake commented Jul 15, 2026 •

Copy link
Copy Markdown
Contributor

Description

Streaming through the Messages to Completions bridge did not set stream_options.include_usage, so OpenAI-compatible backends omitted token usage from their streamed chunks. The recent trailing-usage-chunk fix (#1151) then had no usage-only chunk to flush, so streamed amessages against those providers reported zero input and output tokens (and no cache). Native Anthropic was unaffected, since it overrides _amessages and streams usage directly.

This requests stream_options.include_usage in the streaming branch of messages_params_to_completion_params, so the backend emits the trailing usage-only chunk that the stream wrapper flushes into the closing message_delta. It completes the chain #1151 started: #1151 captures the trailing chunk, this makes sure the trailing chunk is actually produced.

Safe across providers:

  • Providers that do not support stream_options (Cerebras, Ollama, Together, Cohere, Mistral) already strip it in their own _convert_completion_params.
  • The native Anthropic provider never reaches this bridge (it overrides _amessages), so its SDK never sees a stream_options kwarg.

Downstream context: this is the root cause of the zero-metering behavior reported in mozilla-ai/otari#256 (streamed /v1/messages recorded zero tokens and zero cost while non-streaming messages and streaming chat completions metered fine). The old any-llm gateway solved the same class of bug for chat completions in #974; this brings the messages path in line.

PR Type

  • 🐛 Bug Fix

Relevant issues

Complements #1151. Root cause of mozilla-ai/otari#256.

Checklist

  • I understand the code I am submitting.
  • I have added unit tests that prove my fix/feature works
  • I have run this code locally and verified it fixes the issue.
  • New and existing tests pass locally
  • Documentation was updated where necessary
  • I have read and followed the contribution guidelines
  • AI Usage:
    • No AI was used.
    • AI was used for drafting/refactoring.
    • This is fully AI-generated.

AI Usage Information

Testing

Full unit suite passes (1490 passed, 64 skipped), plus ruff and mypy strict clean over src and the touched tests. New tests:

  • test_messages_compat.py: stream_options.include_usage is present when streaming, absent when not streaming and when stream is unset.
  • test_messages.py: the CompletionParams handed to _acompletion actually carries include_usage on the streaming path and omits it on the non-streaming path. This is the coverage the trailing-chunk fix lacked (its tests hand-fed usage chunks, so they could not catch that usage was never requested).

Summary by CodeRabbit

  • Bug Fixes
    • Improved streamed completion token-usage reporting by requesting a final usage-only chunk when streaming is enabled.
    • Filtered OpenAI-specific stream_options from provider requests where unsupported (Watsonx, Groq, xAI, and Azure).
  • Tests
    • Added unit tests validating streaming vs non-streaming forwarding behaviour for the messages→completion bridge.
    • Added dedicated unit tests ensuring Watsonx, Groq, xAI, and Azure conversions drop stream_options.

@njbrake
njbrake temporarily deployed to integration-tests July 15, 2026 18:45 — with GitHub Actions Inactive
@coderabbitai

coderabbitai Bot commented Jul 15, 2026 •

Copy link
Copy Markdown

Review Change Stack

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: f681109d-6b2d-402e-8445-ede0644baddf

📥 Commits

Reviewing files that changed from the base of the PR and between 68bc002 and 0b5315d.

📒 Files selected for processing (1)
  • tests/unit/test_messages.py

Walkthrough

Changes

Streaming message-to-completion conversion now requests a trailing usage chunk when enabled. Azure, Groq, Watsonx, and xAI conversions exclude the OpenAI-specific streaming option, with tests covering conversion, bridge propagation, and provider filtering.

Streaming usage propagation

Layer / File(s) Summary
Completion parameter conversion
src/any_llm/utils/messages_compat.py, tests/unit/test_messages_compat.py
Streaming parameters include stream_options: {"include_usage": True}; non-streaming and unset stream values omit the option.
Asynchronous messages bridge integration
tests/unit/test_messages.py
AnyLLM._amessages tests verify the converted options passed to _acompletion for streaming and non-streaming requests.
Provider payload filtering
src/any_llm/providers/{azure,groq,watsonx,xai}/*, tests/unit/providers/*
Azure, Groq, Watsonx, and xAI conversions remove stream_options from provider payloads, with regression coverage.

Possibly related PRs

Suggested reviewers: tbille

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Title check ✅ Passed The title is concise and accurately summarises the main change: requesting usage on streaming so token metering works.
Description check ✅ Passed The description matches the required template and includes the change summary, PR type, related issue, checklist, testing, and AI usage details.
Docstring Coverage ✅ Passed Docstring coverage is 93.75% which is sufficient. The required threshold is 80.00%.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/messages-stream-include-usage

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@codecov

codecov Bot commented Jul 15, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

Files with missing lines Coverage Δ
src/any_llm/providers/azure/azure.py 80.58% <100.00%> (ø)
src/any_llm/providers/groq/groq.py 83.18% <100.00%> (-9.74%) ⬇️
src/any_llm/providers/watsonx/watsonx.py 80.61% <ø> (ø)
src/any_llm/providers/xai/xai.py 70.96% <ø> (-13.71%) ⬇️
src/any_llm/utils/messages_compat.py 100.00% <100.00%> (ø)

... and 32 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@tests/unit/test_messages.py`:
- Line 445: Move the inline import of ChatCompletion, ChatCompletionMessage, and
Choice from the test body to the module-level import section at the top of
tests/unit/test_messages.py, preserving the existing import usage and removing
the duplicate inline import.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: d57dc3da-457a-417f-b251-c030a8d13bc5

📥 Commits

Reviewing files that changed from the base of the PR and between eeeb4f2 and fdb0486.

📒 Files selected for processing (3)
  • src/any_llm/utils/messages_compat.py
  • tests/unit/test_messages.py
  • tests/unit/test_messages_compat.py

Comment thread tests/unit/test_messages.py Outdated
@njbrake
njbrake temporarily deployed to integration-tests July 15, 2026 19:00 — with GitHub Actions Inactive
@njbrake njbrake added the run-integration-tests Put this label on a PR to trigger the integration test suite: works with forks label Jul 15, 2026
@github-actions github-actions Bot removed the run-integration-tests Put this label on a PR to trigger the integration test suite: works with forks label Jul 15, 2026

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
src/any_llm/providers/watsonx/watsonx.py (1)

74-79: 🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

Apply the filter after merging kwargs. converted_params.update(kwargs) can reintroduce stream_options, so the unsupported field may still reach Watsonx. Filter it after the merge and add coverage for the kwargs path.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/any_llm/providers/watsonx/watsonx.py` around lines 74 - 79, Update the
parameter-building flow around converted_params so kwargs are merged before
filtering unsupported fields, ensuring stream_options cannot be reintroduced by
converted_params.update(kwargs). Preserve the existing reasoning_effort handling
and add coverage verifying stream_options supplied through kwargs is excluded
from the Watsonx request.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Outside diff comments:
In `@src/any_llm/providers/watsonx/watsonx.py`:
- Around line 74-79: Update the parameter-building flow around converted_params
so kwargs are merged before filtering unsupported fields, ensuring
stream_options cannot be reintroduced by converted_params.update(kwargs).
Preserve the existing reasoning_effort handling and add coverage verifying
stream_options supplied through kwargs is excluded from the Watsonx request.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 11aa25b2-4eca-47b0-b86e-2b81372e4b74

📥 Commits

Reviewing files that changed from the base of the PR and between fdb0486 and d6331fa.

📒 Files selected for processing (2)
  • src/any_llm/providers/watsonx/watsonx.py
  • tests/unit/providers/test_watsonx_provider.py

njbrake and others added 3 commits July 15, 2026 19:23
The Messages->Completions bridge streamed without setting
stream_options.include_usage, so OpenAI-compatible backends omitted token
usage from their chunks. The recent trailing-usage-chunk fix (#1151) then had
no usage-only chunk to flush, so streamed amessages against those providers
reported zero input and output tokens. Native Anthropic was unaffected (it
overrides _amessages and streams usage directly).

Request stream_options.include_usage in the streaming branch of
messages_params_to_completion_params so the backend emits the trailing
usage-only chunk that the stream wrapper flushes into the closing
message_delta. Providers that do not support stream_options already strip it in
their own param conversion, and the native Anthropic provider never reaches
this bridge, so the injection is safe across providers.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…payload

The Messages bridge now sets stream_options for every streamed request, but
Watsonx merges its params dict straight into the chat payload without filtering
against TextChatParameters, so the OpenAI-only stream_options field would be
forwarded to the Watsonx API. Exclude it in _convert_completion_params, matching
the other OpenAI-incompatible providers (Cerebras, Cohere, Mistral, Ollama,
Together), so streaming stays unaffected while usage still comes back natively.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Like Watsonx, the Groq and xAI providers pass their params through to a
non-OpenAI SDK client that rejects the OpenAI-only stream_options field. The
Messages bridge now sets stream_options for every streamed request, so
streaming /v1/messages (and any streaming call that carries stream_options)
raised TypeError: create() got an unexpected keyword argument 'stream_options'.
Exclude it in _convert_completion_params, matching the other providers that
already drop it (Cerebras, Cohere, Mistral, Ollama, Together, Watsonx). Caught
by the integration suite (test_messages_streaming[groq], [xai]).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@njbrake
njbrake force-pushed the fix/messages-stream-include-usage branch from d6331fa to e3baea6 Compare July 15, 2026 19:25
@njbrake
njbrake temporarily deployed to integration-tests July 15, 2026 19:25 — with GitHub Actions Inactive
@njbrake njbrake added the run-integration-tests Put this label on a PR to trigger the integration test suite: works with forks label Jul 15, 2026
@github-actions github-actions Bot removed the run-integration-tests Put this label on a PR to trigger the integration test suite: works with forks label Jul 15, 2026
@njbrake
njbrake requested review from khaledosman and tbille July 15, 2026 23:04
The Azure AI Inference SDK does not model the OpenAI-only stream_options
field; complete() forwards unknown kwargs down to the transport, which
rejects them with TypeError: ClientSession._request() got an unexpected
keyword argument 'stream_options'. The Messages bridge now sets
stream_options for every streamed request, so streaming /v1/messages
through the Azure provider would crash. Exclude it in
_convert_completion_params, matching the other providers that already
drop it. Reproduced in isolation against the installed SDK; the
integration suite could not catch this because Azure is not credentialed
in CI.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@njbrake
njbrake temporarily deployed to integration-tests July 15, 2026 23:09 — with GitHub Actions Inactive

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
src/any_llm/providers/azure/azure.py (1)

186-193: 🩺 Stability & Availability | 🟠 Major | ⚡ Quick win

Filter stream_options after merging kwargs. params.model_dump() excludes it, but call_kwargs.update(kwargs) can add it back and send an unsupported field to Azure.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/any_llm/providers/azure/azure.py` around lines 186 - 193, Update the
call-kwargs construction in the relevant Azure provider method to remove or
filter the stream_options key after call_kwargs.update(kwargs), ensuring kwargs
cannot reintroduce this unsupported field before returning. Preserve the
existing model_dump exclusions and other merged parameters.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@tests/unit/providers/test_azure_provider.py`:
- Around line 232-243: Extend
test_convert_completion_params_drops_stream_options to pass stream_options
directly through the kwargs argument of
AzureProvider._convert_completion_params, alongside or separately from the
CompletionParams field, and assert the returned mapping excludes it. Preserve
the existing params-based coverage while adding the extra-kwargs path that
exercises call_kwargs.update(kwargs).

---

Outside diff comments:
In `@src/any_llm/providers/azure/azure.py`:
- Around line 186-193: Update the call-kwargs construction in the relevant Azure
provider method to remove or filter the stream_options key after
call_kwargs.update(kwargs), ensuring kwargs cannot reintroduce this unsupported
field before returning. Preserve the existing model_dump exclusions and other
merged parameters.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: fd48acf9-1531-452e-82ed-eda30b8a3bd0

📥 Commits

Reviewing files that changed from the base of the PR and between e3baea6 and 68bc002.

📒 Files selected for processing (2)
  • src/any_llm/providers/azure/azure.py
  • tests/unit/providers/test_azure_provider.py

Comment thread tests/unit/providers/test_azure_provider.py

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Fixes missing token metering when streaming through the Messages to Completions compatibility bridge by ensuring OpenAI-compatible backends are asked to emit usage during streaming. This aligns the Messages streaming path with the already-fixed trailing usage-only chunk handling so streamed amessages reports correct input, output, and cache token counts.

Changes:

  • Inject stream_options={"include_usage": True} when MessagesParams.stream=True in messages_params_to_completion_params.
  • Filter out stream_options in providers whose SDKs reject unknown OpenAI streaming knobs (xAI, Groq, Watsonx, Azure).
  • Add unit tests covering streaming vs non-streaming propagation of stream_options through both the compat layer and the _amessages bridge.

Reviewed changes

Copilot reviewed 11 out of 11 changed files in this pull request and generated no comments.

Show a summary per file
File Description
tests/unit/test_messages.py Verifies _amessages passes include_usage when streaming and omits it when non-streaming.
tests/unit/test_messages_compat.py Verifies the compat conversion adds stream_options only for stream=True and omits it otherwise.
tests/unit/providers/test_xai_provider.py Asserts xAI conversion drops stream_options.
tests/unit/providers/test_watsonx_provider.py Asserts Watsonx conversion drops stream_options.
tests/unit/providers/test_groq_provider.py Asserts Groq conversion drops stream_options.
tests/unit/providers/test_azure_provider.py Asserts Azure conversion drops stream_options.
src/any_llm/utils/messages_compat.py Implements stream_options.include_usage injection for streaming bridge requests.
src/any_llm/providers/xai/xai.py Drops unsupported stream_options in xAI provider param conversion.
src/any_llm/providers/watsonx/watsonx.py Drops unsupported stream_options in Watsonx provider param conversion.
src/any_llm/providers/groq/groq.py Drops unsupported stream_options in Groq provider param conversion.
src/any_llm/providers/azure/azure.py Drops unsupported stream_options in Azure provider param conversion.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

@khaledosman khaledosman left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the fix end-to-end: verified the bridge chain into the #1151 trailing-chunk flush, and checked the provider blast radius against actual SDK signatures (groq create and xai chat.create have no stream_options parameter, azure complete() does not model it, huggingface-hub 1.3.2 does, and cohere/together/cerebras/ollama/mistral already strip it; bedrock/sagemaker/gemini/lmstudio build kwargs explicitly and ignore it). Groq and xAI still meter streamed usage despite the strip since both surface usage natively on chunks. Touched test files, ruff check and format pass locally; CI is green.

Two minor inline comments below. One observation that fits neither diff line: every BaseOpenAIProvider subclass (perplexity, deepinfra, dashscope, moonshot, ...) now sends stream_options to its remote endpoint on the streamed messages path. The OpenAI SDK accepts it, but backend tolerance can't be verified statically; one run with the run-integration-tests label would derisk that.

🤖 This review was created by Claude Code.

Comment thread src/any_llm/utils/messages_compat.py
Comment thread tests/unit/test_messages.py Outdated
khaledosman
khaledosman previously approved these changes Jul 16, 2026
any_llm.types.completion is a core dependency and the file already imports
from it at the top, so ChatCompletion, ChatCompletionMessage and Choice
belong in the module-level import block rather than inline in a test body.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@njbrake
njbrake temporarily deployed to integration-tests July 16, 2026 12:44 — with GitHub Actions Inactive
@njbrake
njbrake merged commit 891a51f into main Jul 16, 2026
12 checks passed

This branch was previously deployed

1 inactive deployment
integration-tests — 0b5315d9 Deployed Jul 16, 2026 by njbrake via run-docs-tests #2177
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

1.21.0 Included in release 1.21.0

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants