Skip to content

fix(minimax): preserve usage-only streaming chunks - #1249

Merged
njbrake merged 3 commits into
mozilla-ai:mainfrom
zhuziqi97:fix/minimax-stream-usage-chunk
Aug 11, 2026
Merged

njbrake merged 3 commits into
mozilla-ai:mainfrom
zhuziqi97:fix/minimax-stream-usage-chunk

Conversation

@zhuziqi97

@zhuziqi97 zhuziqi97 commented Aug 7, 2026 •

Copy link
Copy Markdown
Contributor

Description

Fixes MiniMax streaming usage loss reported in #1185.

MiniMax returns token usage in a trailing streaming chunk with choices=[].
MinimaxProvider._convert_completion_response_async() previously forwarded
only chunks containing a choice delta, so it discarded this usage-only chunk
before the Messages bridge or downstream stream handler could observe it.

This change keeps the existing MiniMax filtering behavior while also forwarding
chunks where chunk.usage is not None.

The regression test verifies that:

  • normal content chunks remain available;
  • unrelated empty chunks without usage remain filtered;
  • a trailing choices=[] chunk with usage survives conversion;
  • input, output, and total token values remain unchanged after the XML reasoning
    wrapper.

No usage is synthesized, and absent usage is not converted to zero. Non-streaming
behavior and other providers are unchanged.

Validation

Full local unit suite:

1920 passed, 68 skipped, 3 warnings

The warnings are existing provider exception deprecation warnings unrelated to
this change.

Full pre-commit suite passed, including Ruff, formatting, Mypy, codespell, and
repository file checks.

A live MiniMax-M3 streaming request was also verified with the patched provider.
The raw provider usage and normalized result matched:

input_tokens: 191
output_tokens: 27
total_tokens: 218
cached_tokens: 128

No credentials, prompt text, response text, or request headers were retained.

PR Type

  • Bug Fix

Relevant issues

Fixes #1185

Checklist

  • I understand the code I am submitting.
  • I have added unit tests that prove my fix/feature works
  • I have run this code locally and verified it fixes the issue.
  • New and existing tests pass locally
  • Documentation was updated where necessary
  • I have read and followed the contribution guidelines
  • AI Usage:
    • No AI was used.
    • AI was used for drafting/refactoring.
    • This is fully AI-generated.

AI Usage Information

  • AI Model used: OpenAI GPT-5.6 Sol
  • AI Developer Tool used: Codex
  • Any other info you'd like to share: The scope and acceptance criteria were
    directed by the contributor. Codex implemented the change, added the
    regression test, ran the local validation, and drafted this PR description.
    The contributor reviewed the patch before submission.

When answering questions by the reviewer, please respond yourself, do not copy/paste the reviewer comments into an AI system and paste back its answer. We want to discuss with you, not your AI :)

  • I am an AI Agent filling out this form

Summary by CodeRabbit

  • Bug Fixes
    • Improved Minimax streaming reliability by preserving usage-only updates alongside response content.
    • Unrelated or empty streaming chunks are now ignored, preventing unusable data from appearing in streamed responses.
    • Usage token information remains available in terminal updates, even when those updates contain no response text.

@coderabbitai

coderabbitai Bot commented Aug 7, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 85e9b048-a5f2-4458-8134-39dc7033192d

📥 Commits

Reviewing files that changed from the base of the PR and between 66c6fa2 and 269e23b.

📒 Files selected for processing (2)
  • src/any_llm/providers/minimax/minimax.py
  • tests/unit/providers/test_minimax_provider.py

Walkthrough

Minimax streaming now preserves usage-only chunks and filters unrelated objects. Unit tests verify content retention, empty-choice terminal chunks, and prompt, completion, and total token counts.

Changes

Minimax streaming usage

Layer / File(s) Summary
Preserve usage-bearing chunks
src/any_llm/providers/minimax/minimax.py
The streaming filter retains typed chunks with deltas or usage data. It clears unusable choices before conversion.
Validate streamed usage
tests/unit/providers/test_minimax_provider.py
Async tests verify content chunks, unrelated objects, empty chunks, and usage-only terminal chunks with their token counts.

Possibly related PRs

Suggested labels: 1.21.0

Suggested reviewers: tbille

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly identifies the MiniMax bug fix that preserves usage-only streaming chunks.
Description check ✅ Passed The description follows the template, explains the fix, links issue #1185, records validation, and completes the checklist.
Linked Issues check ✅ Passed The implementation and tests satisfy issue #1185 by preserving usage-only chunks while filtering unusable chunks and retaining token counts.
Out of Scope Changes check ✅ Passed The changes are limited to MiniMax streaming conversion and related regression tests, with no unrelated scope identified.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@zhuziqi97

Copy link
Copy Markdown
Contributor Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Aug 7, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@src/any_llm/providers/minimax/minimax.py`:
- Around line 47-48: Flatten the nested conditions in the chunk-processing logic
by combining the OpenAIChatCompletionChunk type check with its
choices/delta-or-usage predicate into a single if statement, preserving the
existing filter behavior. Run the repository pre-commit checks and confirm Ruff
lint and formatting pass.

In `@tests/unit/providers/test_minimax_provider.py`:
- Around line 139-145: Update the test assertions for the first content chunk in
the Minimax provider conversion to verify that result[0].usage is None, while
preserving the existing usage-only chunk assertions for result[1].
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 2915796c-fe6e-46e7-8d43-6eead6294b00

📥 Commits

Reviewing files that changed from the base of the PR and between b4bb154 and 51e66b7.

📒 Files selected for processing (2)
  • src/any_llm/providers/minimax/minimax.py
  • tests/unit/providers/test_minimax_provider.py

Comment thread src/any_llm/providers/minimax/minimax.py Outdated
Comment thread tests/unit/providers/test_minimax_provider.py
@zhuziqi97

Copy link
Copy Markdown
Contributor Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Aug 7, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@zhuziqi97
zhuziqi97 marked this pull request as ready for review August 7, 2026 07:33
The usage guard let Minimax's terminal chunk through whole. That chunk
carries a `message` instead of a `delta`, so the OpenAI SDK parses
`delta` as None and chunk validation raises, which is the failure mozilla-ai#657
introduced this filter to avoid.

Yield the usage while dropping the unusable choices, and cover the real
wire shape with a regression test built through the SDK deserializer.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@njbrake
njbrake temporarily deployed to integration-tests August 10, 2026 20:28 — with GitHub Actions Inactive
@njbrake njbrake added the run-integration-tests Put this label on a PR to trigger the integration test suite: works with forks label Aug 10, 2026
@github-actions github-actions Bot removed the run-integration-tests Put this label on a PR to trigger the integration test suite: works with forks label Aug 10, 2026
@codecov

codecov Bot commented Aug 10, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 66.66667% with 2 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
src/any_llm/providers/minimax/minimax.py 66.66% 1 Missing and 1 partial ⚠️
Files with missing lines Coverage Δ
src/any_llm/providers/minimax/minimax.py 88.23% <66.66%> (-1.35%) ⬇️

... and 22 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@njbrake njbrake left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Thank you for the contribution!

@njbrake
njbrake merged commit ac69e42 into mozilla-ai:main Aug 11, 2026
16 of 19 checks passed
@github-actions github-actions Bot added the 1.25.0 Included in release 1.25.0 label Aug 11, 2026
HareeshBahuleyan added a commit that referenced this pull request Sep 14, 2026
…unk (#1392)

## Description

Native Anthropic streaming reports token usage on the `message_stop`
event, which arrives after the `message_delta` event that carries
`stop_reason`. The chunk converter attached a synthetic choice (`index
0`, empty delta, `finish_reason: None`) to every chunk, including that
usage-only one. OpenAI-compatible providers put final usage on a
trailing chunk with `choices: []` (OpenAI documents this for
`stream_options.include_usage`), and the MiniMax fix in #1249 aligned
MiniMax to the same shape. Code that picks up usage with `if not
chunk.choices` therefore worked for OpenAI and silently saw no usage for
Anthropic.

Observed on `main` (last two chunks of a streamed `claude-sonnet-4-6`
reply, null fields dropped):

```
8 {'choices': [{'delta': {}, 'finish_reason': 'stop', 'index': 0}]}
9 {'choices': [{'delta': {}, 'index': 0}], 'usage': {'completion_tokens': 5, 'prompt_tokens': 13, 'total_tokens': 18}}
```

With this change chunk 9 becomes `{'choices': [], 'usage': {...}}`,
identical in shape to OpenAI's trailing usage chunk. Chunks 1 to 8 are
unchanged.

Change: `_create_openai_chunk_from_anthropic_chunk` returns from the
`MessageStopEvent` branch before the choice is appended. The stop event
carries no delta and no stop reason, so nothing is lost.

Tests: the two existing `message_stop` usage tests now assert `choices
== []`; two new converter tests pin that the usage chunk arrives after
the `finish_reason` chunk with empty choices, and that a raw
`message_stop` without an accumulated message yields neither choices nor
usage. A third test drives `acompletion(stream=True)` through a fake SDK
message stream and reads the chunks the way an OpenAI-style consumer
does (content and `finish_reason` from chunks with choices, usage from
the chunk without). On `main` that test fails with `usage is None`; text
and `finish_reason` already arrive correctly.

Validation:

- `uv run pytest tests/unit`: 2435 passed, 69 skipped
- `uv run pre-commit run --all-files`: clean (ruff, ruff format, mypy
strict, codespell)
- Live streams against Anthropic (`claude-sonnet-4-6`) and OpenAI
(`gpt-5-nano`, `include_usage`) printed chunk by chunk; usage chunk
shapes match after the fix
- `uv run pytest tests/integration -k anthropic` with real keys: 23
passed, 7 skipped. The skips are the existing capability skips
(Anthropic has no batch, responses, moderation, or embeddings endpoints;
one test targets Claude on Bedrock only, see #1184)

Not changed, same pattern: Bedrock also attaches a synthetic choice to
its `metadata` usage chunk (`src/any_llm/providers/bedrock/utils.py`).
Gemini attaches usage to every content chunk and never emits a
usage-only chunk.

## PR Type

- 🐛 Bug Fix

## Relevant issues

None open. Same shape as the MiniMax fix in #1249.

## Checklist
- [x] I understand the code I am submitting.
- [x] I have added unit tests that prove my fix/feature works
- [x] I have run this code locally and verified it fixes the issue.
- [x] New and existing tests pass locally
- [x] Documentation was updated where necessary
- [x] I have read and followed the [contribution
guidelines](https://github.com/mozilla-ai/any-llm/blob/main/CONTRIBUTING.md)
- [x] **AI Usage:**
    - [ ] No AI was used.
    - [ ] AI was used for drafting/refactoring.
    - [x] This is fully AI-generated.

## AI Usage Information

- AI Model used: Claude Fable 5.1 (`claude-fable-5-1`)
- AI Developer Tool used: Claude Code
- Any other info you'd like to share: The bug was reported by an
automated agent; the fix, tests, and this description were produced by
Claude Code and reviewed by the submitter, who will answer reviewer
questions personally.

- [x] I am an AI Agent filling out this form (check box if true)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Corrected streaming responses so the final usage-only chunk no longer
includes an empty choice.
* Improved compatibility with OpenAI-style streaming consumers by
delivering usage after the finish reason.
* Ensured message-stop events produce the expected empty choices list,
including when no message content has been accumulated.
* Standardised streaming completion behaviour so consumers receive text,
a single stop reason, and final usage in the correct order.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Hareesh <hareeshbahuleyan@gmail.com>

This branch was previously deployed

1 inactive deployment
integration-tests — 269e23be Deployed Aug 10, 2026 by njbrake via run-docs-tests #2352
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

1.25.0 Included in release 1.25.0

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Streamed messages through minimax still meter zero tokens (usage-only chunk dropped by chunk filter)

2 participants