Skip to content

fix(anthropic): leave choices empty on the usage-only message_stop chunk - #1

Draft
ranzhh wants to merge 14 commits into
mainfrom
feature/anthropic-usage-chunk-empty-choices
Draft

ranzhh wants to merge 14 commits into
mainfrom
feature/anthropic-usage-chunk-empty-choices

Conversation

@ranzhh

@ranzhh ranzhh commented Sep 9, 2026 •

Copy link
Copy Markdown
Owner

Description

Native Anthropic streaming reports token usage on the message_stop event, which arrives after the message_delta event that carries stop_reason. The chunk converter attached a synthetic choice (index 0, empty delta, finish_reason: None) to every chunk, including that usage-only one. OpenAI-compatible providers put final usage on a trailing chunk with choices: [] (OpenAI documents this for stream_options.include_usage), and the MiniMax fix in mozilla-ai#1249 aligned MiniMax to the same shape. Code that picks up usage with if not chunk.choices therefore worked for OpenAI and silently saw no usage for Anthropic.

Observed on main (last two chunks of a streamed claude-sonnet-4-6 reply, null fields dropped):

8 {'choices': [{'delta': {}, 'finish_reason': 'stop', 'index': 0}]}
9 {'choices': [{'delta': {}, 'index': 0}], 'usage': {'completion_tokens': 5, 'prompt_tokens': 13, 'total_tokens': 18}}

With this change chunk 9 becomes {'choices': [], 'usage': {...}}, identical in shape to OpenAI's trailing usage chunk. Chunks 1 to 8 are unchanged.

Change: _create_openai_chunk_from_anthropic_chunk returns from the MessageStopEvent branch before the choice is appended. The stop event carries no delta and no stop reason, so nothing is lost.

Tests: the two existing message_stop usage tests now assert choices == []; two new converter tests pin that the usage chunk arrives after the finish_reason chunk with empty choices, and that a raw message_stop without an accumulated message yields neither choices nor usage. A third test drives acompletion(stream=True) through a fake SDK message stream and reads the chunks the way an OpenAI-style consumer does (content and finish_reason from chunks with choices, usage from the chunk without). On main that test fails with usage is None; text and finish_reason already arrive correctly.

Validation:

  • uv run pytest tests/unit: 2435 passed, 69 skipped
  • uv run pre-commit run --all-files: clean (ruff, ruff format, mypy strict, codespell)
  • Live streams against Anthropic (claude-sonnet-4-6) and OpenAI (gpt-5-nano, include_usage) printed chunk by chunk; usage chunk shapes match after the fix
  • uv run pytest tests/integration -k anthropic with real keys: 23 passed, 7 skipped. The skips are the existing capability skips (Anthropic has no batch, responses, moderation, or embeddings endpoints; one test targets Claude on Bedrock only, see [BUG] output_format not working with Bedrock > Claude mozilla-ai/any-llm#1184)

Not changed, same pattern: Bedrock also attaches a synthetic choice to its metadata usage chunk (src/any_llm/providers/bedrock/utils.py). Gemini attaches usage to every content chunk and never emits a usage-only chunk.

PR Type

  • 🐛 Bug Fix

Relevant issues

None open. Same shape as the MiniMax fix in mozilla-ai#1249.

Checklist

  • I understand the code I am submitting.
  • I have added unit tests that prove my fix/feature works
  • I have run this code locally and verified it fixes the issue.
  • New and existing tests pass locally
  • Documentation was updated where necessary
  • I have read and followed the contribution guidelines
  • AI Usage:
    • No AI was used.
    • AI was used for drafting/refactoring.
    • This is fully AI-generated.

AI Usage Information

  • AI Model used: Claude Fable 5.1 (claude-fable-5-1)

  • AI Developer Tool used: Claude Code

  • Any other info you'd like to share: The bug was reported by an automated agent; the fix, tests, and this description were produced by Claude Code and reviewed by the submitter, who will answer reviewer questions personally.

  • I am an AI Agent filling out this form (check box if true)

ranzhh and others added 14 commits September 8, 2026 12:20
Native Anthropic streaming reports token usage on the message_stop chunk,
after the chunk that carries finish_reason. The converter attached a
synthetic empty-delta choice to that chunk, so consumers that detect the
trailing usage chunk by an empty choices list, as they can for every
OpenAI-compatible provider, saw no usage for Anthropic.

Return early from the message_stop branch with choices left empty, matching
the shape MiniMax was aligned to in mozilla-ai#1249.
…zilla-ai#1380)

## Description

Validate and normalize V4 Chat extensions while preserving caller-owned
dictionaries and explicit wire overrides.

## Verification

2444 passed, 69 skipped, 4 warnings; all-files pre-commit passed
locally.

No live-provider tests were run. Upstream CI awaits maintainer approval.

## PR Type

Provider contract and tests.

## Checklist

- [ ] I understand the code I am submitting. (Human author confirmation
pending.)
- [x] I have added unit tests that prove my fix/feature works
- [x] I have run this code locally and verified it fixes the issue.
(Offline tests.)
- [x] New and existing tests pass locally
- [x] Documentation was updated where necessary
- [x] I have read and followed the [contribution
guidelines](https://github.com/mozilla-ai/any-llm/blob/main/CONTRIBUTING.md)
- [x] **AI Usage:**
    - [ ] No AI was used.
    - [x] AI was used for drafting/refactoring.
    - [ ] This is fully AI-generated.

## AI Usage Information

- AI Model used: GPT-6 Astra Medium
- AI Developer Tool used: Amp
- [x] I am an AI Agent filling out this form (check box if true)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added support for mapping reasoning effort levels to DeepSeek’s
thinking modes.
* Added validation for DeepSeek user identifiers before requests are
sent.
* Preserved reasoning content across tool-assisted conversations,
including assistant responses without tool calls.

* **Bug Fixes**
* Improved handling of provider defaults when reasoning effort is
unspecified or automatic.
* Unsupported reasoning effort values now return a clear request error.
* Improved request-field filtering and settings merging without
modifying caller-provided values.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Hareesh <hareeshbahuleyan@gmail.com>
## Description

### Why

Together's agent-loop integration test intermittently returns HTTP 400
input validation errors with `openai/gpt-oss-20b`. Together documents
Qwen 3.5 9B for agentic multi-step function calling.

### What changed

Use `Qwen/Qwen3.5-9B` for Together's general integration tests while
keeping GPT-OSS for reasoning coverage.

### Notes

- `uv run pytest tests/unit -q`: 2,447 passed, 69 skipped.
- Hooks for `tests/conftest.py` pass. Full pre-commit is blocked locally
by an `httpx` versus `httpx2` mypy type conflict in unchanged
`tests/unit/providers/test_openai_exceptions.py:188`.
- Targeted live Together CI run:
`test_agent_loop_sequential_tool_calls[together]` passed on its first
attempt in 15.93 seconds ([run
34337514731](https://github.com/mozilla-ai/any-llm/actions/runs/34337514731)).

## PR Type

- 🐛 Bug Fix

## Relevant issues

CI run: https://github.com/mozilla-ai/any-llm/actions/runs/34332374589

## Checklist

- [x] I understand the code I am submitting.
- [ ] I have added unit tests that prove my fix/feature works
- [ ] I have run this code locally and verified it fixes the issue.
- [x] New and existing tests pass locally
- [x] Documentation was updated where necessary
- [x] I have read and followed the [contribution
guidelines](https://github.com/mozilla-ai/any-llm/blob/main/CONTRIBUTING.md)
- [x] **AI Usage:**
    - [ ] No AI was used.
    - [ ] AI was used for drafting/refactoring.
    - [x] This is fully AI-generated.

## AI Usage Information

- AI Model used: OpenAI GPT-5.6 Sol
- AI Developer Tool used: Pi
- Any other info you'd like to share: The model selection was verified
against Together's official model catalog and agentic function-calling
documentation.

When answering questions by the reviewer, please respond yourself, do
not copy/paste the reviewer comments into an AI system and paste back
its answer. We want to discuss with you, not my AI :)

- [x] I am an AI Agent filling out this form (check box if true)



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

- **Tests**
- Updated the test configuration to use the Qwen 3.5 9B model for
Together AI provider coverage.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Drive acompletion(stream=True) through a fake SDK message stream and read
the chunks the way callers written against OpenAI do: content and
finish_reason from chunks with choices, usage from the chunk without. On
main the loop ends with usage still None.
## Description

Add native structured-output support to Otari's Anthropic-compatible
Messages path.

- Route Pydantic and dataclass output types through Otari's native
`message()` API using Anthropic-compatible transformed JSON schemas.
- Normalize raw `output_config` dictionaries before forwarding them to
Otari.
- Preserve `context_management`, `betas`, and `cache_control` when
structured output is requested.
- Support structured-output streaming for providers that opt into the
capability.
- Add the same native streaming capability to the Anthropic provider
path.
- Update Messages API documentation and add streaming and non-streaming
coverage.

Previously, Otari structured-output requests fell back through the
Messages-to-Completions bridge. That path could not retain
Anthropic-specific fields and rejected structured output combined with
context management or beta features.

## PR Type

- 🆕 New Feature

## Relevant issues

Related to mozilla-ai/octonous#4903.

## Verification

- `uv run pytest tests/unit/providers/test_otari_provider.py -q --reruns
0`: 49 passed
- `uv run pytest tests/unit/providers/test_anthropic_messages.py -q -k
'output_format or output_config' --reruns 0`: 7 passed
- Commit-time lint, formatting, mypy, codespell, and repository hygiene
hooks passed.
- The existing local Anthropic `httpx`/`httpx2` full-suite mismatch is
being handled separately in mozilla-ai#1371.
- Live Otari integration testing was not run because Otari credentials
were unavailable locally.

#### Checklist

- [x] I understand the code I am submitting.
- [x] I have added unit tests that prove my fix/feature works
- [x] I have run this code locally and verified it fixes the issue.
- [ ] New and existing tests pass locally
- [x] Documentation was updated where necessary
- [x] I have read and followed the [contribution
guidelines](https://github.com/mozilla-ai/any-llm/blob/main/CONTRIBUTING.md)
- [x] **AI Usage:**
    - [ ] No AI was used.
    - [ ] AI was used for drafting/refactoring.
    - [x] This is fully AI-generated.

## AI Usage Information

- AI Model used: GPT-5 Codex
- AI Developer Tool used: Codex
- Any other info you'd like to share: The implementation was developed
test-first and verified with focused provider suites. The separate
Anthropic dependency mismatch discovered during full-suite verification
is tracked independently in mozilla-ai#1371.

When answering questions by the reviewer, please respond yourself, do
not copy/paste the reviewer comments into an AI system and paste back
the answer. We want to discuss with you, not your AI :)

- [x] I am an AI Agent filling out this form (check box if true)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added support for streaming schema-constrained Messages responses
where supported.
- Anthropic and Otari providers now support structured output during
streaming.
- Otari structured-output requests use the native Messages endpoint and
support typed schemas and configuration dictionaries.
- Schema-less output configuration dictionaries now return regular
message responses.

- **Documentation**
- Clarified structured-output and streaming behaviour across the API
documentation.

- **Tests**
- Added coverage for streaming structured output, provider-specific
options, schema conversion, and native endpoint handling.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
## Description

Keep Gemini thinking capability handling while preserving the existing
reasoning-effort interface and wire defaults.

Compatibility is preserved for `minimal=256`, `xhigh`/`max` budget
aliases (`32768`) and level aliases (`HIGH`). `None` and `none` still
send `include_thoughts=false`, which hides summaries without disabling
thinking; `auto` leaves caller-supplied configuration untouched.
Unlisted/custom model IDs retain permissive version-based routing.
Model-specific budget limits remain provider-validated rather than
globally clamped.

Deliberate changes: known Gemini 3 models use native thinking levels,
including known models below 3.5 and their numeric revisions. Gemini 3.1
Pro maps `minimal` to `low`. Known unsupported combinations raise
`UnsupportedParameterError` locally: `minimal` on 3.8/3.7 Flash, and
`low`/`medium` on 3.1 Flash Lite Image. Google recommends thinking
levels for Gemini 3 while retaining budget compatibility; these changes
can affect reasoning allocation compared with the earlier budget
mapping. See [Google's thinking controls and model
limits](https://ai.google.dev/gemini-api/docs/generate-content/thinking).

## Verification

2441 passed, 69 skipped, 4 warnings; all-files pre-commit passed
locally.

No live-provider tests were run. Upstream CI awaits maintainer approval.

## PR Type

Provider contract and tests.

## Checklist

- [ ] I understand the code I am submitting. (Human author confirmation
pending.)
- [x] I have added unit tests that prove my fix/feature works
- [x] I have run this code locally and verified it fixes the issue.
(Offline tests.)
- [x] New and existing tests pass locally
- [x] Documentation was updated where necessary
- [x] I have read and followed the [contribution
guidelines](https://github.com/mozilla-ai/any-llm/blob/main/CONTRIBUTING.md)
- [x] **AI Usage:**
    - [ ] No AI was used.
    - [x] AI was used for drafting/refactoring.
    - [ ] This is fully AI-generated.

## AI Usage Information

- AI Model used: GPT-6 Astra Medium
- AI Developer Tool used: Amp
- [x] I am an AI Agent filling out this form (check box if true)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Improved Gemini reasoning controls with model-specific thinking levels
and effort settings.
- Added support for Gemini 3.1 Pro’s minimal reasoning option and budget
limits for Gemini 2.5 models.
- Expanded response-format handling for structured data, JSON schemas,
JSON objects, plain text, and unset formats.

- **Bug Fixes**
- Improved validation for unsupported model and reasoning-effort
combinations.
- Preserved distinct handling for disabled, automatic, and unspecified
reasoning settings.

- **Chores**
- Updated optional Gemini and Vertex AI integrations to require a newer
Google GenAI SDK.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Hareesh <hareeshbahuleyan@gmail.com>
## Description

Normalize Ollama completion finish reasons before constructing any-llm's
response types. An SDK response with a missing or null `done_reason`
currently raises `ValidationError` in nonstreaming conversion; `load`
and `unload` raise in both converters. These lifecycle responses are
documented as completed, empty chat responses.

Use `stop` for missing terminal reasons, lifecycle reasons, and unknown
reasons, following the existing provider fallback convention. Preserve
all previously accepted OpenAI finish reasons and nonstreaming tool-call
precedence. In streams, a chunk with no reason remains unfinished unless
`done` is true, in which case it receives the terminal fallback.

Tests exercise the real Ollama SDK's JSON and NDJSON decoding through a
mocked HTTP transport, plus typed SDK tool-call responses. On the base
commit, 10 new cases fail and 15 pass; with the fix, all 79 Ollama unit
tests pass. The full unit suite passes with 2,456 passed and 69 existing
skips. Repository-wide pre-commit checks pass, including mypy.

Live validation now passes on Ollama 0.33.3 with llama3.2:1b: all four
existing async/parallel/streaming integration tests passed with no
skips. Real stop and length results survive both conversions; native SDK
load/unload responses fail on the original converters and normalize to
stop with this fix. Those lifecycle requests use the native SDK because
any-llm does not accept an empty messages list. Missing/unknown reasons
remain controlled-transport coverage, not observed live responses. No
timestamp handling changes are included.

## PR Type

- 🐛 Bug Fix

## Relevant issues

None found for this normalization defect in the live open issue and PR
checks.

## Checklist

- [ ] I understand the code I am submitting. (Human author review
pending.)
- [x] I have added unit tests that prove my fix/feature works
- [x] I have run this code locally and verified it fixes the issue.
- [x] New and existing tests pass locally (2,456 unit tests; all four
selected live Ollama integration tests pass)
- [x] Documentation was updated where necessary (normalization semantics
documented in the helper docstring)
- [x] I have read and followed the [contribution
guidelines](https://github.com/mozilla-ai/any-llm/blob/main/CONTRIBUTING.md)
- [x] **AI Usage:**
    - [ ] No AI was used.
    - [ ] AI was used for drafting/refactoring.
    - [x] This is fully AI-generated.

## AI Usage Information

- AI Model used: GPT-6
- AI Developer Tool used: Codex
- Any other info you'd like to share: Codex researched the current
source and documentation, implemented the fix and regression tests, ran
validation, and drafted this body. Human author review is pending.
Controlled SDK transport tests are supplemented by real local inference
and native SDK lifecycle probes; no user/private data was used.

When answering questions by the reviewer, please respond yourself, do
not copy/paste the reviewer comments into an AI system and paste back
its answer. We want to discuss with you, not your AI :)

- [x] I am an AI Agent filling out this form (check box if true)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Normalised completion and streaming finish reasons from Ollama for
consistent results.
  * Unknown or missing terminal reasons now default to `stop`.
* Tool-call completions now report the appropriate `tool_calls` finish
reason.
* **Tests**
* Added coverage for streaming, non-streaming, and tool-call
finish-reason handling.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
## Description

Gemini media responses currently discard `inline_data` parts during
conversion to the OpenAI-compatible response shape. This makes
image-only responses appear to have no choices and loses media from
text-plus-image responses.

This change converts inline media to data URLs in both non-streaming and
streaming response paths, exposes them as `images`, and emits a choice
when a candidate contains media without text or tool calls.

## PR Type

- 🐛 Bug Fix

## Relevant issues

Fixes mozilla-ai#1295

## Testing

- `python -m pytest tests/unit/providers/test_gemini_provider.py -q -k
'skips_parts or image_only'`
- `ruff check src/any_llm/providers/gemini/utils.py
tests/unit/providers/test_gemini_provider.py`
- `git diff --check`

The focused tests pass locally. The full provider module includes
existing integration-style cases that exceeded the local command
timeout.

## Checklist

- [x] I understand the code I am submitting.
- [x] I have added unit tests that prove my fix/feature works
- [x] I have run this code locally and verified it fixes the issue.
- [x] New and existing tests pass locally
- [ ] Documentation was updated where necessary
- [x] I have read and followed the [contribution
guidelines](https://github.com/mozilla-ai/any-llm/blob/main/CONTRIBUTING.md)
- [x] **AI Usage:**
    - [ ] No AI was used.
    - [ ] AI was used for drafting/refactoring.
    - [x] This is fully AI-generated.

## AI Usage Information

- AI Model used: GPT-5
- AI Developer Tool used: Codex
- Any other info you'd like to share: The patch and tests were reviewed
locally before submission.

- [x] I am an AI Agent filling out this form (check box if true)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Gemini inline images are now returned as OpenAI-compatible `data:`
URLs.
  - Image content is supported in both standard and streaming responses.
  - Image-only responses now produce a valid response choice.

- **Bug Fixes**
- Inline image parts are no longer omitted when converting Gemini
responses.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
…ozilla-ai#1358)

## Description

The Bedrock provider maps only three of the Converse API's nine
`stopReason` values, so a guardrail-blocked or context-overflowed
response reaches the caller as an ordinary `finish_reason="stop"`.

`src/any_llm/providers/bedrock/utils.py:517` (non-streaming):

```python
finish_reason: Literal["stop", "length"] = "length" if stop_reason == "max_tokens" else "stop"
```

and `:611-617` (streaming) handles `max_tokens` and `tool_use` and sends
everything else to `"stop"`.

The visible failure is the structured-output guard in
`src/any_llm/any_llm.py:798-806`. When a Bedrock Guardrail blocks a
response, Converse returns `stopReason="guardrail_intervened"` with the
guardrail's blocked-message text as content. Because that arrives as
`finish_reason="stop"`, the `content_filter` branch never fires and the
guardrail's prose is handed to `parse_json_content()` instead, so the
caller gets a pydantic `ValidationError` from an unrelated layer rather
than the `ContentFilterFinishReasonError` this library raises for every
other provider. `content_filtered` behaves the same way, and
`model_context_window_exceeded` reports as `"stop"` rather than
`"length"`.

This is the same fix already merged for gemini (mozilla-ai#1202), zai (mozilla-ai#1204),
cohere (mozilla-ai#1301) and anthropic (mozilla-ai#1306/mozilla-ai#1328). Bedrock was the last
provider still hardcoding a partial mapping, so this follows
`ANTHROPIC_STOP_REASON_TO_FINISH_REASON` in shape and naming:

```python
BEDROCK_STOP_REASON_TO_FINISH_REASON: dict[str, _FinishReason] = {
    "end_turn": "stop",
    "max_tokens": "length",
    "model_context_window_exceeded": "length",
    "tool_use": "tool_calls",
    "content_filtered": "content_filter",
    "guardrail_intervened": "content_filter",
}
```

`stop_sequence`, `malformed_model_output` and `malformed_tool_use` have
no OpenAI counterpart and keep falling through to the `"stop"` default,
as do any values a future service model adds. Both call sites now go
through one `_map_stop_reason()` helper, which also lets the `cast` at
the non-streaming call site go away.

The nine-value enum is not taken from the AWS docs prose. It is read out
of the installed botocore service model (`bedrock-runtime`, shape
`StopReason`), which is the same source the SDK validates against:

```python
>>> import botocore.session
>>> botocore.session.Session().get_service_model("bedrock-runtime").shape_for("StopReason").enum
['end_turn', 'tool_use', 'max_tokens', 'stop_sequence', 'guardrail_intervened', 'content_filtered',
 'malformed_model_output', 'malformed_tool_use', 'model_context_window_exceeded']
```

The tests read the enum from that same service model and parametrize
over it, so a botocore upgrade that adds a stop reason fails the suite
instead of silently defaulting the new reason to `"stop"`.

## Reproduction

```python
import asyncio
from unittest.mock import Mock

from pydantic import BaseModel

from any_llm.exceptions import ContentFilterFinishReasonError
from any_llm.providers.bedrock import BedrockProvider


class City(BaseModel):
    name: str


# What Converse returns when a Bedrock Guardrail blocks the response.
blocked = {
    "output": {"message": {"content": [{"text": "Sorry, I cannot answer that."}]}},
    "stopReason": "guardrail_intervened",
}

client = Mock()
client.converse.return_value = blocked
provider = BedrockProvider(client=client)

print("finish_reason:", provider._convert_completion_response(blocked).choices[0].finish_reason)

try:
    asyncio.run(
        provider.acompletion(
            model="us.anthropic.claude-sonnet-4-20250514-v1:0",
            messages=[{"role": "user", "content": "Hello"}],
            response_format=City,
        )
    )
except Exception as exc:
    print("raised:", type(exc).__name__)
```

Before:

```
finish_reason: stop
raised: ValidationError
```

After:

```
finish_reason: content_filter
raised: ContentFilterFinishReasonError
```

## Tests

Added to `tests/unit/providers/test_aws_provider.py`:

- `test_convert_response_maps_every_bedrock_stop_reason` and
`test_streaming_chunk_maps_every_bedrock_stop_reason`, parametrized over
the full botocore `StopReason` enum, asserting both paths agree.
- `test_convert_response_without_stop_reason_finishes_as_stop`.
-
`test_guardrail_blocked_structured_output_raises_content_filter_error`,
the end-to-end case above.

On the unfixed tree these fail:

```
FAILED test_convert_response_maps_every_bedrock_stop_reason[tool_use]
FAILED test_convert_response_maps_every_bedrock_stop_reason[guardrail_intervened]
FAILED test_convert_response_maps_every_bedrock_stop_reason[content_filtered]
FAILED test_convert_response_maps_every_bedrock_stop_reason[model_context_window_exceeded]
FAILED test_streaming_chunk_maps_every_bedrock_stop_reason[guardrail_intervened]
FAILED test_streaming_chunk_maps_every_bedrock_stop_reason[content_filtered]
FAILED test_streaming_chunk_maps_every_bedrock_stop_reason[model_context_window_exceeded]
FAILED test_guardrail_blocked_structured_output_raises_content_filter_error
8 failed, 12 passed
```

With the fix, `uv run pytest tests/unit/providers/test_aws_provider.py`
→ 116 passed.

The `[tool_use]` case is the one behaviour change beyond the three
broken reasons: a `stopReason="tool_use"` response that carries no
`toolUse` block now reports `"tool_calls"` instead of `"stop"`. Both
real `tool_use` shapes (a genuine tool call, and the synthetic
`any_llm_structured_output` unwrap) return earlier in
`_convert_response` and are untouched.

## PR Type

- 🐛 Bug Fix

## Relevant issues

No open issue. Same fix as mozilla-ai#1202 / mozilla-ai#1204 / mozilla-ai#1301 / mozilla-ai#1328, applied to the
remaining provider.

## Checklist

- [x] I understand the code I am submitting.
- [x] I have added unit tests that prove my fix/feature works
- [x] I have run this code locally and verified it fixes the issue.
- [x] New and existing tests pass locally
- [x] Documentation was updated where necessary (no user-facing API
changed)
- [x] I have read and followed the [contribution
guidelines](https://github.com/mozilla-ai/any-llm/blob/main/CONTRIBUTING.md)
- [x] **AI Usage:**
    - [ ] No AI was used.
    - [ ] AI was used for drafting/refactoring.
    - [x] This is fully AI-generated.

## Verification

- `uv run pre-commit run --all-files` (ruff, ruff-format, mypy,
codespell) clean.
- `uv run pytest tests/unit/providers/test_aws_provider.py` → 116
passed.
- `uv run pytest tests/unit` → 2308 passed, 69 skipped, 16 failed. All
16 failures are in `test_anthropic_messages.py` /
`test_anthropic_provider.py` and reproduce identically on an unmodified
`main` in this environment: `TypeError: Invalid 'http_client' argument;
Expected an instance of httpx2.AsyncClient but got <class
'httpx.AsyncClient'>`, from what a local `uv sync --all-extras -U`
resolves. Nothing bedrock-related.
- No integration run: I do not have AWS Bedrock credentials, and the two
Bedrock paths this touches are pure response converters exercised by the
unit tests above against service-model-sourced stop reasons.

## AI Usage Information

- AI Model used: Claude Opus 5
- AI Developer Tool used: Claude Code
- Any other info you'd like to share: The stop-reason enum was read from
the installed botocore service model rather than the AWS docs, and the
tests read it from the same place so the list cannot drift.

- [x] I am an AI Agent filling out this form (check box if true)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Improved Amazon Bedrock response handling for length limits, tool
calls, and content-filter outcomes.
* Standardised finish reasons across streaming and non-streaming
responses.
  * Unknown or unsupported stop reasons now safely default to `stop`.
* Structured-output requests blocked by a Bedrock guardrail now report a
content-filter error correctly.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…a-ai#1327)

## Description

The agent-loop integration tests hand-add a `name` field to every tool
message they build. OpenAI tool messages have no such key
(`ChatCompletionToolMessageParam` is `role`/`content`/`tool_call_id`;
`name` belongs to the deprecated `role="function"` shape), so the suite
exercised a shape no real caller sends, and the hand-added name
short-circuited the resolution path mozilla-ai#1318 fixed. Deleting the two lines
makes the loops run on the real shape. To be precise about what that
buys: it removes a false-safe path rather than turning these tests into
a regression guard for mozilla-ai#1318 itself, since the bug was a wrong answer,
not an error, and these tests do not assert answer content. The
`name`-first and fallback branches stay pinned by unit tests (gemini,
ollama, cohere).

The key was never load-bearing anywhere: no converter requires it,
cohere's `_patch_messages` exists to strip it, and the OpenAI-compatible
providers were passing it through to their APIs verbatim. gemini
resolves the name from `tool_call_id` since mozilla-ai#1318 (ollama has the same
map, but sits in the local-provider set this file skips). Live run of
`tests/integration/test_agent_loop.py` on this branch with the keys
available here: 12 passed, 70 skipped; the six failures (azureopenai,
fireworks, sagemaker on both loop tests) fail identically on
`upstream/main` with the same environment, for billing/endpoint reasons
unrelated to this change.

## PR Type

- 🧪 Tests

## Relevant issues

Follow-up to mozilla-ai#1318 (second bullet of the review's follow-up notes).

## Checklist
<!-- If this checklist is deleted from the PR submission it will be
immediately closed -->
- [x] I understand the code I am submitting.
- [ ] I have added unit tests that prove my fix/feature works (test-only
change; the edited tests are the deliverable)
- [x] I have run this code locally and verified it fixes the issue.
- [x] New and existing tests pass locally
- [x] Documentation was updated where necessary
- [x] I have read and followed the [contribution
guidelines](https://github.com/mozilla-ai/any-llm/blob/main/CONTRIBUTING.md)
- [x] **AI Usage:**
    - [ ] No AI was used.
    - [x] AI was used for drafting/refactoring.
    - [ ] This is fully AI-generated.

## AI Usage Information

- AI Model used: Claude (Fable 5)
- AI Developer Tool used: Claude Code
- Any other info you'd like to share:

- [ ] I am an AI Agent filling out this form (check box if true)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Tests**
* Improved integration coverage for agent conversations involving tool
results.
* Added checks that final responses accurately reflect returned
information.
* Strengthened validation that conversations complete successfully after
tool interactions.
* Updated test coverage for parallel and sequential tool calls,
including streamlined tool-result messages.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
…ozilla-ai#1391)

## Description

The agent-loop tests introduced in mozilla-ai#1327 reject valid conversations: a
model may request Paris and London weather on separate turns, and a
final answer may contain an empty tool-call list. The old sequential
loop also silently passed when it exhausted its iteration budget.

Both scenarios now share a bounded loop that executes intermediate tool
calls, checks that the required calls occurred, and recognizes both
`None` and `[]` as completion. Final answers must still reference the
weather result, and incomplete conversations fail at the iteration
limit. The multiple-call test accepts calls across turns; tool selection
remains unchanged. Both tool-message shapes, with and without `name`,
remain covered.

Unit tests cover calls across turns, batched calls, both completion
shapes, premature answers, and iteration exhaustion using standard
mocks.

## Verification

- Unit suite: 2,541 passed, 69 skipped. Focused helper suite: 7 passed.
- Live agent-loop tests with local credentials: both scenarios passed
for Anthropic, Fireworks, Gemini, Hugging Face, Mistral, OpenAI, and
xAI, covering 14 provider cases. Gemini's multiple-call case initially
returned HTTP 503 (model experiencing high demand), then passed a
targeted retry.
- The initial live run reported 20 passed, 80 skipped, and 3 failed.
Besides the Gemini 503, both MZAI cases failed before an API request
because the existing model fixture lacks an MZAI entry. Skips cover
missing credentials/configuration and the existing
local-provider/Perplexity exclusions. Providers without local
credentials still need CI verification.
- All pre-commit checks pass for the changed files. Full-repository
checks pass except the existing mypy error at
`tests/unit/providers/test_openai_exceptions.py:188`: an
`httpx2.Response` is passed to the installed OpenAI SDK 2.53.0, which
expects `httpx.Response`.

## PR Type

- 🐛 Bug Fix

## Relevant issues

Regression from mozilla-ai#1327 and [integration run
34586285143](https://github.com/mozilla-ai/any-llm/actions/runs/34586285143/job/103221196155).

## Checklist

- [x] I understand the code I am submitting.
- [x] I have added unit tests that prove my fix/feature works
- [x] I have run this code locally and verified it fixes the issue.
- [ ] New and existing tests pass locally
- [x] Documentation was updated where necessary
- [x] I have read and followed the [contribution
guidelines](https://github.com/mozilla-ai/any-llm/blob/main/CONTRIBUTING.md)
- [x] **AI Usage:**
    - [ ] No AI was used.
    - [ ] AI was used for drafting/refactoring.
    - [x] This is fully AI-generated.

## AI Usage Information

- AI Model used: GPT-5.6 Sol (initial implementation), GPT-6
(simplification and verification)
- AI Developer Tool used: Pi (initial implementation), Codex (follow-up)
- Any other info you'd like to share: AI inspected the failing CI logs,
simplified the test helper, and ran unit and live-provider tests under
maintainer direction.

When answering questions by the reviewer, please respond yourself, do
not copy/paste the reviewer comments into an AI system and paste back
its answer. We want to discuss with you, not your AI :)

- [x] I am an AI Agent filling out this form (check box if true)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Tests**
  - Expanded agent-loop coverage for sequential and multiple tool calls.
- Added checks for optional tool-name propagation, premature responses,
and iteration-limit failures.
- Introduced reusable test utilities and asynchronous mock-client setup
to improve consistency across agent-loop tests.
  - Renamed tests to better reflect their coverage.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
…zilla-ai#1387)

## Description

google-genai's async client attaches an `aiohttp.ClientResponse` to the
errors it raises, and aiohttp spells the status `status`, not
`status_code`. `_extract_status_code` only read `status_code`, so every
async Gemini 4xx/5xx classified by message alone and came back with
`status_code=None`. This reads both spellings.

Split out of mozilla-ai#1294, whose gemini half landed in mozilla-ai#1381.

Tests: one new unit test in `tests/unit/test_exception_handler.py` that
builds a response carrying `status` and no `status_code`. It fails on
main (`assert None == 400`) and passes here.
`tests/unit/test_exception_handler.py`, 86 passed. Pre-commit clean.

## PR Type

- 🐛 Bug Fix

## Relevant issues

Split out of mozilla-ai#1294.

## Checklist

- [x] I understand the code I am submitting.
- [x] I have added unit tests that prove my fix/feature works
- [x] I have run this code locally and verified it fixes the issue.
- [x] New and existing tests pass locally
- [x] Documentation was updated where necessary (not applicable)
- [x] I have read and followed the [contribution
guidelines](https://github.com/mozilla-ai/any-llm/blob/main/CONTRIBUTING.md)
- [x] **AI Usage:**
    - [ ] No AI was used.
    - [x] AI was used for drafting/refactoring.
    - [ ] This is fully AI-generated.

## AI Usage Information

- AI Model used: Claude (Fable 5.1)
- AI Developer Tool used: Claude Code
- Any other info you'd like to share:

- [ ] I am an AI Agent filling out this form (check box if true)

https://claude.ai/code/session_01546kUvB5GcyCkhQVSjpjbk

Co-authored-by: Hareesh <hareeshbahuleyan@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants