Skip to content

fix(bedrock): map the full Converse stopReason set to finish_reason - #1358

Merged
javiermtorres merged 1 commit into
mozilla-ai:mainfrom
emecii:fix/bedrock-stop-reason-mapping
Sep 11, 2026
Merged

javiermtorres merged 1 commit into
mozilla-ai:mainfrom
emecii:fix/bedrock-stop-reason-mapping

Conversation

@emecii

@emecii emecii commented Aug 31, 2026 •

Copy link
Copy Markdown
Contributor

Description

The Bedrock provider maps only three of the Converse API's nine stopReason values, so a guardrail-blocked or context-overflowed response reaches the caller as an ordinary finish_reason="stop".

src/any_llm/providers/bedrock/utils.py:517 (non-streaming):

finish_reason: Literal["stop", "length"] = "length" if stop_reason == "max_tokens" else "stop"

and :611-617 (streaming) handles max_tokens and tool_use and sends everything else to "stop".

The visible failure is the structured-output guard in src/any_llm/any_llm.py:798-806. When a Bedrock Guardrail blocks a response, Converse returns stopReason="guardrail_intervened" with the guardrail's blocked-message text as content. Because that arrives as finish_reason="stop", the content_filter branch never fires and the guardrail's prose is handed to parse_json_content() instead, so the caller gets a pydantic ValidationError from an unrelated layer rather than the ContentFilterFinishReasonError this library raises for every other provider. content_filtered behaves the same way, and model_context_window_exceeded reports as "stop" rather than "length".

This is the same fix already merged for gemini (#1202), zai (#1204), cohere (#1301) and anthropic (#1306/#1328). Bedrock was the last provider still hardcoding a partial mapping, so this follows ANTHROPIC_STOP_REASON_TO_FINISH_REASON in shape and naming:

BEDROCK_STOP_REASON_TO_FINISH_REASON: dict[str, _FinishReason] = {
    "end_turn": "stop",
    "max_tokens": "length",
    "model_context_window_exceeded": "length",
    "tool_use": "tool_calls",
    "content_filtered": "content_filter",
    "guardrail_intervened": "content_filter",
}

stop_sequence, malformed_model_output and malformed_tool_use have no OpenAI counterpart and keep falling through to the "stop" default, as do any values a future service model adds. Both call sites now go through one _map_stop_reason() helper, which also lets the cast at the non-streaming call site go away.

The nine-value enum is not taken from the AWS docs prose. It is read out of the installed botocore service model (bedrock-runtime, shape StopReason), which is the same source the SDK validates against:

>>> import botocore.session
>>> botocore.session.Session().get_service_model("bedrock-runtime").shape_for("StopReason").enum
['end_turn', 'tool_use', 'max_tokens', 'stop_sequence', 'guardrail_intervened', 'content_filtered',
 'malformed_model_output', 'malformed_tool_use', 'model_context_window_exceeded']

The tests read the enum from that same service model and parametrize over it, so a botocore upgrade that adds a stop reason fails the suite instead of silently defaulting the new reason to "stop".

Reproduction

import asyncio
from unittest.mock import Mock

from pydantic import BaseModel

from any_llm.exceptions import ContentFilterFinishReasonError
from any_llm.providers.bedrock import BedrockProvider


class City(BaseModel):
    name: str


# What Converse returns when a Bedrock Guardrail blocks the response.
blocked = {
    "output": {"message": {"content": [{"text": "Sorry, I cannot answer that."}]}},
    "stopReason": "guardrail_intervened",
}

client = Mock()
client.converse.return_value = blocked
provider = BedrockProvider(client=client)

print("finish_reason:", provider._convert_completion_response(blocked).choices[0].finish_reason)

try:
    asyncio.run(
        provider.acompletion(
            model="us.anthropic.claude-sonnet-4-20250514-v1:0",
            messages=[{"role": "user", "content": "Hello"}],
            response_format=City,
        )
    )
except Exception as exc:
    print("raised:", type(exc).__name__)

Before:

finish_reason: stop
raised: ValidationError

After:

finish_reason: content_filter
raised: ContentFilterFinishReasonError

Tests

Added to tests/unit/providers/test_aws_provider.py:

  • test_convert_response_maps_every_bedrock_stop_reason and test_streaming_chunk_maps_every_bedrock_stop_reason, parametrized over the full botocore StopReason enum, asserting both paths agree.
  • test_convert_response_without_stop_reason_finishes_as_stop.
  • test_guardrail_blocked_structured_output_raises_content_filter_error, the end-to-end case above.

On the unfixed tree these fail:

FAILED test_convert_response_maps_every_bedrock_stop_reason[tool_use]
FAILED test_convert_response_maps_every_bedrock_stop_reason[guardrail_intervened]
FAILED test_convert_response_maps_every_bedrock_stop_reason[content_filtered]
FAILED test_convert_response_maps_every_bedrock_stop_reason[model_context_window_exceeded]
FAILED test_streaming_chunk_maps_every_bedrock_stop_reason[guardrail_intervened]
FAILED test_streaming_chunk_maps_every_bedrock_stop_reason[content_filtered]
FAILED test_streaming_chunk_maps_every_bedrock_stop_reason[model_context_window_exceeded]
FAILED test_guardrail_blocked_structured_output_raises_content_filter_error
8 failed, 12 passed

With the fix, uv run pytest tests/unit/providers/test_aws_provider.py → 116 passed.

The [tool_use] case is the one behaviour change beyond the three broken reasons: a stopReason="tool_use" response that carries no toolUse block now reports "tool_calls" instead of "stop". Both real tool_use shapes (a genuine tool call, and the synthetic any_llm_structured_output unwrap) return earlier in _convert_response and are untouched.

PR Type

  • 🐛 Bug Fix

Relevant issues

No open issue. Same fix as #1202 / #1204 / #1301 / #1328, applied to the remaining provider.

Checklist

  • I understand the code I am submitting.
  • I have added unit tests that prove my fix/feature works
  • I have run this code locally and verified it fixes the issue.
  • New and existing tests pass locally
  • Documentation was updated where necessary (no user-facing API changed)
  • I have read and followed the contribution guidelines
  • AI Usage:
    • No AI was used.
    • AI was used for drafting/refactoring.
    • This is fully AI-generated.

Verification

  • uv run pre-commit run --all-files (ruff, ruff-format, mypy, codespell) clean.
  • uv run pytest tests/unit/providers/test_aws_provider.py → 116 passed.
  • uv run pytest tests/unit → 2308 passed, 69 skipped, 16 failed. All 16 failures are in test_anthropic_messages.py / test_anthropic_provider.py and reproduce identically on an unmodified main in this environment: TypeError: Invalid 'http_client' argument; Expected an instance of httpx2.AsyncClient but got <class 'httpx.AsyncClient'>, from what a local uv sync --all-extras -U resolves. Nothing bedrock-related.
  • No integration run: I do not have AWS Bedrock credentials, and the two Bedrock paths this touches are pure response converters exercised by the unit tests above against service-model-sourced stop reasons.

AI Usage Information

  • AI Model used: Claude Opus 5

  • AI Developer Tool used: Claude Code

  • Any other info you'd like to share: The stop-reason enum was read from the installed botocore service model rather than the AWS docs, and the tests read it from the same place so the list cannot drift.

  • I am an AI Agent filling out this form (check box if true)

Summary by CodeRabbit

  • Bug Fixes
    • Improved Amazon Bedrock response handling for length limits, tool calls, and content-filter outcomes.
    • Standardised finish reasons across streaming and non-streaming responses.
    • Unknown or unsupported stop reasons now safely default to stop.
    • Structured-output requests blocked by a Bedrock guardrail now report a content-filter error correctly.

@coderabbitai

coderabbitai Bot commented Aug 31, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Team

Run ID: 587d0c7e-9854-4a4f-b781-7429d565ee63

📥 Commits

Reviewing files that changed from the base of the PR and between 9e976ca and 229b5db.

📒 Files selected for processing (2)
  • src/any_llm/providers/bedrock/utils.py
  • tests/unit/providers/test_aws_provider.py

Included review availability: Your plan provides up to 8 included reviews per hour; 7 remain after this review.


Walkthrough

Bedrock stop reasons now use a shared mapping for streaming and non-streaming responses. Tests cover service-model reasons, missing stop reasons, and guardrail-blocked structured output.

Changes

Bedrock stop-reason mapping

Layer / File(s) Summary
Stop-reason contract
src/any_llm/providers/bedrock/utils.py
Adds the shared Bedrock-to-OpenAI finish-reason mapping and defaults unknown or non-string reasons to stop.
Response conversion
src/any_llm/providers/bedrock/utils.py
Applies the shared mapping to non-streaming and streaming response conversion.
Mapping and guardrail validation
tests/unit/providers/test_aws_provider.py
Tests service-model stop reasons, the missing-reason fallback, both conversion paths, and guardrail-blocked structured output.

Suggested reviewers: tbille

Merge Risk: ⚪ Minimal · up to 96314

Bedrock responses now report tool calls, token limits, content filtering, and guardrail intervention through the appropriate finish reasons in both streaming and non-streaming flows. The implementation has focused coverage for these behaviors, with no remaining current-head merge risk identified.

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly and concisely describes the main change: mapping the full Bedrock Converse stopReason set to finish_reason.
Description check ✅ Passed The description follows the required template and provides clear scope, rationale, mappings, tests, verification results, known unrelated failures, and AI usage details.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 10 functions across 2 files.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@emecii
emecii temporarily deployed to integration-tests September 2, 2026 10:14 — with GitHub Actions Inactive
@javiermtorres

Copy link
Copy Markdown
Contributor

@emecii please check the conflicts. Otherwise LGTM.

@codecov

codecov Bot commented Sep 2, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

Files with missing lines Coverage Δ
src/any_llm/providers/bedrock/utils.py 89.39% <100.00%> (+0.28%) ⬆️

... and 1 file with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@emecii
emecii force-pushed the fix/bedrock-stop-reason-mapping branch from 9e976ca to 229b5db Compare September 5, 2026 02:10
@javiermtorres
javiermtorres force-pushed the fix/bedrock-stop-reason-mapping branch from 229b5db to 05612c7 Compare September 11, 2026 07:48
@javiermtorres
javiermtorres deployed to integration-tests September 11, 2026 07:49 — with GitHub Actions Active
The Converse API reports nine stop reasons but the provider mapped three,
sending everything else to "stop". A guardrail block
("guardrail_intervened"), a content filter hit ("content_filtered") and a
context overflow ("model_context_window_exceeded") therefore looked like a
normal completion.

The visible failure is the structured-output guard in any_llm.py: with
finish_reason="stop" the content_filter branch never fires, so the
guardrail's blocked-message prose reaches parse_json_content() and the
caller gets a pydantic ValidationError instead of
ContentFilterFinishReasonError.

Add BEDROCK_STOP_REASON_TO_FINISH_REASON, mirroring
ANTHROPIC_STOP_REASON_TO_FINISH_REASON, and route both the streaming and
non-streaming paths through it. "stop_sequence", "malformed_model_output"
and "malformed_tool_use" have no OpenAI counterpart and keep falling
through to "stop".

The tests read the stopReason enum from botocore's bedrock-runtime service
model and parametrize over it, so a botocore upgrade that adds a reason
fails the suite rather than silently defaulting it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@javiermtorres
javiermtorres force-pushed the fix/bedrock-stop-reason-mapping branch from 05612c7 to 9631447 Compare September 11, 2026 08:01
@javiermtorres
javiermtorres deployed to integration-tests September 11, 2026 08:02 — with GitHub Actions Active
@javiermtorres
javiermtorres merged commit 25ac9a2 into mozilla-ai:main Sep 11, 2026
14 checks passed
@github-actions github-actions Bot added the 1.28.0 Included in release 1.28.0 label Sep 18, 2026

This branch was successfully deployed

1 active deployment
integration-tests — 96314476 Deployed Sep 11, 2026 by javiermtorres via run-docs-tests #2919
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

1.28.0 Included in release 1.28.0

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants