Skip to content

fix(mistral): clamp reasoning_effort, drop prompt_mode, replay thinking traces - #1241

Merged
njbrake merged 1 commit into
mainfrom
fix/mistral-reasoning-effort-and-thinking-replay
Aug 5, 2026
Merged

njbrake merged 1 commit into
mainfrom
fix/mistral-reasoning-effort-and-thinking-replay

Conversation

@njbrake

@njbrake njbrake commented Aug 5, 2026 •

Copy link
Copy Markdown
Contributor

Description

Fixes the two test_reasoning.py::*[mistral] failures on main (run 31011944484), and one related bug found while verifying.

reasoning_effort clamping. Mistral accepts only "high" or "none", so passing any_llm's finer-grained levels through returned 400 reasoning_effort low is not supported for this model. Every explicit level now collapses onto "high". Bifrost does the same thing for Mistral, and both bifrost and litellm clamp rather than propagate a predictable 400 on other restricted providers (Gemini Pro, Bedrock Nova).

prompt_mode removed. #1236 added prompt_mode="reasoning", which Mistral documents only for the deprecated magistral models. It now returns 400 Reasoning prompt mode is not enabled for this model on every model, magistral included. An explicit reasoning_effort is sufficient to get a thinking block; I verified that against mistral-medium-3-5, mistral-small-latest, magistral-medium-latest and magistral-small-latest. The reasoning test model moves to mistral-medium-3-5 since magistral is deprecated.

Thinking replay. Mistral requires the full assistant message, thinking trace included, on later turns and warns that dropping it degrades quality. We parsed traces out but never folded them back in, so multi-turn reasoning silently lost them. _patch_messages now rebuilds the ThinkChunk, carrying the replay signature through extra_content as the Anthropic provider already does.

PR Type

  • 🐛 Bug Fix

Relevant issues

None filed; caught by the post-merge integration run above.

Checklist

  • I understand the code I am submitting.
  • I have added unit tests that prove my fix/feature works
  • I have run this code locally and verified it fixes the issue.
  • New and existing tests pass locally
  • Documentation was updated where necessary
  • I have read and followed the contribution guidelines
  • AI Usage:
    • No AI was used.
    • AI was used for drafting/refactoring.
    • This is fully AI-generated.

Verified locally with a real MISTRAL_API_KEY: both previously failing tests pass, the full Mistral integration selection is green (21 passed, 6 skipped), tests/unit is green (1881 passed), and multi-turn replay was confirmed end to end to carry the trace into turn 2.

AI Usage Information

  • AI Model used: Claude Opus 5

  • AI Developer Tool used: Claude Code

  • Any other info you'd like to share: Behavior was established by probing the live Mistral API rather than from docs alone, which is what showed prompt_mode now failing on magistral too. Cross-checked the clamping approach against litellm and bifrost before landing it.

  • I am an AI Agent filling out this form (check box if true)

Summary by CodeRabbit

  • New Features

    • Improved Mistral reasoning support by preserving thinking signatures in standard and streaming responses.
    • Enabled reliable replay of reasoning across multi-turn conversations and tool interactions.
  • Bug Fixes

    • Normalised explicit reasoning requests to Mistral’s supported high-effort setting.
    • Prevented unsupported reasoning values and unwanted prompt-mode changes from being sent.

Mistral accepts a reasoning_effort of only "high" or "none", so sending
any_llm's finer-grained levels verbatim failed with "reasoning_effort low
is not supported for this model". Collapse every explicit level onto
"high" instead, matching how other gateways clamp restricted effort
scales.

Also stop sending prompt_mode="reasoning". Mistral documents that param
only for the deprecated native-reasoning magistral models and now
rejects it on every model, magistral included, with "Reasoning prompt
mode is not enabled for this model". An explicit reasoning_effort is
sufficient to get a thinking block on every current model, verified
against magistral-medium-latest, magistral-small-latest,
mistral-small-latest and mistral-medium-3-5.

Point the reasoning test model at mistral-medium-3-5, since the
magistral models are deprecated.

Separately, preserve reasoning traces across turns. Mistral requires the
full assistant message, thinking trace included, to be replayed on
subsequent turns, and warns that dropping it degrades output quality.
any_llm normalizes the trace onto message.reasoning, so fold it back
into a ThinkChunk when converting messages, carrying the replay
signature through extra_content the way the Anthropic provider already
does.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@njbrake
njbrake temporarily deployed to integration-tests August 5, 2026 14:27 — with GitHub Actions Inactive
@njbrake njbrake added the run-integration-tests Put this label on a PR to trigger the integration test suite: works with forks label Aug 5, 2026
@github-actions github-actions Bot removed the run-integration-tests Put this label on a PR to trigger the integration test suite: works with forks label Aug 5, 2026
@coderabbitai

coderabbitai Bot commented Aug 5, 2026 •

Copy link
Copy Markdown

Review Change Stack

Walkthrough

Changes

Mistral reasoning parameters now collapse explicit efforts to "high" and no longer set prompt_mode automatically. Response and streaming conversion preserve thinking signatures. Assistant reasoning is rebuilt for multi-turn Mistral requests. Tests cover conversion, replay, signatures, and tool messages.

Mistral reasoning support

Layer / File(s) Summary
Reasoning effort conversion
src/any_llm/providers/mistral/mistral.py, tests/conftest.py, tests/unit/providers/test_mistral_provider.py
Explicit reasoning efforts convert to "high". "auto" and "none" remain unset. Explicit prompt_mode values are forwarded.
Response and stream signature conversion
src/any_llm/providers/mistral/utils.py, tests/unit/providers/test_mistral_provider.py
Normal and streaming thinking chunks expose reasoning text and signatures through extra_content.
Assistant reasoning replay
src/any_llm/providers/mistral/utils.py, tests/unit/providers/test_mistral_provider.py
Assistant reasoning is rebuilt as Mistral thinking and text chunks before message patching. Tool-message handling remains supported.

Possibly related PRs

Suggested reviewers: tbille

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly summarises the three main Mistral changes: clamping reasoning effort, removing prompt mode, and replaying thinking traces.
Description check ✅ Passed The description covers the required sections, explains the changes, identifies the bug-fix type, records verification results, and completes the checklist.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/mistral-reasoning-effort-and-thinking-replay

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@njbrake
njbrake merged commit f929e2d into main Aug 5, 2026
17 of 18 checks passed
@njbrake
njbrake deleted the fix/mistral-reasoning-effort-and-thinking-replay branch August 5, 2026 14:28
@codecov

codecov Bot commented Aug 5, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 97.82609% with 1 line in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
src/any_llm/providers/mistral/utils.py 97.67% 0 Missing and 1 partial ⚠️
Files with missing lines Coverage Δ
src/any_llm/providers/mistral/mistral.py 99.41% <100.00%> (-0.01%) ⬇️
src/any_llm/providers/mistral/utils.py 77.04% <97.67%> (+10.49%) ⬆️

... and 1 file with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 4

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@src/any_llm/providers/mistral/mistral.py`:
- Around line 107-115: Update the comments around _convert_completion_params in
src/any_llm/providers/mistral/mistral.py#L107-L115 to state only that the
provider does not add prompt_mode automatically, without claiming
caller-supplied prompt_mode is never forwarded. Update the corresponding test
description/comments in tests/unit/providers/test_mistral_provider.py#L390-L395
to state that default conversion does not add prompt_mode; no behavioral test
change is required.

In `@src/any_llm/providers/mistral/utils.py`:
- Around line 501-506: Update the message-processing loop around
_build_mistral_assistant_content so it never rebinds the loop variable msg. Keep
msg unchanged and assign any modified assistant message to a separate
patched_msg before appending it to processed_msg; otherwise append the original
message. Run Ruff lint/format and strict mypy afterward.

In `@tests/unit/providers/test_mistral_provider.py`:
- Around line 104-109: Replace the ANN401-violating Any annotation on the
reasoning parameter of test_patch_messages_rebuilds_thinking_chunk_for_replay
with str | dict[str, str], matching the two parametrized input shapes. Run Ruff
formatting/linting and strict mypy to verify the test passes all checks.
- Around line 1263-1288: Add a test alongside
test_create_openai_chunk_captures_thinking_signature covering a ThinkChunk with
thinking content but no signature. Build the same streaming event without the
signature field, call _create_openai_chunk_from_mistral_chunk, and assert the
resulting delta.extra_content is None while preserving the reasoning content
assertion.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 3819a32b-571b-483c-8286-763dd5e87c32

📥 Commits

Reviewing files that changed from the base of the PR and between aec66ff and 08a93e1.

📒 Files selected for processing (4)
  • src/any_llm/providers/mistral/mistral.py
  • src/any_llm/providers/mistral/utils.py
  • tests/conftest.py
  • tests/unit/providers/test_mistral_provider.py

Comment on lines +107 to +115
# An explicit reasoning_effort is all Mistral needs to emit a `thinking` block; every
# model defaults to not reasoning, so "auto" (leave the choice to Mistral) and "none"
# both mean send nothing. Note that prompt_mode="reasoning" must not be sent alongside
# it: Mistral documents that param only for the deprecated native-reasoning magistral
# models and now rejects it on every model, magistral included, with "Reasoning prompt
# mode is not enabled for this model".
reasoning_effort = converted_params.pop("reasoning_effort", None)
mistral_effort = _REASONING_EFFORT_TO_MISTRAL.get(reasoning_effort) if reasoning_effort else None
if mistral_effort is not None:
converted_params["reasoning_effort"] = mistral_effort
converted_params["prompt_mode"] = "reasoning"
if reasoning_effort is not None and reasoning_effort not in ("auto", "none"):
converted_params["reasoning_effort"] = _MISTRAL_REASONING_EFFORT

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win

Describe automatic prompt_mode removal accurately.

The implementation forwards caller-supplied prompt_mode through converted_params.update(kwargs). The current comments state that prompt_mode must never accompany reasoning_effort. Limit the wording to the provider not adding prompt_mode automatically.

  • src/any_llm/providers/mistral/mistral.py#L107-L115: state that _convert_completion_params does not add prompt_mode.
  • tests/unit/providers/test_mistral_provider.py#L390-L395: state that the default conversion does not add prompt_mode.
📍 Affects 2 files
  • src/any_llm/providers/mistral/mistral.py#L107-L115 (this comment)
  • tests/unit/providers/test_mistral_provider.py#L390-L395
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/any_llm/providers/mistral/mistral.py` around lines 107 - 115, Update the
comments around _convert_completion_params in
src/any_llm/providers/mistral/mistral.py#L107-L115 to state only that the
provider does not add prompt_mode automatically, without claiming
caller-supplied prompt_mode is never forwarded. Update the corresponding test
description/comments in tests/unit/providers/test_mistral_provider.py#L390-L395
to state that default conversion does not add prompt_mode; no behavioral test
change is required.

Comment on lines 501 to 506
for i, msg in enumerate(messages):
if msg.get("role") == "assistant":
rebuilt_content = _build_mistral_assistant_content(msg)
if rebuilt_content is not None:
msg = {**msg, "content": rebuilt_content}
processed_msg.append(msg)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win

Do not rebind the loop variable.

Ruff reports PLW2901 because line 505 reassigns msg. Keep msg as the input message and append a separate patched_msg.

As per coding guidelines, ruff lint and format, plus strict mypy, must pass before review.

🧰 Tools
🪛 Ruff (0.16.1)

[warning] 505-505: for loop variable msg overwritten by assignment target

(PLW2901)

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/any_llm/providers/mistral/utils.py` around lines 501 - 506, Update the
message-processing loop around _build_mistral_assistant_content so it never
rebinds the loop variable msg. Keep msg unchanged and assign any modified
assistant message to a separate patched_msg before appending it to
processed_msg; otherwise append the original message. Run Ruff lint/format and
strict mypy afterward.

Sources: Coding guidelines, Linters/SAST tools

Comment on lines +104 to +109
@pytest.mark.parametrize(
"reasoning",
["thought hard", {"content": "thought hard"}],
ids=["serialized-string", "object-form"],
)
def test_patch_messages_rebuilds_thinking_chunk_for_replay(reasoning: Any) -> None:

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win

Replace Any in the test parameter.

Ruff reports ANN401 on line 109. Use str | dict[str, str] for reasoning, because these are the only parametrised shapes.

As per coding guidelines, ruff lint and format, plus strict mypy, must pass before review.

🧰 Tools
🪛 Ruff (0.16.1)

[warning] 109-109: Dynamically typed expressions (typing.Any) are disallowed in reasoning

(ANN401)

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/unit/providers/test_mistral_provider.py` around lines 104 - 109,
Replace the ANN401-violating Any annotation on the reasoning parameter of
test_patch_messages_rebuilds_thinking_chunk_for_replay with str | dict[str,
str], matching the two parametrized input shapes. Run Ruff formatting/linting
and strict mypy to verify the test passes all checks.

Sources: Coding guidelines, Linters/SAST tools

Comment on lines +1263 to +1288
def test_create_openai_chunk_captures_thinking_signature() -> None:
"""Streaming deltas surface the thinking signature the same way as non-streaming."""
pytest.importorskip("mistralai")
from mistralai.client.models import TextChunk, ThinkChunk

from any_llm.providers.mistral.utils import _create_openai_chunk_from_mistral_chunk

choice = Mock()
choice.index = 0
choice.delta.content = [ThinkChunk(thinking=[TextChunk(text="step one")], signature="sig-abc")]
choice.delta.role = "assistant"
choice.delta.tool_calls = None
choice.finish_reason = None

event = Mock()
event.data.id = "chatcmpl-abc"
event.data.created = 1_700_000_000
event.data.model = "mistral-medium-3-5"
event.data.choices = [choice]
event.data.usage = None

chunk = _create_openai_chunk_from_mistral_chunk(event)

assert chunk.choices[0].delta.reasoning is not None
assert chunk.choices[0].delta.reasoning.content == "step one"
assert chunk.choices[0].delta.extra_content == {"mistral": {"signature": "sig-abc"}}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win

Test the unsigned streaming path.

Lines 283-284 add a signature-present branch. Lines 311-316 must omit extra_content when Mistral does not send a signature. Add a streaming ThinkChunk without a signature and assert that delta.extra_content is None.

As per coding guidelines, test every new branch, including edge paths.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/unit/providers/test_mistral_provider.py` around lines 1263 - 1288, Add
a test alongside test_create_openai_chunk_captures_thinking_signature covering a
ThinkChunk with thinking content but no signature. Build the same streaming
event without the signature field, call _create_openai_chunk_from_mistral_chunk,
and assert the resulting delta.extra_content is None while preserving the
reasoning content assertion.

Source: Coding guidelines

This branch was previously deployed

1 inactive deployment
integration-tests — 08a93e15 Deployed Aug 5, 2026 by njbrake via run-docs-tests #2315
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

1.24.0 Included in release 1.24.0

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant