Skip to content

fix(anthropic-adapter): translate stop_sequences and disabled thinking for non-Claude targets - #34589

Merged
tin-berri merged 5 commits into
litellm_internal_stagingfrom
litellm_lit4798_glm_stop_thinking
Jul 29, 2026
Merged

tin-berri merged 5 commits into
litellm_internal_stagingfrom
litellm_lit4798_glm_stop_thinking

Conversation

@tin-berri

@tin-berri tin-berri commented Jul 25, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Claude Code's classifier sends /v1/messages with stop_sequences to non-Claude targets (e.g. Fireworks GLM)
  • Fireworks rejects stop_sequences outright (HTTP 400), since its OpenAI-compatible API only accepts stop
  • thinking: {type: "disabled"} not mapped to reasoning_effort: "none", so the model still burns budget reasoning

How it solves it:

  • Adapter now translates stop_sequences -> stop for all non-Anthropic-native targets (mirrors how thinking is already handled)
  • thinking: {type: "disabled"} now maps to reasoning_effort: "none" instead of None
  • Guarded the reasoning_auto_summary wrapping so a disabled-thinking "none" stays a plain string instead of becoming {"effort": "none", "summary": "detailed"} (there's no reasoning trace to summarize when thinking is disabled, and non-Claude providers expect a plain string)

Relevant issues

  • Adapter's translatable-param list omitted stop_sequences, so it fell through to verbatim copy-through instead of translation
  • translate_anthropic_thinking_to_reasoning_effort returned None for the "disabled" case instead of "none"
  • Making "disabled" truthy surfaced a latent second issue: it now reaches the reasoning_auto_summary wrapping logic, which needed an explicit guard

Linear ticket

Resolves LIT-4798

Pre-Submission checklist

  • I have added meaningful tests
  • My PR passes all CI/CD checks (e.g., lint, format, unit tests)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review

Screenshots / Proof of Fix

Live proxy against a local stub standing in for Fireworks (no Fireworks credentials in this environment; stub enforces the same strict "extra field" rejection Fireworks applies).

Before fix (commit 78348fd1c7):

$ curl .../v1/messages -d '{"model":"glm-classifier","max_tokens":100,"messages":[...],"stop_sequences":["</block>"]}'
{"error":{"message":"litellm.BadRequestError: Fireworks_aiException - {\"error\":{\"message\":\"Extra inputs are not permitted: stop_sequences=['</block>']\" ...}}"}}
HTTP_STATUS:400

$ curl .../v1/messages -d '{"model":"glm-classifier","max_tokens":100,"messages":[...],"thinking":{"type":"disabled"}}'
{"content":[{"type":"text","text":"stop=None reasoning_effort=None"}]}
HTTP_STATUS:200

After fix (commit adeb46b157, stop_sequences + disabled-thinking core fix):

$ curl .../v1/messages -d '{"model":"glm-classifier","max_tokens":100,"messages":[...],"stop_sequences":["</block>"]}'
{"content":[{"type":"text","text":"stop=['</block>'] reasoning_effort=None"}]}
HTTP_STATUS:200

$ curl .../v1/messages -d '{"model":"glm-classifier","max_tokens":100,"messages":[...],"thinking":{"type":"disabled"}}'
{"content":[{"type":"text","text":"stop=None reasoning_effort='none'"}]}
HTTP_STATUS:200

Follow-up fix (commit a6ce7b10f9, mutation-tested unit repro/verify since it's an interaction with a global opt-in flag rather than a Fireworks-specific wire issue):

# reasoning_auto_summary=True, thinking={"type": "disabled"}
# before: {"effort": "none", "summary": "detailed"}   <- would be rejected as reasoning_effort by non-Claude providers
# after:  "none"

Type

🐛 Bug Fix


Note

Low Risk
Scoped to the experimental Anthropic→OpenAI adapter and covered by new unit tests; behavior change is intentional for non-Claude routing with no auth or data-path impact.

Overview
Fixes Anthropic /v1/messages pass-through when the upstream target is OpenAI-compatible (e.g. Fireworks) instead of native Claude.

Stop sequences: stop_sequences is now a translatable param and is mapped to OpenAI stop during translate_anthropic_to_openai, so providers that reject unknown stop_sequences no longer get a verbatim copy-through. Empty lists are ignored.

Disabled thinking: thinking: {type: "disabled"} now becomes reasoning_effort: "none" for non-Claude models (previously dropped as None), so reasoning can actually be turned off.

Reasoning summary wrapping: Duplicated summary/auto-summary logic is centralized in _apply_reasoning_summary_wrapping. When thinking is disabled, "none" stays a plain string even if reasoning_auto_summary is on—avoiding {"effort": "none", "summary": "detailed"} that non-Claude backends reject.

Unit tests cover stop translation, disabled thinking, and the auto-summary interaction.

Reviewed by Cursor Bugbot for commit d478b99. Bugbot is set up for automated code reviews on this repo. Configure here.

@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai

@greptile-apps

greptile-apps Bot commented Jul 25, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

Updates the experimental Anthropic messages adapter to:

  • Translate stop_sequences into the OpenAI-compatible stop parameter.
  • Map disabled thinking to reasoning_effort: "none".
  • Keep disabled reasoning effort as a plain string when automatic summaries are enabled.
  • Add focused adapter and handler regression tests for these behaviors.

Confidence Score: 5/5

The PR appears safe to merge.

No blocking failures remain.

Important Files Changed

Filename Overview
litellm/llms/anthropic/experimental_pass_through/adapters/transformation.py Adds stop-sequence translation and consistently preserves disabled-thinking semantics across both reasoning translation paths.
tests/test_litellm/llms/anthropic/experimental_pass_through/adapters/test_anthropic_experimental_pass_through_adapters_transformation.py Adds focused unit coverage for stop translation, empty stop sequences, disabled thinking, and automatic-summary interaction.
tests/test_litellm/llms/anthropic/experimental_pass_through/messages/test_anthropic_experimental_pass_through_messages_handler.py Extends handler-level coverage for disabled-thinking translation and downstream request behavior.

Reviews (4): Last reviewed commit: "fix(anthropic-adapter): drop redundant c..." | Re-trigger Greptile

@codecov

codecov Bot commented Jul 25, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai

…g for non-Claude targets

Claude Code's auto-mode classifier sends stop_sequences and thinking:
{type: disabled} on /v1/messages. The Anthropic adapter passed
stop_sequences through unchanged instead of mapping it to OpenAI's stop,
which Fireworks' OpenAI-compatible endpoint rejects with HTTP 400. It also
dropped disabled thinking instead of mapping it to reasoning_effort: none,
so the model spent its output budget on reasoning it was told to skip.

Resolves LIT-4798
…in string

Guard against reasoning_auto_summary wrapping "none" into a dict when
thinking is disabled — there's no reasoning trace to summarize, and
non-Claude providers (e.g. Fireworks) expect reasoning_effort as a
plain string.
@tin-berri
tin-berri force-pushed the litellm_lit4798_glm_stop_thinking branch from a6ce7b1 to 9da21f3 Compare July 25, 2026 01:48
@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai

Codecov flagged the empty-list early-return in
_translate_stop_sequences_to_openai as an uncovered line in the diff —
add a regression test asserting stop_sequences=[] does not set
new_kwargs["stop"].
@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

…ling gap

translate_thinking_for_model duplicated the same summary/auto_summary
wrapping logic as _translate_thinking_to_openai without the
disabled-thinking guard, so it could still wrap "none" into an
{effort, summary} dict when reasoning_auto_summary is enabled (caught
by Cursor Bugbot). Extract the wrapping rule into one shared
_apply_reasoning_summary_wrapping helper used by both call sites so
this invariant can't drift apart again.
@tin-berri

Copy link
Copy Markdown
Contributor Author

Good catch — translate_thinking_for_model duplicated the same wrapping logic without the disabled-thinking guard. Rather than patch that one call site, extracted the shared rule into _apply_reasoning_summary_wrapping so both paths use one implementation (5072590).

…ne budget

_apply_reasoning_summary_wrapping already returns Any, so wrapping its
dict-literal returns in cast(Any, ...) was a no-op that only inflated the
LIT006 cast-count budget the lint gate enforces.
@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit d478b99. Configure here.

@codspeed

codspeed Bot commented Jul 25, 2026

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_lit4798_glm_stop_thinking (d478b99) with litellm_internal_staging (7b019cf)

Open in CodSpeed

@tin-berri
tin-berri enabled auto-merge July 27, 2026 18:17

@mateo-berri mateo-berri left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM; thanks!

@tin-berri
tin-berri disabled auto-merge July 29, 2026 00:42
@tin-berri
tin-berri merged commit 32a4377 into litellm_internal_staging Jul 29, 2026
79 of 80 checks passed
@tin-berri
tin-berri deleted the litellm_lit4798_glm_stop_thinking branch July 29, 2026 00:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants