Skip to content

chore(deps): litellm 1.98.0 -> 1.98.1, so Azure GPT-6 models take reasoning_effort - #5

Merged
johndmendonca merged 1 commit into
johndmendonca:mainfrom
RubenBranco:chore/litellm-1.98.1
Sep 25, 2026
Merged

johndmendonca merged 1 commit into
johndmendonca:mainfrom
RubenBranco:chore/litellm-1.98.1

Conversation

@RubenBranco

Copy link
Copy Markdown

Why

litellm 1.98.0 recognises only GPT-5 names as Azure reasoning models. For azure/gpt-6-luna, gpt-6-sol and gpt-6-astra it rejects reasoning_effort with UnsupportedParamsError. tau2 sets litellm.drop_params = True, so here litellm drops the effort without an error, and GPT-6 runs at its default effort. litellm fixed the name match in BerriAI/litellm#39631 and backported the fix to 1.98.1 (BerriAI/litellm#43130). On 1.98.1, tau2's agent calls, which always carry tools, go to the Azure Responses API with reasoning.effort.

1.98.1 is a patch on the 1.98 line. It stays below 1.100, where litellm's hosted_vllm path drops reasoning_content from tool-loop history.

Changes

  • uv.lock moves litellm from 1.98.0 to 1.98.1. No other package moves.
  • pyproject.toml is unchanged, because litellm>=1.98.0 already allows 1.98.1.

Evidence

No model was called. A local stub HTTP server recorded every request litellm sent and returned canned replies.

Tests. 190 passed, 17 failed and 1 xfailed, on main and on this branch. The same 17 tests fail on both with AuthenticationError: they call a real LLM API, and the run had no provider credentials. The command was uv run --locked --with pytest --with pytest-asyncio pytest tests. As in #4, it excluded tests/test_domains/test_banking_knowledge, tests/test_gym, tests/test_streaming and tests/test_voice, which need optional dependencies.

vLLM requests are unchanged. LLMAgent on hosted_vllm/glm-5.2-fp8 ran a 3-request airline tool loop. The requests were a user turn, the result of a get_user_details call from the airline environment, and a second user turn. Every stub reply carried reasoning_content. The 3 request bodies are byte-identical on 1.98.0 and 1.98.1. The second and third requests replay 1 and 2 assistant messages, each with reasoning and reasoning_content.

GPT-6 now gets the effort. Each call used azure/gpt-6-luna with reasoning_effort="high" and max_completion_tokens=16000.

Call 1.98.0 1.98.1
LLMAgent, 14 tools chat completions, no effort sent /openai/responses with reasoning.effort: high, max_output_tokens: 16000, model gpt-6-luna
generate(), no tools chat completions, no effort sent chat completions with reasoning_effort: high, max_completion_tokens: 16000
generate(), no tools, litellm.drop_params = False UnsupportedParamsError chat completions with reasoning_effort: high, max_completion_tokens: 16000

Every result above is the same with litellm's live model cost map and with LITELLM_LOCAL_MODEL_COST_MAP=True.

🤖 Generated with Claude Code

…soning_effort.

- uv.lock: litellm 1.98.0 -> 1.98.1; no other package moves
- pyproject.toml unchanged: the litellm>=1.98.0 floor already allows 1.98.1

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@RubenBranco
RubenBranco marked this pull request as ready for review September 25, 2026 12:32
@johndmendonca
johndmendonca merged commit a8cf25a into johndmendonca:main Sep 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants