Skip to content

chore(deps): litellm 1.98.0 -> 1.98.1, so Azure GPT-6 models take reasoning_effort - #2

Merged
zeppombal merged 1 commit into
zeppombal:mainfrom
RubenBranco:chore/litellm-1.98.1
Sep 25, 2026
Merged

zeppombal merged 1 commit into
zeppombal:mainfrom
RubenBranco:chore/litellm-1.98.1

Conversation

@RubenBranco

Copy link
Copy Markdown

Why

litellm 1.98.0 recognises only GPT-5 names as Azure reasoning models. For azure/gpt-6-luna, gpt-6-sol and gpt-6-astra it rejects reasoning_effort with UnsupportedParamsError, so every collect call for a GPT-6 model with a reasoning effort fails. litellm fixed the name match in BerriAI/litellm#39631 and backported the fix to 1.98.1 (BerriAI/litellm#43130).

1.98.1 is a patch on the 1.98 line. It stays below 1.100, where litellm's hosted_vllm path drops reasoning_content from tool-loop history.

Changes

  • uv.lock moves litellm from 1.98.0 to 1.98.1. No other package moves.
  • pyproject.toml is unchanged, because litellm>=1.50 already allows 1.98.1.

Evidence

No model was called. A local stub HTTP server recorded every request litellm sent and returned canned replies.

Tests. The repo has no test suite. uv sync --locked succeeds on main and on this branch.

vLLM requests are unchanged. collect_one ran on the first 2 questions of questions.v2.json for the bare model name glm-5.2-fp8, which LiteLLMClient sends as hosted_vllm/glm-5.2-fp8. It used the settings Dawn renders into config.v2.json: no system prompt, no temperature, reasoning effort off, max_tokens: 24576 and retries: 1. The 2 request bodies are byte-identical on 1.98.0 and 1.98.1.

GPT-6 now gets the effort. Each call used azure/gpt-6-luna with effort high.

Call 1.98.0 1.98.1
collect_one, max_tokens=16000 error row: litellm call failed (attempt 1/1): litellm.UnsupportedParamsError: azure does not support parameters: ['reasoning_effort'] chat completions with reasoning_effort: high, max_completion_tokens: 16000
litellm.completion, max_completion_tokens=16000 UnsupportedParamsError chat completions with reasoning_effort: high, max_completion_tokens: 16000

Every result above is the same with litellm's live model cost map and with LITELLM_LOCAL_MODEL_COST_MAP=True.

🤖 Generated with Claude Code

…soning_effort.

- uv.lock: litellm 1.98.0 -> 1.98.1; no other package moves
- pyproject.toml unchanged: the litellm>=1.50 floor already allows 1.98.1

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@RubenBranco
RubenBranco marked this pull request as ready for review September 25, 2026 12:32
@zeppombal
zeppombal merged commit 9bfef25 into zeppombal:main Sep 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants