Skip to content

ROB-1797 fix bug of losing temp arg - #698

Merged
arikalon1 merged 5 commits into
masterfrom
ROB-1979-bedrock-think-and-temp
Jul 30, 2025
Merged

arikalon1 merged 5 commits into
masterfrom
ROB-1979-bedrock-think-and-temp

Conversation

@RoiGlinik

Copy link
Copy Markdown
Collaborator

fixes a bug where temperature is poped and then lost for 2+ complete iterations.

seen on bedrock with thinking arguments .

@coderabbitai

coderabbitai Bot commented Jul 23, 2025 •

Copy link
Copy Markdown
Contributor

Walkthrough

The changes adjust how the temperature parameter is handled when invoking language model completions. In holmes/core/llm.py, the fallback to a global TEMPERATURE constant is removed and temperature is sourced directly from method arguments or self.args. In holmes/core/tool_calling_llm.py, the TEMPERATURE environment variable is now explicitly passed to LLM completion calls.

Changes

Cohort / File(s) Change Summary
Core LLM logic
holmes/core/llm.py
Removed import of TEMPERATURE. Changed logic to set and pass temperature in LLM completion calls, eliminating fallback to a global constant.
Tool-calling LLM integration
holmes/core/tool_calling_llm.py
Added explicit use of the TEMPERATURE environment variable as an argument in all LLM completion calls within ToolCallingLLM.

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~15 minutes

Note

⚡️ Unit Test Generation is now available in beta!

Learn more here, or try it out under "Finishing Touches" below.


📜 Recent review details

Configuration used: CodeRabbit UI
Review profile: CHILL
Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between a6fed22 and 29a15e5.

📒 Files selected for processing (1)
  • holmes/core/tool_calling_llm.py (4 hunks)
🚧 Files skipped from review as they are similar to previous changes (1)
  • holmes/core/tool_calling_llm.py
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (3)
  • GitHub Check: llm_evals
  • GitHub Check: Pre-commit checks
  • GitHub Check: Pre-commit checks
✨ Finishing Touches
  • 📝 Generate Docstrings
🧪 Generate unit tests
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch ROB-1979-bedrock-think-and-temp

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share
🪧 Tips

Chat

There are 3 ways to chat with CodeRabbit:

  • Review comments: Directly reply to a review comment made by CodeRabbit. Example:
    • I pushed a fix in commit <commit_id>, please review it.
    • Explain this complex logic.
    • Open a follow-up GitHub issue for this discussion.
  • Files and specific lines of code (under the "Files changed" tab): Tag @coderabbitai in a new review comment at the desired location with your query. Examples:
    • @coderabbitai explain this code block.
    • @coderabbitai modularize this function.
  • PR comments: Tag @coderabbitai in a new PR comment to ask questions about the PR branch. For the best results, please provide a very specific query, as very limited context is provided in this mode. Examples:
    • @coderabbitai gather interesting stats about this repository and render them as a table. Additionally, render a pie chart showing the language distribution in the codebase.
    • @coderabbitai read src/utils.ts and explain its main purpose.
    • @coderabbitai read the files in the src/scheduler package and generate a class diagram using mermaid and a README in the markdown format.
    • @coderabbitai help me debug CodeRabbit configuration file.

Support

Need help? Create a ticket on our support page for assistance with any issues or questions.

Note: Be mindful of the bot's finite context window. It's strongly recommended to break down tasks such as reading entire modules into smaller chunks. For a focused discussion, use review comments to chat about specific files and their changes, instead of using the PR comments.

CodeRabbit Commands (Invoked using PR comments)

  • @coderabbitai pause to pause the reviews on a PR.
  • @coderabbitai resume to resume the paused reviews.
  • @coderabbitai review to trigger an incremental review. This is useful when automatic reviews are disabled for the repository.
  • @coderabbitai full review to do a full review from scratch and review all the files again.
  • @coderabbitai summary to regenerate the summary of the PR.
  • @coderabbitai generate docstrings to generate docstrings for this PR.
  • @coderabbitai generate sequence diagram to generate a sequence diagram of the changes in this PR.
  • @coderabbitai generate unit tests to generate unit tests for this PR.
  • @coderabbitai resolve resolve all the CodeRabbit review comments.
  • @coderabbitai configuration to show the current CodeRabbit configuration for the repository.
  • @coderabbitai help to get help.

Other keywords and placeholders

  • Add @coderabbitai ignore anywhere in the PR description to prevent this PR from being reviewed.
  • Add @coderabbitai summary to generate the high-level summary at a specific location in the PR description.
  • Add @coderabbitai anywhere in the PR title to generate the title automatically.

CodeRabbit Configuration File (.coderabbit.yaml)

  • You can programmatically configure CodeRabbit by adding a .coderabbit.yaml file to the root of your repository.
  • Please see the configuration documentation for more information.
  • If your editor has YAML language server enabled, you can add the path at the top of this file to enable auto-completion and validation: # yaml-language-server: $schema=https://coderabbit.ai/integrations/schema.v2.json

Documentation and Community

  • Visit our Documentation for detailed information on how to use CodeRabbit.
  • Join our Discord Community to get help, request features, and share feedback.
  • Follow us on X/Twitter for updates and announcements.

@RoiGlinik
RoiGlinik requested a review from arikalon1 July 30, 2025 06:47

@arikalon1 arikalon1 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nice work

@github-actions

Copy link
Copy Markdown
Contributor

Results of HolmesGPT evals

  • ask_holmes: 24/37 test cases were successful, 0 regressions, 1 skipped, 12 mock failures
Test suite Test case Status
ask 01_how_many_pods ✅
ask 02_what_is_wrong_with_pod 🔧
ask 03_what_is_the_command_to_port_forward 🔧
ask 04_related_k8s_events ↪️
ask 05_image_version 🔧
ask 09_crashpod ✅
ask 10_image_pull_backoff 🔧
ask 11_init_containers ✅
ask 14_pending_resources ✅
ask 15_failed_readiness_probe ✅
ask 17_oom_kill ✅
ask 18_crash_looping_v2 ✅
ask 19_detect_missing_app_details 🔧
ask 24_misconfigured_pvc 🔧
ask 28_permissions_error ✅
ask 29_events_from_alert_manager ✅
ask 39_failed_toolset 🔧
ask 41_setup_argo0 ✅
ask 41_setup_argo1 ✅
ask 42_dns_issues_steps_new_tools 🔧
ask 43_current_datetime_from_prompt ✅
ask 45_fetch_deployment_logs_simple ✅
ask 51_logs_summarize_errors 🔧
ask 53_logs_find_term ✅
ask 54_not_truncated_when_getting_pods 🔧
ask 59_label_based_counting ✅
ask 60_count_less_than 🔧
ask 61_exact_match_counting ✅
ask 63_fetch_error_logs_no_errors ✅
ask 77_liveness_probe_misconfiguration 🔧
ask 79_configmap_mount_issue 🔧
ask 83_secret_not_found 🔧
ask 86_configmap_like_but_secret 🔧
ask 89_runbook_missing_cloudwatch 🔧
ask 90_runbook_basic_selection 🔧
ask 24a_misconfigured_pvc_basic 🔧
ask 13a_pending_node_selector_basic 🔧

Legend

  • ✅ the test was successful
  • ↪️ the test was skipped
  • ⚠️ the test failed but is known to be flaky or known to fail
  • 🔧 the test failed due to mock data issues (not a code regression)
  • ❌ the test failed and should be fixed before merging the PR

@arikalon1
arikalon1 merged commit 0521cd0 into master Jul 30, 2025
@arikalon1
arikalon1 deleted the ROB-1979-bedrock-think-and-temp branch July 30, 2025 08:25
aantn pushed a commit that referenced this pull request May 6, 2026
## Human part

Hey! I noticed when using Opus 4.7 with Holmes that it would complain
that `temperature` was not supported (this is via Bedrock btw). Trying
to remove `temperature` or add `drop_params` did nothing.

## Summary

`DefaultLLM.completion()` at
[`holmes/core/llm.py:564`](https://github.com/HolmesGPT/holmesgpt/blob/master/holmes/core/llm.py#L564)
called

```python
self.args.setdefault("temperature", temperature)
```

unconditionally. With the signature default `temperature:
Optional[float] = None`, this materialized a `temperature=None` key in
`self.args` (and therefore in the kwargs forwarded to
`litellm.completion` via `**self.args`) even when the caller passed
nothing.

LiteLLM's `drop_params=True` (passed by both call sites —
`holmes/core/tool_calling_llm.py:1077` and
`holmes/core/truncation/compaction.py:133`) only strips params a
provider is KNOWN not to support. Newer Bedrock Anthropic endpoints
(e.g. Claude Opus 4.7) nominally support temperature but reject
`temperature=None`, so the request fails. Setting `temperature: null` in
`modelList` didn't help either: the null flowed into `self.args`, the
subsequent `setdefault` was a no-op, and the bogus key still went
through.

This PR fixes the kwargs-construction bug at that line. It does **not**
expand Holmes' responsibility for provider capability decisions —
LiteLLM's `drop_params` stays the authority for real provider-specific
stripping.

## The fix (3 lines at `holmes/core/llm.py:564`)

```python
# Strip a pre-existing `temperature: None` (e.g. from `temperature: null` in
# modelList) before applying the caller's value, so setdefault() is not blocked
# by a null sentinel and so no `temperature=None` leaks to providers that reject
# it (e.g. Bedrock Anthropic Opus 4.7). Preserves PR #698: when args holds a real
# temperature, setdefault is a no-op and the persisted value survives.
if self.args.get("temperature", ...) is None:
    self.args.pop("temperature", None)
if temperature is not None:
    self.args.setdefault("temperature", temperature)
```

Order matters: the pop runs first to strip any `None` already in
`self.args`. Only then does `setdefault` apply the caller's value. If we
ran `setdefault` first, a config `None` would block the caller's real
temperature (because `setdefault` is a no-op when the key exists) and
the subsequent pop would silently drop it.

## Behavior matrix

| # | Caller `temperature` | `self.args` before | Forwarded to
`litellm.completion` | Case |
|---|---|---|---|---|
| 1 | `0.7` | `{}` | `temperature=0.7` | Normal call (unchanged) |
| 2 | `0.0` | `{}` | `temperature=0.0` | Falsy but valid (unchanged) |
| 3 | `None` | `{"temperature": 0.5}` | `temperature=0.5` | **PR #698
guard** — persisted temperature survives |
| 4 | `None` | `{}` | no `temperature` key | **Fixed** — no bogus `None`
materialized |
| 5 | `None` | `{"temperature": None}` | no `temperature` key |
**Fixed** — `temperature: null` in modelList no longer leaks |
| 6 | `0.5` | `{"temperature": None}` | `temperature=0.5` | **Fixed** —
caller value no longer silently dropped by config null |
| 7 | `0.5` | `{"temperature": 0.7}` | `temperature=0.7` | Persisted
wins over caller (PR #698 precedence, unchanged) |

## Prior art

- #698 (ROB-1797) ADDED the `setdefault` to fix "temperature is popped
and then lost for 2+ complete iterations, seen on bedrock with thinking
arguments." This PR preserves that behavior — row 3 in the matrix is the
exact regression that #698 fixed, and a test locks it in.
- #808 upgraded LiteLLM so `drop_params` handles provider-specific
stripping. This PR keeps that contract intact; it only prevents Holmes
from manufacturing a `None` key that LiteLLM's drop logic isn't
guaranteed to catch.

## Tests

New file `tests/core/test_llm_completion_temperature.py` with 7 tests,
one per matrix row. Each test mocks `holmes.core.llm.litellm.completion`
and asserts on the `call_args.kwargs` the mock received — no network, no
real LLM calls.

Verified locally:

| | master | this PR |
|---|---|---|
| Rows 1, 2, 3, 7 | ✅ pass | ✅ pass |
| Rows 4, 5, 6    | ❌ fail | ✅ pass |
| `poetry run pytest tests/core -m "not llm"` | — | 389 passed, 14
skipped (env-gated), 0 regressions |

Note: row 4 also regressed on master — the old
`setdefault("temperature", None)` materialized a `None` key even when
`self.args` was empty. The fix handles rows 4, 5, and 6 in one stroke.

## Scope

One bug, one PR. Intentionally does **not**:
- Generalize to `top_p`, `max_tokens`, `stop`, or other params — if
those have sibling bugs they should be separate PRs.
- Add a new config surface (no `strip_temperature` flag, no modelList
schema change).
- Change LiteLLM's role as the authority for provider-specific param
stripping.

## Test plan

- [x] 7 new unit tests pass on the fix; rows 4, 5, 6 fail on master
(verified)
- [x] `poetry run pytest tests/core -m "not llm" --no-cov` — no
regressions
- [ ] Maintainer review — any preference for fix location or scope

cc / relevant commit authors: @RoiGlinik (#698), @aantn (#808)

---------

Signed-off-by: alam0rt <sam@samlockart.com>
moshemorad added a commit that referenced this pull request Jun 14, 2026
## Problem

A customer reported Holmes responses getting cut off mid-answer. Their
chat metadata showed:

```json
"max_completion_tokens_per_call": 4096,
"finish_reason": "length",
"max_output_tokens": 64000,
"max_tokens": 1000000
```

`completion_tokens` landed at exactly 4096 with `finish_reason:
"length"` — the model hit a hard 4096 output cap, even though Holmes
computed (and reported) a 64000-token output budget.

**Root cause:** `get_maximum_output_token()` is used to reserve output
space during input budgeting and compaction
(`input_context_window_limiter.py`, `compaction.py`), but it was never
sent on the actual request — `DefaultLLM.completion()` passed no
`max_tokens` to litellm. litellm then falls back to provider defaults.
For Anthropic-family models, litellm resolves the default from its cost
map; when the model name isn't in the map (proxy aliases, custom
gateways — the same situation that makes users configure
`max_context_size` by hand), it falls back to
`DEFAULT_ANTHROPIC_CHAT_MAX_TOKENS = 4096`. Long answers get silently
truncated while the metadata claims a 64000 budget.

## Fix

`DefaultLLM.completion()` now always sends an explicit `max_tokens`,
using the same value input budgeting already reserves. Precedence:

1. Explicit `max_tokens` / `max_completion_tokens` in model args —
always wins (and a user-set `max_completion_tokens` blocks injection so
no conflicting pair is sent).
2. `OVERRIDE_MAX_OUTPUT_TOKEN` env var — now actually reaches the
request instead of only affecting compaction math.
3. Computed: `min(64000, context_window / 5)`, capped by the model's
`max_output_tokens` from litellm's cost map when known.

`max_tokens: null` / `max_completion_tokens: null` config sentinels are
stripped, mirroring the existing `temperature: null` handling (PR #698
semantics).

Provider safety: litellm 1.83.7 translates `max_tokens` per provider —
`max_completion_tokens` for OpenAI o-series/gpt-5 reasoning models,
`maxTokens` for Bedrock converse, `maxOutputTokens` for Gemini — so
sending it is safe across providers (verified against the pinned
litellm).

## Changes

- `holmes/core/llm.py` — inject `max_tokens` in
`DefaultLLM.completion()` (all three call sites benefit: agentic loop,
compaction, fast-model summarization)
- `tests/core/test_llm_completion_max_tokens.py` — new behavior-matrix
tests, including a reproduction of the customer scenario (unknown model
+ `max_context_size: 1000000` → `max_tokens: 64000`, not 4096)
- `tests/core/test_llm_completion_temperature.py`,
`tests/core/test_llm_completion_cache_control.py` — test helpers that
bypass `__init__` now set `max_context_size`
- `docs/reference/context-management.md` — document the output token
limit and its resolution order

## Testing

- `tests/core/test_llm_completion_max_tokens.py` — 8 new tests, all pass
- Full non-LLM suite: **2487 passed, 91 skipped** (skips are
missing-credential environment skips, pre-existing)

## Note for operators

Models unknown to litellm's cost map now receive a computed `max_tokens`
instead of none. If the computed value exceeds what the upstream model
actually supports, the provider may reject the request with an explicit
error instead of silently truncating at its default — set
`OVERRIDE_MAX_OUTPUT_TOKEN` or `max_tokens` in model args to the correct
value for the model.

https://claude.ai/code/session_01LPu4dG5LMjRgQgpBUsWhKw

---
_Generated by [Claude
Code](https://claude.ai/code/session_01LPu4dG5LMjRgQgpBUsWhKw)_

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Release Notes

* **Bug Fixes**
* LLM completion requests now consistently include an explicit
output-token budget to help prevent mid-response truncation.
* Output-token cap resolution is updated with clearer precedence and
correct behavior when `max_tokens` and `max_completion_tokens` are both
present.

* **Documentation**
* Context management docs now include an “Output Token Limit” section
explaining the enforced cap and its resolution order.

* **Tests**
* Added coverage to verify output-token limit injection, stripping,
environment overrides, and model-specific capping.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Claude <noreply@anthropic.com>
Co-authored-by: Claude <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants