Skip to content

Move exception tracebacks to debug level to improve error visibility - #1570

Open
aantn wants to merge 1 commit into
masterfrom
claude/fix-env-var-visibility-ILOFO
Open

aantn wants to merge 1 commit into
masterfrom
claude/fix-env-var-visibility-ILOFO

Conversation

@aantn

@aantn aantn commented Feb 15, 2026 •

Copy link
Copy Markdown
Collaborator

Tracebacks from LLM API errors (e.g. litellm) were printed to console
via logging.error(exc_info=True), often spanning 50+ lines with chained
exceptions. This buried useful info (env var hints, warnings) above the
traceback where users couldn't see it. Now tracebacks are logged at
debug level (visible with -v flag) and only the clean error message is
shown by default. A tip about -v is shown in interactive mode.

https://claude.ai/code/session_01W597xJtLq8XHce8i6F9LKg
Signed-off-by: Claude noreply@anthropic.com

Summary by CodeRabbit

Release Notes

  • Chores
    • Enhanced error messages to include exception details for better clarity.
    • Reorganized debug output to provide full tracebacks at verbose log level.
    • Added guidance to use the verbose flag (-v) for detailed error diagnostics.

Tracebacks from LLM API errors (e.g. litellm) were printed to console
via logging.error(exc_info=True), often spanning 50+ lines with chained
exceptions. This buried useful info (env var hints, warnings) above the
traceback where users couldn't see it. Now tracebacks are logged at
debug level (visible with -v flag) and only the clean error message is
shown by default. A tip about -v is shown in interactive mode.

https://claude.ai/code/session_01W597xJtLq8XHce8i6F9LKg
Signed-off-by: Claude <noreply@anthropic.com>
@netlify

netlify Bot commented Feb 15, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit deaf601
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/69923dc49970440008b750df
😎 Deploy Preview https://deploy-preview-1570--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@github-actions

github-actions Bot commented Feb 15, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker image ready for c13ed42 (built in 4m 4s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use this tag to pull the image for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:c13ed42
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:c13ed42 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:c13ed42
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:c13ed42

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:c13ed42

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:c13ed42

@github-actions

github-actions Bot commented Feb 15, 2026 •

Copy link
Copy Markdown
Contributor

✅ Results of HolmesGPT evals

Automatically triggered by commit deaf601 on branch claude/fix-env-var-visibility-ILOFO

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 31.8s 5 11 $0.2358
✅ 101_loki_historical_logs_pod_deleted 44.1s 6 11 $0.2704
✅ 111_pod_names_contain_service 31.4s 5 11 $0.2233
✅ 112_find_pvcs_by_uuid 40.6s 7 10 $0.2847
✅ 12_job_crashing 25.7s 4 7 $0.2048
✅ 176_network_policy_blocking_traffic_no_runbooks 40.8s 6 16 $0.2812
✅ 24_misconfigured_pvc 32.3s 5 13 $0.2278
✅ 43_current_datetime_from_prompt 5.7s 1 — $0.1073
✅ 61_exact_match_counting 13.2s 3 2 $0.1454
Total 29.5s avg 4.7 avg 10.1 avg $1.9808
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/fix-env-var-visibility-ILOFO -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, easy, elasticsearch, embeds, fast, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/fix-env-var-visibility-ILOFO -f markers=regression -f filter=

@coderabbitai

coderabbitai Bot commented Feb 15, 2026 •

Copy link
Copy Markdown
Contributor

Walkthrough

Reorganizes error logging across multiple modules to move tracebacks from ERROR level to DEBUG level while explicitly including exception messages in error-level logs. Control flow and user-facing error messages remain unchanged. Adds exception chaining in some paths and user-guidance tips in interactive mode.

Changes

Cohort / File(s) Summary
LLM and Tool Call Error Handling
holmes/core/tool_calling_llm.py, holmes/core/tools.py
Replaces exc_info=True on error logs with dedicated debug-level traceback logs. Augments error messages to include exception strings and adds exception chaining in Azure BadRequestError path.
Main Command Handlers
holmes/main.py
Shifts traceback logging from error to debug level across fetch command handlers (alertmanager, jira, github, pagerduty, opsgenie, ticket) while explicitly including exception messages in error logs.
Interactive Mode
holmes/interactive.py
Moves tracebacks to debug level and augments error messages with user-guidance tips suggesting -v flag for verbose output in save_conversation_to_file and main interactive loop.
Stream Utilities
holmes/utils/stream.py
Replaces error-level traceback logging with separate debug-level traceback logs in stream formatters (stream_investigate_formatter, stream_chat_formatter) while preserving error handling logic.

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~12 minutes

Possibly related PRs

Suggested reviewers

  • arikalon1
  • moshemorad
🚥 Pre-merge checks | ✅ 3 | ❌ 1
❌ Failed checks (1 warning)
Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 57.14% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title directly and concisely summarizes the primary change: moving exception tracebacks from error to debug level. It accurately reflects the main objective of the PR without being vague or misleading.
Merge Conflict Detection ✅ Passed ✅ No merge conflicts detected when merging into master

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions

Copy link
Copy Markdown
Contributor

🔬 CLI Performance Benchmark

🟡 Startup Time (no LLM)

Measures holmes version execution time (imports + initialization)

Metric PR Master Change
Cold Start 10.42s 9.65s +8.0%
Warm Mean 4.79s 4.51s +6.1%
Warm Min 4.76s 4.48s
Warm Max 4.81s 4.55s

🟡 Full CLI with LLM

Measures holmes ask execution time (OpenRouter + Haiku 4.5)

Metric PR Master Change
Cold Start 28.12s 14.13s +99.0%
Warm Mean 7.41s 6.92s +7.2%
Warm Min 7.16s 6.56s
Warm Max 7.59s 7.29s

PR: c13ed426 | Master: 181daec0 | Iterations: 5

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Fix all issues with AI agents
In `@holmes/interactive.py`:
- Line 1438: The console.print call uses an unnecessary f-string prefix for a
plain literal; locate the console.print call that currently reads
console.print(f"[dim]Tip: Use -v for more details[/dim]") and remove the leading
"f" so it becomes a normal string literal; update the call in the interactive.py
function/method where console.print is used and re-run the linter (Ruff) to
ensure the F541 warning is resolved.
🧹 Nitpick comments (2)
holmes/interactive.py (1)

1079-1081: Inconsistent exc_info=e vs exc_info=True across the PR.

This file passes the exception instance (exc_info=e), while tools.py and tool_calling_llm.py use exc_info=True. Both work in Python 3, but within an except block exc_info=True is idiomatic and consistent. Same applies to line 1436.

Suggested fix for consistency
-        logging.debug("Full traceback:", exc_info=e)
+        logging.debug("Full traceback:", exc_info=True)
holmes/main.py (1)

475-476: Same exc_info=e vs exc_info=True inconsistency noted here.

All six debug-traceback call sites in this file use exc_info=e. Consider using exc_info=True for consistency with tools.py and tool_calling_llm.py. Both are valid, but uniformity across the codebase aids readability.

Comment thread holmes/interactive.py
console.print(f"[bold {ERROR_COLOR}]Error: {e}[/bold {ERROR_COLOR}]")
logging.debug("Full traceback for interactive mode error:", exc_info=e)
console.print(f"\n[bold {ERROR_COLOR}]Error: {e}[/bold {ERROR_COLOR}]")
console.print(f"[dim]Tip: Use -v for more details[/dim]")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Remove extraneous f prefix — no placeholders in this string.

Ruff F541 flags this correctly. The string has no interpolation.

Proposed fix
-            console.print(f"[dim]Tip: Use -v for more details[/dim]")
+            console.print("[dim]Tip: Use -v for more details[/dim]")
🧰 Tools
🪛 Ruff (0.15.0)

[error] 1438-1438: f-string without any placeholders

Remove extraneous f prefix

(F541)

🤖 Prompt for AI Agents
In `@holmes/interactive.py` at line 1438, The console.print call uses an
unnecessary f-string prefix for a plain literal; locate the console.print call
that currently reads console.print(f"[dim]Tip: Use -v for more details[/dim]")
and remove the leading "f" so it becomes a normal string literal; update the
call in the interactive.py function/method where console.print is used and
re-run the linter (Ruff) to ensure the F541 warning is resolved.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants