Skip to content

Remove additional instructions support - #1477

Merged
naomi-robusta merged 2 commits into
masterfrom
claude/add-env-var-check-Wker9
Feb 5, 2026
Merged

naomi-robusta merged 2 commits into
masterfrom
claude/add-env-var-check-Wker9

Conversation

@naomi-robusta

@naomi-robusta naomi-robusta commented Feb 4, 2026 •

Copy link
Copy Markdown
Collaborator

Remove additional_instructions support

Summary by CodeRabbit

  • Breaking Changes
    • Removed the additionalInstructions configuration from tool and toolset definitions. Output guidance and post-processing instructions are no longer applied; tool outputs are returned as-is. Update any custom toolset configurations to remove these blocks.

@linux-foundation-easycla

linux-foundation-easycla Bot commented Feb 4, 2026 •

Copy link
Copy Markdown

CLA Not Signed

@github-actions

github-actions Bot commented Feb 4, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

📜 Run @ 2650c91 (#21709914696)

✅ Results of HolmesGPT evals

Automatically triggered by commit 2650c91 on branch claude/add-env-var-check-Wker9

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 38.1s 6 10 $0.2360
✅ 101_loki_historical_logs_pod_deleted 50.8s 5 11 $0.2622
✅ 111_pod_names_contain_service 31.0s 4 10 $0.2000
✅ 112_find_pvcs_by_uuid 31.4s 5 7 $0.2198
✅ 12_job_crashing 36.7s 5 12 $0.2417
✅ 176_network_policy_blocking_traffic_no_runbooks 49.5s 7 16 $0.2915
✅ 24_misconfigured_pvc 35.4s 5 13 $0.2273
✅ 43_current_datetime_from_prompt 6.7s 1 — $0.1060
✅ 61_exact_match_counting 17.6s 4 4 $0.1586
Total 33.0s avg 4.7 avg 10.4 avg $1.9431
📜 Run @ 0b070c9 (#21670750389)

✅ Results of HolmesGPT evals

Automatically triggered by commit 0b070c9 on branch claude/add-env-var-check-Wker9

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 37.2s 5 12 $0.2420
✅ 101_loki_historical_logs_pod_deleted 66.2s 8 16 $0.3503
✅ 111_pod_names_contain_service 34.1s 5 11 $0.2186
✅ 112_find_pvcs_by_uuid 43.0s 7 9 $0.2787
✅ 12_job_crashing 31.6s 5 9 $0.2136
✅ 176_network_policy_blocking_traffic_no_runbooks 49.4s 6 17 $0.2870
✅ 24_misconfigured_pvc 41.2s 6 14 $0.2449
✅ 43_current_datetime_from_prompt 6.1s 1 — $0.1050
✅ 61_exact_match_counting 17.6s 4 4 $0.1599
Total 36.3s avg 5.2 avg 11.5 avg $2.1000
📜 Run @ 1727cce (#21670552574)

✅ Results of HolmesGPT evals

Automatically triggered by commit 1727cce on branch claude/add-env-var-check-Wker9

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 34.0s 5 11 $0.2313
✅ 101_loki_historical_logs_pod_deleted 36.1s 5 8 $0.2213
✅ 111_pod_names_contain_service 29.0s 4 9 $0.2004
✅ 112_find_pvcs_by_uuid 42.2s 7 10 $0.2792
✅ 12_job_crashing 46.2s 7 17 $0.3017
✅ 176_network_policy_blocking_traffic_no_runbooks 56.8s 8 18 $0.3273
✅ 24_misconfigured_pvc 42.7s 7 14 $0.2485
✅ 43_current_datetime_from_prompt 5.2s 1 — $0.1050
✅ 61_exact_match_counting 13.1s 3 2 $0.1417
Total 33.9s avg 5.2 avg 11.1 avg $2.0565
📜 Run @ 1c1160b (#21666478936)

✅ Results of HolmesGPT evals

Automatically triggered by commit 1c1160b on branch claude/add-env-var-check-Wker9

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 35.9s 5 13 $0.2472
✅ 101_loki_historical_logs_pod_deleted 41.9s 5 10 $0.2371
✅ 111_pod_names_contain_service 34.3s 5 11 $0.2213
✅ 112_find_pvcs_by_uuid 36.6s 6 6 $0.2207
✅ 12_job_crashing 36.0s 5 12 $0.2394
✅ 176_network_policy_blocking_traffic_no_runbooks 43.1s 5 15 $0.2688
✅ 24_misconfigured_pvc 32.4s 5 13 $0.2258
✅ 43_current_datetime_from_prompt 5.5s 1 — $0.1054
✅ 61_exact_match_counting 17.5s 4 4 $0.1594
Total 31.5s avg 4.6 avg 10.5 avg $1.9251
📜 Run @ c037393 (#21666237870)

✅ Results of HolmesGPT evals

Automatically triggered by commit c037393 on branch claude/add-env-var-check-Wker9

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 33.9s 5 10 $0.2241
✅ 101_loki_historical_logs_pod_deleted 47.3s 6 11 $0.2547
✅ 111_pod_names_contain_service 35.1s 5 11 $0.2218
✅ 112_find_pvcs_by_uuid 51.3s 10 9 $0.2886
✅ 12_job_crashing 41.8s 6 14 $0.2666
✅ 176_network_policy_blocking_traffic_no_runbooks 43.4s 6 15 $0.2751
✅ 24_misconfigured_pvc 33.5s 5 13 $0.2219
✅ 43_current_datetime_from_prompt 5.1s 1 — $0.1047
✅ 61_exact_match_counting 17.6s 4 4 $0.1594
Total 34.3s avg 5.3 avg 10.9 avg $2.0168

✅ Results of HolmesGPT evals

Automatically triggered by commit de6bead on branch claude/add-env-var-check-Wker9

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 34.3s 5 11 $0.2317
✅ 101_loki_historical_logs_pod_deleted 38.5s 5 9 $0.2345
✅ 111_pod_names_contain_service 34.4s 5 11 $0.2242
✅ 112_find_pvcs_by_uuid 32.2s 6 6 $0.2235
✅ 12_job_crashing 34.2s 5 12 $0.2451
✅ 176_network_policy_blocking_traffic_no_runbooks 48.0s 7 16 $0.2856
✅ 24_misconfigured_pvc 36.1s 6 13 $0.2360
✅ 43_current_datetime_from_prompt 5.3s 1 — $0.1056
✅ 61_exact_match_counting 16.8s 4 4 $0.1595
Total 31.1s avg 4.9 avg 10.2 avg $1.9458
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/add-env-var-check-Wker9 -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, easy, elasticsearch, embeds, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/add-env-var-check-Wker9 -f markers=regression -f filter=

@github-actions

github-actions Bot commented Feb 4, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker image ready for ea2aa9c (built in 5m 15s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use this tag to pull the image for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:ea2aa9c
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:ea2aa9c me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:ea2aa9c
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:ea2aa9c

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:ea2aa9c

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:ea2aa9c

@coderabbitai

coderabbitai Bot commented Feb 4, 2026 •

Copy link
Copy Markdown
Contributor

Walkthrough

Removed the additional_instructions feature across the tools system: the field and all code that applied or propagated post-processing instructions were deleted from Tool, Toolset, and YAMLToolsetFromConfig; tools now return raw output without instruction-based modifications. (50 words)

Changes

Cohort / File(s) Summary
Core tools implementation
holmes/core/tools.py
Deleted additional_instructions fields from Tool, Toolset, and ToolsetYamlFromConfig. Removed __apply_additional_instructions helper and related error handling. Tool._invoke and YAMLTool._invoke no longer apply instructions and always return raw_output. Preprocessing no longer injects or forwards additional instructions.
Documentation & examples
docs/data-sources/custom-toolsets.md, holmes/core/toolset_manager.py
Removed additionalInstructions blocks from toolset YAML examples and docstring examples; documentation no longer references the removed field.

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~12 minutes

Possibly related PRs

Suggested reviewers

  • arikalon1
  • moshemorad
  • Avi-Robusta
🚥 Pre-merge checks | ✅ 2 | ❌ 1
❌ Failed checks (1 warning)
Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title 'Remove additional instructions support' directly and clearly summarizes the main change: eliminating the additional_instructions feature from Tool, Toolset, and ToolsetYamlFromConfig classes.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@netlify

netlify Bot commented Feb 4, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit de6bead
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/698481526661a30008d6bad1
😎 Deploy Preview https://deploy-preview-1477--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@github-actions

github-actions Bot commented Feb 4, 2026 •

Copy link
Copy Markdown
Contributor

🔬 CLI Performance Benchmark

🟡 Startup Time (no LLM)

Measures holmes version execution time (imports + initialization)

Metric PR Master Change
Cold Start 9.99s 8.45s +18.2%
Warm Mean 4.70s 3.95s +19.0%
Warm Min 4.60s 3.93s
Warm Max 4.82s 3.99s

🟡 Full CLI with LLM

Measures holmes ask execution time (OpenRouter + Haiku 4.5)

Metric PR Master Change
Cold Start 26.61s 23.58s +12.8%
Warm Mean 7.76s 6.47s +19.9%
Warm Min 7.55s 6.43s
Warm Max 8.04s 6.49s

PR: ea2aa9c8 | Master: 90483b4b | Iterations: 5

@naomi-robusta
naomi-robusta force-pushed the claude/add-env-var-check-Wker9 branch from 9a01f55 to d07a816 Compare February 4, 2026 09:34

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Fix all issues with AI agents
In `@holmes/core/tools.py`:
- Around line 501-514: The logger currently emits the full contents of
self.additional_instructions (see the logger.info call near
enable_additional_instructions and the __apply_additional_instructions usage),
which may leak secrets; change the logging to avoid printing the instruction
string itself and instead log safe metadata such as the toolset name
(toolset_name), existence flag, and/or instruction length or a fixed
placeholder, and keep the call sites around enable_additional_instructions and
the error branch unchanged otherwise so output_with_instructions is produced by
__apply_additional_instructions when allowed.

Comment thread holmes/core/tools.py Outdated
@naomi-robusta
naomi-robusta force-pushed the claude/add-env-var-check-Wker9 branch 2 times, most recently from c037393 to 8e3cb92 Compare February 4, 2026 09:42
@naomi-robusta
naomi-robusta marked this pull request as ready for review February 4, 2026 09:42
@naomi-robusta
naomi-robusta force-pushed the claude/add-env-var-check-Wker9 branch from 8e3cb92 to 1c1160b Compare February 4, 2026 09:44

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Fix all issues with AI agents
In `@holmes/common/env_vars.py`:
- Around line 155-156: load_bool currently relies on json.loads which mis-parses
common boolean strings ("yes"/"no") and returns ints for "1"/"0", causing
type-safety issues for ENABLE_ADDITIONAL_INSTRUCTIONS; update load_bool to
accept and normalize common boolean representations (e.g., case-insensitive
"true"/"false", "yes"/"no", "1"/"0", and actual bools) and return Optional[bool]
consistently (None when unset/invalid), and add an explicit type annotation to
the constant as ENABLE_ADDITIONAL_INSTRUCTIONS: Optional[bool] =
load_bool("ENABLE_ADDITIONAL_INSTRUCTIONS", False).
🧹 Nitpick comments (1)
holmes/common/env_vars.py (1)

155-156: Add type annotation to maintain consistency with Python typing guidelines.

The load_bool() function returns Optional[bool], so the suggested type annotation is correct. However, note that many similar flags in this file (e.g., ENABLE_TELEMETRY, DEVELOPMENT_MODE, ROBUSTA_AI) also use load_bool() without type annotations. Consider applying this annotation across all module-level globals for consistency with the coding guideline requiring type hints throughout Python code.

♻️ Suggested annotation
-ENABLE_ADDITIONAL_INSTRUCTIONS = load_bool("ENABLE_ADDITIONAL_INSTRUCTIONS", False)
+ENABLE_ADDITIONAL_INSTRUCTIONS: Optional[bool] = load_bool(
+    "ENABLE_ADDITIONAL_INSTRUCTIONS", False
+)

Comment thread holmes/common/env_vars.py Outdated
@naomi-robusta
naomi-robusta force-pushed the claude/add-env-var-check-Wker9 branch from 1c1160b to 1727cce Compare February 4, 2026 11:58
@naomi-robusta naomi-robusta changed the title Add feature flag for additional instructions in tool invocation Remove additional instructions support Feb 4, 2026
@naomi-robusta
naomi-robusta force-pushed the claude/add-env-var-check-Wker9 branch 3 times, most recently from 4aa576d to 0b070c9 Compare February 4, 2026 12:04
Signed-off-by: Naomi Caren <naomi@robusta.dev>
@naomi-robusta
naomi-robusta force-pushed the claude/add-env-var-check-Wker9 branch from 0b070c9 to 2650c91 Compare February 5, 2026 11:34
@naomi-robusta
naomi-robusta merged commit 7faf9ee into master Feb 5, 2026
18 of 20 checks passed
@naomi-robusta
naomi-robusta deleted the claude/add-env-var-check-Wker9 branch February 5, 2026 12:04
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants