Skip to content

Clarify command execution policy in bash instructions - #1767

Closed
aantn wants to merge 2 commits into
masterfrom
claude/add-kubectl-exec-8IhvR
Closed

aantn wants to merge 2 commits into
masterfrom
claude/add-kubectl-exec-8IhvR

Conversation

@aantn

@aantn aantn commented Mar 14, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Updated the bash instructions template to provide clearer guidance on command execution policies, distinguishing between pre-approved commands, other commands that require user approval, and blocked commands.

Key Changes

  • Renamed "Allowed Commands" section to "Pre-approved Commands" for better clarity
  • Added new "Other Commands" section that explicitly explains:
    • Commands not in the pre-approved list can still be used
    • They will be sent to the user for approval before execution
    • Users should not refuse to run commands just because they're not pre-approved
  • Updated "Blocked Commands" section description to clarify these commands cannot be executed "even with user approval"

Implementation Details

These changes improve the user experience by:

  • Making the three-tier command policy more explicit (pre-approved → approval-required → blocked)
  • Reducing confusion about whether non-pre-approved commands can be executed
  • Encouraging users to attempt commands that aren't pre-approved, with the understanding that approval will be requested

https://claude.ai/code/session_01PUmi6foSSLQJJNrS8Z8yxH

Summary by CodeRabbit

  • Documentation
    • Updated command handling instructions to clarify the distinction between pre-approved and other commands.
    • Added explanatory guidance on using non-pre-approved commands with user approval.
    • Enhanced descriptions of blocked commands to clarify that blocking applies regardless of approval status.

The bash instructions prompt listed "Allowed Commands" which the LLM
interpreted as the ONLY commands it could use. Rename to "Pre-approved
Commands" and add an explicit "Other Commands" section explaining that
non-listed commands can still be attempted and will trigger user approval.

https://claude.ai/code/session_01PUmi6foSSLQJJNrS8Z8yxH
Signed-off-by: Claude <noreply@anthropic.com>
@claude

claude Bot commented Mar 14, 2026

Copy link
Copy Markdown
Contributor

Claude Code Review

This repository is configured for manual code reviews. Comment @claude review to trigger a review.

@github-actions

github-actions Bot commented Mar 14, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

📜 Run @ 75757f8 (#23096436528)

✅ Results of HolmesGPT evals

Automatically triggered by commit 75757f8 on branch claude/add-kubectl-exec-8IhvR

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 10/10 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 26.8s 5 10 $0.2293 102,439 100,694 22,612 1,745 601 76,892 23,802 — —
✅ 101_loki_historical_logs_pod_deleted 46.9s 7 14 $0.3253 162,381 159,337 27,670 3,044 653 129,352 29,985 — —
✅ 111_pod_names_contain_service 27.9s 5 9 $0.2219 100,415 98,643 22,123 1,772 450 76,291 22,352 — —
✅ 112_find_pvcs_by_uuid 19.1s 4 4 $0.1861 77,365 76,218 20,736 1,147 519 55,470 20,748 — —
✅ 12_job_crashing 29.1s 5 14 $0.2531 112,429 110,463 25,220 1,966 545 84,427 26,036 — —
✅ 176_network_policy_blocking_traffic_no_runbooks 38.0s 7 16 $0.3012 157,556 154,933 26,566 2,623 734 127,156 27,777 — —
✅ 227_count_configmaps_per_namespace[0] 25.0s 6 10 $0.2122 113,358 111,911 20,582 1,447 461 90,951 20,960 — —
✅ 24_misconfigured_pvc 31.3s 6 14 $0.2491 121,189 119,055 22,904 2,134 630 95,152 23,903 — —
✅ 43_current_datetime_from_prompt 4.3s 1 — $0.1086 17,051 16,943 16,943 108 108 0 16,943 — —
✅ 61_exact_match_counting 13.6s 4 4 $0.1514 71,194 70,747 18,223 447 229 52,512 18,235 — —
Total 26.2s avg 5.0 avg 10.6 avg $2.2381 1,035,377 1,018,944 27,670 16,433 734 788,203 230,741 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 59 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)

✅ Results of HolmesGPT evals

Automatically triggered by commit 5edc640 on branch claude/add-kubectl-exec-8IhvR

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 10/10 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 28.3s 6 12 $0.2516 123,653 121,732 23,714 1,921 710 96,614 25,118 — —
✅ 101_loki_historical_logs_pod_deleted 39.8s 7 14 $0.2916 150,909 148,095 25,071 2,814 654 122,356 25,739 — —
✅ 111_pod_names_contain_service 23.1s 4 8 $0.2028 78,773 77,224 21,487 1,549 616 55,301 21,923 — —
✅ 112_find_pvcs_by_uuid 20.4s 5 4 $0.3114 95,133 93,968 20,727 1,165 287 53,037 40,931 — —
✅ 12_job_crashing 29.1s 5 13 $0.2549 110,195 108,043 24,851 2,152 583 82,259 25,784 — —
✅ 176_network_policy_blocking_traffic_no_runbooks 42.9s 8 16 $0.4074 177,082 174,619 26,156 2,463 438 129,128 45,491 — —
✅ 227_count_configmaps_per_namespace[0] 20.1s 5 10 $0.1993 94,915 93,584 20,458 1,331 595 72,774 20,810 — —
✅ 24_misconfigured_pvc 31.7s 7 13 $0.2596 142,842 140,743 22,920 2,099 459 116,711 24,032 — —
✅ 43_current_datetime_from_prompt 3.8s 1 — $0.1088 17,083 16,972 16,972 111 111 0 16,972 — —
✅ 61_exact_match_counting 12.2s 4 4 $0.1514 71,295 70,853 18,246 442 225 52,595 18,258 — —
Total 25.1s avg 5.2 avg 10.4 avg $2.4389 1,061,880 1,045,833 26,156 16,047 710 780,775 265,058 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 59 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/add-kubectl-exec-8IhvR -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
tags: regression

Or with more options (one per line):

/eval
model: gpt-4o
tags: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
tags: regression
Option Description
model Model(s) to test (default: same as automatic runs)
tags Pytest tags / markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

Option 3: Add PR labels to include extra evals in automatic regression runs:

Label Effect
evals-tag-<name> Run tests with tag <name> alongside regression
evals-id-<name> Run a specific eval by test ID

Examples: evals-tag-easy, evals-id-09_crashpod

🏷️ Valid tags

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/add-kubectl-exec-8IhvR -f markers=regression -f filter=

@github-actions

github-actions Bot commented Mar 14, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker images ready for ded8b9be (built in 6m 17s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use these tags to pull the images for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:ded8b9be
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:ded8b9be me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:ded8b9be
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:ded8b9be
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:ded8b9be
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:ded8b9be me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:ded8b9be
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:ded8b9be

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:ded8b9be \
  --set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set operator.image=holmes-operator-dev:ded8b9be

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:ded8b9be \
  --set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.operator.image=holmes-operator-dev:ded8b9be

@coderabbitai

coderabbitai Bot commented Mar 14, 2026 •

Copy link
Copy Markdown
Contributor

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: fd8afdd3-4ae7-4d15-bb4f-0bf413ab99f0

📥 Commits

Reviewing files that changed from the base of the PR and between 0cc94a1 and 5edc640.

📒 Files selected for processing (1)
  • holmes/plugins/toolsets/bash/bash_instructions.jinja2

Walkthrough

Updated bash_instructions.jinja2 template to reframe command authorization terminology and guidance: renames "Allowed Commands" to "Pre-approved Commands," introduces an "Other Commands" section for non-pre-approved commands with user approval, clarifies that blocked commands cannot be executed regardless of user approval, and adds explanatory text promoting pre-approved commands and built-in tools.

Changes

Cohort / File(s) Summary
Bash Instructions Documentation
holmes/plugins/toolsets/bash/bash_instructions.jinja2
Renamed "Allowed Commands" section to "Pre-approved Commands"; added new "Other Commands" section explaining usage of non-pre-approved commands with user approval; extended "Blocked Commands" phrasing to clarify blocking persists regardless of user approval; added descriptive paragraph emphasizing preference for pre-approved commands and built-in tools.

Estimated code review effort

🎯 1 (Trivial) | ⏱️ ~3 minutes

Possibly related PRs

Suggested reviewers

  • Sheeproid
🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately summarizes the main change: clarifying command execution policy in bash instructions. It is concise, specific, and directly relates to the primary objective of the PR.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

📝 Coding Plan
  • Generate coding plan for human review comments

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@netlify

netlify Bot commented Mar 14, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit 5edc640
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/69b5cf2dee362d0008b3bc37
😎 Deploy Preview https://deploy-preview-1767--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

Tweak the "Other Commands" language so the LLM favors pre-approved
commands and built-in tools to minimize user interruptions, while still
knowing it can use non-listed commands when necessary.

https://claude.ai/code/session_01PUmi6foSSLQJJNrS8Z8yxH
Signed-off-by: Claude <noreply@anthropic.com>
@github-actions

github-actions Bot commented Mar 14, 2026 •

Copy link
Copy Markdown
Contributor

🔬 CLI Performance Benchmark

🟡 Startup Time (no LLM)

Measures holmes version execution time (imports + initialization)

Metric PR Master Change
Cold Start 11.14s 11.16s -0.2%
Warm Mean 5.23s 4.71s +11.1%
Warm Min 5.16s 4.65s
Warm Max 5.32s 4.79s

🟡 Full CLI with LLM

Measures holmes ask execution time (OpenRouter + Haiku 4.5)

Metric PR Master Change
Cold Start 21.41s 31.65s -32.4%
Warm Mean 7.74s 7.32s +5.8%
Warm Min 7.70s 6.93s
Warm Max 7.84s 8.02s

PR: ded8b9be | Master: 0cc94a16 | Iterations: 5

@aantn aantn closed this Mar 14, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants