Skip to content

ROB-320 Claude/remove edit command step2 - #2202

Merged
RoiGlinik merged 3 commits into
masterfrom
claude/remove-edit-command-step2
Jun 18, 2026
Merged

RoiGlinik merged 3 commits into
masterfrom
claude/remove-edit-command-step2

Conversation

@RoiGlinik

@RoiGlinik RoiGlinik commented Jun 17, 2026 •

Copy link
Copy Markdown
Collaborator

Summary by CodeRabbit

Release Notes

  • Removed Features

    • Removed tool command editing capability during tool approval workflows. Tools now execute with their originally proposed commands without modification options.
  • Bug Fixes

    • Improved backward compatibility to silently handle legacy approval payloads containing command override values, preventing validation errors from older clients.

Three-part plan for hardening tool approval against the forged-approval
primitive reported in GHSA-6m4w-cmhp-f95f:

- tool-approval-tickets.md: original draft (raw HMAC), kept for context
- step-2-remove-edit-command.md: remove the edit-command flow so an
  approved tool call cannot be mutated between approval and execution
- step-3-tool-approval-tickets-jwt.md: signed approval tickets via JWT
  (HS256/PyJWT, 7-day TTL, collapsed reason codes)

Step 1 (deny_list-always-runs in _invoke) was considered and dropped:
requires_approval already filters DENIED commands before they reach the
user, so after Steps 2+3 close the mutation+forgery paths, the deny
check at _invoke with user_approved=True is unreachable in the
legitimate flow.

Signed-off-by: Roi Glinik <groi.tech@gmail.com>
Closes the post-approval mutation primitive: today the approval UI can
substitute any value for a bash tool_call's command argument after the
user approves the original, and the executor runs the substituted
command with user_approved=True (so bash_toolset's deny list is
skipped). An attacker who can submit a tool_decisions blob (forged
pending_approval, or a legitimate but malicious user) can take any
benign approved bash command and swap in rm -rf, kubectl delete
namespace, token exfil, etc.

Pure removal:
- holmes/core/models.py: drop ToolApprovalDecision.edit_command
- holmes/core/tool_calling_llm.py: drop the splice block in
  _execute_tool_decisions
- tests/test_tool_decision_edit_command.py: deleted
- tests/core/conversations_worker/integration/test_conversation_integration.py:
  drop test_approval_with_edit_command
- tests/test_edit_command_removed.py: regression test that an old
  client still POSTing edit_command (a) is silently accepted (Pydantic
  v2 default extra="ignore" drops the field) and (b) the original
  command runs, not the substituted one

FE coordination: the Robusta frontend's matching PR removes the inline
edit affordance and stops sending edit_command. Rollout order is
FE-first; this Holmes change is safe to deploy after the FE rolls.

Per specs/step-2-remove-edit-command.md.

Signed-off-by: Roi Glinik <groi.tech@gmail.com>
The three step-2/step-3 design docs were carried in the first commit
on this branch for review context. Removing them now — they live
locally outside the repo and the merged PR should reflect only the
code change. No content lost.

Signed-off-by: Roi Glinik <groi.tech@gmail.com>
@github-actions

github-actions Bot commented Jun 17, 2026 •

Copy link
Copy Markdown
Contributor

✅ Results of HolmesGPT evals

Automatically triggered by commit 551586c on branch claude/remove-edit-command-step2

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 14/14 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions Denied commands Src
✅ 09_crashpod 29.6s 4 8 $0.2101 69,983 68,304 19,535 1,679 716 48,195 20,109 125 — — src
✅ 101_loki_historical_logs_pod_deleted 39.4s 4 9 $0.2256 69,684 67,413 19,507 2,271 752 47,360 20,053 184 — — src
✅ 112_find_pvcs_by_uuid 13.9s 2 2 $0.1545 33,127 32,271 17,949 856 446 14,319 17,952 279 — — src
✅ 12_job_crashing 32.6s 4 10 $0.2224 72,605 70,697 20,624 1,908 559 49,787 20,910 192 — — src
✅ 176_network_policy_blocking_traffic_no_skills 30.3s 4 10 $0.2124 69,696 68,001 20,586 1,695 537 47,410 20,591 262 — — src
✅ 227_count_configmaps_per_namespace[0] 16.3s 3 6 $0.1491 46,347 45,637 16,653 710 437 28,980 16,657 29 — — src
✅ 243_pod_names_contain_service 26.4s 3 5 $0.1732 48,034 46,671 17,321 1,363 550 29,071 17,600 236 — — src
✅ 24_misconfigured_pvc 28.0s 4 10 $0.2029 67,906 66,372 18,917 1,534 590 46,480 19,892 48 — — src
✅ 254_elasticsearch_dr_test_log_check 62.4s 10 12 $0.2766 131,541 128,062 17,004 3,479 778 110,456 17,606 260 — — src
✅ 259_wrong_cluster_logs_confusion 66.8s 8 10 $0.2705 109,043 105,282 16,868 3,761 1,440 87,912 17,370 629 — — src
✅ 260_global_es_remote_cluster_logs 64.4s 9 13 $0.2891 130,867 127,277 17,874 3,590 863 107,928 19,349 201 — — src
✅ 43_current_datetime_from_prompt 5.2s 1 — $0.1017 14,429 14,306 14,306 123 123 0 14,306 79 — — src
✅ 51_logs_summarize_errors 21.6s 3 2 $0.1565 46,815 45,960 17,026 855 489 28,930 17,030 33 — — src
✅ 61_exact_match_counting 8.2s 2 1 $0.1146 29,148 28,927 14,641 221 152 14,283 14,644 34 — — src
Total 31.8s avg 4.4 avg 7.5 avg $2.7594 939,225 915,180 20,624 24,045 1,440 661,111 254,069 2,591 — —
Benchmark Comparison Details

Master baseline: latest master-* experiment (post-merge regression eval)
Status: 14 test/model combinations loaded

Benchmark baseline: latest ci-benchmark experiment on master
Status: 17 test/model combinations loaded

Time comparison (seconds):

Test case This branch master (1d ago) Δ vs master benchmark (3d ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 29.6s 28.0s ±0% — —
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 39.4s 42.5s ±0% — —
112_find_pvcs_by_uuid (opus-4.6) 📄 13.9s 12.4s ↑12% — —
12_job_crashing (opus-4.6) 📄 32.6s 37.3s ↓13% — —
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 30.3s 30.3s ±0% — —
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 16.3s 16.0s ±0% — —
243_pod_names_contain_service (opus-4.6) 📄 26.4s 32.0s ↓17% — —
24_misconfigured_pvc (opus-4.6) 📄 28.0s 32.3s ↓13% — —
254_elasticsearch_dr_test_log_check (opus-4.6) 📄 62.4s 68.7s ±0% — —
259_wrong_cluster_logs_confusion (opus-4.6) 📄 66.8s 62.8s ±0% — —
260_global_es_remote_cluster_logs (opus-4.6) 📄 64.4s 56.6s ↑14% — —
43_current_datetime_from_prompt (opus-4.6) 📄 5.2s 3.9s ↑33% — —
51_logs_summarize_errors (opus-4.6) 📄 21.6s 21.0s ±0% — —
61_exact_match_counting (opus-4.6) 📄 8.2s 7.5s ±0% — —
Total (all, n=14) 31.8s 32.2s — — —
Comparable (m=14, b=0) 31.8s 32.2s ±0% — —

Cost comparison:

Test case This branch master (1d ago) Δ vs master benchmark (3d ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 $0.2101 $0.2074 ±0% — —
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 $0.2256 $0.2300 ±0% — —
112_find_pvcs_by_uuid (opus-4.6) 📄 $0.1545 $0.1415 ±0% — —
12_job_crashing (opus-4.6) 📄 $0.2224 $0.2516 ↓12% — —
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 $0.2124 $0.2191 ±0% — —
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 $0.1491 $0.1500 ±0% — —
243_pod_names_contain_service (opus-4.6) 📄 $0.1732 $0.1936 ↓11% — —
24_misconfigured_pvc (opus-4.6) 📄 $0.2029 $0.2216 ±0% — —
254_elasticsearch_dr_test_log_check (opus-4.6) 📄 $0.2766 $0.3151 ↓12% — —
259_wrong_cluster_logs_confusion (opus-4.6) 📄 $0.2705 $0.2583 ±0% — —
260_global_es_remote_cluster_logs (opus-4.6) 📄 $0.2891 $0.2304 ↑25% — —
43_current_datetime_from_prompt (opus-4.6) 📄 $0.1017 $0.1017 ±0% — —
51_logs_summarize_errors (opus-4.6) 📄 $0.1565 $0.1563 ±0% — —
61_exact_match_counting (opus-4.6) 📄 $0.1146 $0.1146 ±0% — —
Total (all, n=14) $0.1971 $0.1994 — — —
Comparable (m=14, b=0) $0.1971 $0.1994 ±0% — —

Total tokens comparison:

Test case This branch master (1d ago) Δ vs master benchmark (3d ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 69,983 68,566 ±0% — —
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 69,684 70,455 ±0% — —
112_find_pvcs_by_uuid (opus-4.6) 📄 33,127 31,640 ±0% — —
12_job_crashing (opus-4.6) 📄 72,605 94,410 ↓23% — —
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 69,696 87,509 ↓20% — —
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 46,347 46,343 ±0% — —
243_pod_names_contain_service (opus-4.6) 📄 48,034 50,198 ±0% — —
24_misconfigured_pvc (opus-4.6) 📄 67,906 69,335 ±0% — —
254_elasticsearch_dr_test_log_check (opus-4.6) 📄 131,541 132,887 ±0% — —
259_wrong_cluster_logs_confusion (opus-4.6) 📄 109,043 106,101 ±0% — —
260_global_es_remote_cluster_logs (opus-4.6) 📄 130,867 68,015 ↑92% — —
43_current_datetime_from_prompt (opus-4.6) 📄 14,429 14,428 ±0% — —
51_logs_summarize_errors (opus-4.6) 📄 46,815 46,746 ±0% — —
61_exact_match_counting (opus-4.6) 📄 29,148 29,148 ±0% — —
Total (all, n=14) 67,088 65,413 — — —
Comparable (m=14, b=0) 67,088 65,413 ±0% — —

Cached tokens comparison:

Test case This branch master (1d ago) Δ vs master benchmark (3d ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 48,195 46,150 ±0% — —
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 47,360 48,004 ±0% — —
112_find_pvcs_by_uuid (opus-4.6) 📄 14,319 14,319 ±0% — —
12_job_crashing (opus-4.6) 📄 49,787 68,997 ↓28% — —
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 47,410 65,459 ↓28% — —
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 28,980 28,979 ±0% — —
243_pod_names_contain_service (opus-4.6) 📄 29,071 29,481 ±0% — —
24_misconfigured_pvc (opus-4.6) 📄 46,480 46,078 ±0% — —
254_elasticsearch_dr_test_log_check (opus-4.6) 📄 110,456 108,065 ±0% — —
259_wrong_cluster_logs_confusion (opus-4.6) 📄 87,912 86,113 ±0% — —
260_global_es_remote_cluster_logs (opus-4.6) 📄 107,928 48,444 ↑123% — —
43_current_datetime_from_prompt (opus-4.6) 📄 — — — — —
51_logs_summarize_errors (opus-4.6) 📄 28,930 28,925 ±0% — —
61_exact_match_counting (opus-4.6) 📄 14,283 14,283 ±0% — —
Total (all, n=14) 47,222 45,236 — — —
Comparable (m=13, b=0) 50,855 48,715 ±0% — —

Turns comparison:

Test case This branch master (1d ago) Δ vs master benchmark (3d ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 4 4 ±0% — —
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 4 4 ±0% — —
112_find_pvcs_by_uuid (opus-4.6) 📄 2 2 ±0% — —
12_job_crashing (opus-4.6) 📄 4 5 ↓20% — —
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 4 5 ↓20% — —
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 3 3 ±0% — —
243_pod_names_contain_service (opus-4.6) 📄 3 3 ±0% — —
24_misconfigured_pvc (opus-4.6) 📄 4 4 ±0% — —
254_elasticsearch_dr_test_log_check (opus-4.6) 📄 10 9 ↑11% — —
259_wrong_cluster_logs_confusion (opus-4.6) 📄 8 8 ±0% — —
260_global_es_remote_cluster_logs (opus-4.6) 📄 9 5 ↑80% — —
43_current_datetime_from_prompt (opus-4.6) 📄 1 1 ±0% — —
51_logs_summarize_errors (opus-4.6) 📄 3 3 ±0% — —
61_exact_match_counting (opus-4.6) 📄 2 2 ±0% — —
Total (all, n=14) 4.4 4.1 — — —
Comparable (m=14, b=0) 4.4 4.1 ±0% — —

Tool calls comparison:

Test case This branch master (1d ago) Δ vs master benchmark (3d ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 8 7 ↑14% — —
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 9 8 ↑12% — —
112_find_pvcs_by_uuid (opus-4.6) 📄 2 2 ±0% — —
12_job_crashing (opus-4.6) 📄 10 12 ↓17% — —
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 10 10 ±0% — —
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 6 6 ±0% — —
243_pod_names_contain_service (opus-4.6) 📄 5 7 ↓29% — —
24_misconfigured_pvc (opus-4.6) 📄 10 11 ±0% — —
254_elasticsearch_dr_test_log_check (opus-4.6) 📄 12 14 ↓14% — —
259_wrong_cluster_logs_confusion (opus-4.6) 📄 10 9 ↑11% — —
260_global_es_remote_cluster_logs (opus-4.6) 📄 13 7 ↑86% — —
43_current_datetime_from_prompt (opus-4.6) 📄 — — — — —
51_logs_summarize_errors (opus-4.6) 📄 2 2 ±0% — —
61_exact_match_counting (opus-4.6) 📄 1 1 ±0% — —
Total (all, n=14) 7.0 7.4 — — —
Comparable (m=13, b=0) 7.5 7.4 ±0% — —

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/remove-edit-command-step2 -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
tags: regression

Or with more options (one per line):

/eval
model: gpt-4o
tags: regression
id: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
tags: regression
Option Description
model Model(s) to test (default: same as automatic runs)
tags Pytest tags / markers (no default - runs all tests!)
id Eval ID / pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

Option 3: Add PR labels to include extra evals (applies to both automatic runs and /eval comments):

Label Effect
evals-tag-<name> Run tests with tag <name> alongside regression
evals-id-<name> Run a specific eval by test ID
evals-model-<name> Override the model (use model list name, e.g. sonnet-4.5)

Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5

🏷️ Valid tags

benchmark, chain-of-causation, compaction, confluence, context_window, conversation_worker, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, images, integration, kafka, kubernetes, leaked-information, logs, loki, manual, mcp, medium, metrics, multi-cluster, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, skills, slackbot, storage, token-limit, toolset-limitation, traces, transparency, victorialogs

🤖 Valid models

deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, fable-5, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, gpt-5.5, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, opus-4.7, opus-4.8, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/remove-edit-command-step2 -f markers=regression -f filter=

@coderabbitai

coderabbitai Bot commented Jun 17, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

Walkthrough

The PR removes the edit_command feature from the tool approval workflow. The edit_command: Optional[str] field is dropped from ToolApprovalDecision, the corresponding argument-rewrite logic is removed from ToolCallingLLM._execute_tool_decisions, existing edit_command tests are deleted, and new regression tests verify backward compatibility with older clients.

Changes

Remove edit_command from tool approval flow

Layer / File(s) Summary
Remove edit_command field and execution logic
holmes/core/models.py, holmes/core/tool_calling_llm.py
edit_command field removed from ToolApprovalDecision; the branch in _execute_tool_decisions that parsed JSON args, overwrote command, and persisted the change into conversation history is deleted. Approved decisions now proceed directly with the original assistant tool call arguments.
Delete old tests, add backward-compat regression tests
tests/test_tool_decision_edit_command.py, tests/core/conversations_worker/integration/test_conversation_integration.py, tests/test_edit_command_removed.py
test_tool_decision_edit_command.py (4 unit tests) and TestToolApproval.test_approval_with_edit_command integration test are removed. tests/test_edit_command_removed.py is added with two regression tests: one confirms ToolApprovalDecision.model_validate silently drops an incoming edit_command field, and another confirms the original assistant-provided command runs even when edit_command is included in the decision payload.

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~10 minutes

Possibly related PRs

  • HolmesGPT/holmesgpt#919: Introduced the ToolApprovalDecision model and the approval workflow that this PR modifies by removing edit_command.
  • HolmesGPT/holmesgpt#1990: Added the edit_command field and the edit-and-persist logic in _execute_tool_decisions that this PR removes.

Suggested reviewers

  • moshemorad
🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 40.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title references a specific task (ROB-320) and describes a sequential step in removing the edit_command feature, which matches the changeset's removal of edit_command functionality across models, tool execution, and tests.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions

github-actions Bot commented Jun 17, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker images ready for 9698e36f5 (built in 4m 27s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use these tags to pull the images for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:9698e36f5
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:9698e36f5 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:9698e36f5
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:9698e36f5
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:9698e36f5
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:9698e36f5 me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:9698e36f5
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:9698e36f5

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:9698e36f5 \
  --set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set operator.image=holmes-operator-dev:9698e36f5

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:9698e36f5 \
  --set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.operator.image=holmes-operator-dev:9698e36f5

@netlify

netlify Bot commented Jun 17, 2026

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit 551586c
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/6a32c867266a9a0008386da8
😎 Deploy Preview https://deploy-preview-2202--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
tests/test_edit_command_removed.py (1)

1-103: ⚡ Quick win

Move this regression test under the tests/core/... hierarchy.

This module validates behavior in holmes/core/models.py and holmes/core/tool_calling_llm.py, but it currently sits at tests/test_edit_command_removed.py. Please place it under the matching tests/core/... tree to keep test/source structure aligned.

As per coding guidelines: “Tests must match source structure under tests/.”

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/test_edit_command_removed.py` around lines 1 - 103, The test module
containing the test_edit_command_in_payload_is_silently_dropped_by_pydantic and
test_original_command_runs_even_when_edit_command_was_sent tests is currently
located at the root level of the tests directory. Since these tests validate
behavior in holmes.core.models and holmes.core.tool_calling_llm, move this
entire test file to align with the source code structure by placing it in the
core subdirectory of tests, matching the hierarchy of the source modules being
tested.

Source: Coding guidelines

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Nitpick comments:
In `@tests/test_edit_command_removed.py`:
- Around line 1-103: The test module containing the
test_edit_command_in_payload_is_silently_dropped_by_pydantic and
test_original_command_runs_even_when_edit_command_was_sent tests is currently
located at the root level of the tests directory. Since these tests validate
behavior in holmes.core.models and holmes.core.tool_calling_llm, move this
entire test file to align with the source code structure by placing it in the
core subdirectory of tests, matching the hierarchy of the source modules being
tested.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: e8fda671-d5eb-43c1-9dad-024eaed16bf2

📥 Commits

Reviewing files that changed from the base of the PR and between ca1a427 and 551586c.

📒 Files selected for processing (5)
  • holmes/core/models.py
  • holmes/core/tool_calling_llm.py
  • tests/core/conversations_worker/integration/test_conversation_integration.py
  • tests/test_edit_command_removed.py
  • tests/test_tool_decision_edit_command.py
💤 Files with no reviewable changes (4)
  • holmes/core/models.py
  • holmes/core/tool_calling_llm.py
  • tests/test_tool_decision_edit_command.py
  • tests/core/conversations_worker/integration/test_conversation_integration.py

@RoiGlinik
RoiGlinik merged commit 6f7dcdc into master Jun 18, 2026
19 of 22 checks passed
@RoiGlinik
RoiGlinik deleted the claude/remove-edit-command-step2 branch June 18, 2026 08:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants