Skip to content

Add behavior_controls to /api/chat for controlling prompt components - #1479

Merged
Sheeproid merged 6 commits into
masterfrom
prompt-control
Feb 9, 2026
Merged

Sheeproid merged 6 commits into
masterfrom
prompt-control

Conversation

@Sheeproid

@Sheeproid Sheeproid commented Feb 4, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Adds a behavior_controls field to the /api/chat endpoint that allows clients to override prompt components at request time. The primary use case is disabling the TodoWrite planning phase for faster responses.

Changes

  • holmes/core/models.py: Added behavior_controls: Optional[Dict[str, bool]] to ChatRequest
  • holmes/core/prompt.py:
    • Added is_component_enabled() function with precedence: env var > API override > default
    • Refactored build_system_prompt and build_user_prompt to use local closure pattern
  • holmes/core/conversations.py: Pass overrides through build_chat_messages
  • server.py: Convert string keys to PromptComponent enum with graceful handling of unknown keys
  • tests/core/test_prompt.py: Added tests for is_component_enabled

Usage

{
  "ask": "What's wrong with my pod?",
  "behavior_controls": {
    "todowrite_instructions": false,
    "todowrite_reminder": false
  }
}

Notes

- Keys are case-insensitive and map 1:1 to PromptComponent enum values
- Unknown keys are ignored with a warning log (forward compatible)
- Env var ENABLED_PROMPTS always takes precedence over API overrides


<!-- This is an auto-generated comment: release notes by coderabbit.ai -->
## Summary by CodeRabbit

* **New Features**
* Per-request controls to selectively enable/disable individual prompt components via a new request field; environment settings still take precedence and unknown keys are ignored with a warning.
* New fast-mode CLI flag that disables specific prompt components for quicker runs and surfaces those choices through the interactive and non-interactive ask flows; applied overrides are logged.

* **Tests**
* Added tests covering per-component enablement, override precedence, and env-var interactions.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Tomer Keshet <tomer@robusta.dev>
@linux-foundation-easycla

linux-foundation-easycla Bot commented Feb 4, 2026 •

Copy link
Copy Markdown

CLA Not Signed

@netlify

netlify Bot commented Feb 4, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit 1bf3202
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/6989afe1bfa6a60008896a5e
😎 Deploy Preview https://deploy-preview-1479--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@github-actions

github-actions Bot commented Feb 4, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

📜 Run @ ce5e5c6 (#21794250430)

✅ Results of HolmesGPT evals

Automatically triggered by commit ce5e5c6 on branch prompt-control

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 35.1s 6 11 $0.2367
✅ 101_loki_historical_logs_pod_deleted 51.5s 7 12 $0.2979
✅ 111_pod_names_contain_service 38.2s 6 12 $0.2400
✅ 112_find_pvcs_by_uuid 26.2s 4 5 $0.1957
✅ 12_job_crashing 43.8s 6 12 $0.2462
✅ 176_network_policy_blocking_traffic_no_runbooks 44.0s 6 16 $0.2794
✅ 24_misconfigured_pvc 33.8s 5 13 $0.2270
✅ 43_current_datetime_from_prompt 5.4s 1 — $0.1051
✅ 61_exact_match_counting 14.2s 3 2 $0.1447
Total 32.5s avg 4.9 avg 10.4 avg $1.9727
📜 Run @ 632deb9 (#21677940072)

✅ Results of HolmesGPT evals

Automatically triggered by commit 632deb9 on branch prompt-control

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 42.6s 6 11 $0.2370
✅ 101_loki_historical_logs_pod_deleted 53.7s 6 11 $0.2689
✅ 111_pod_names_contain_service 44.8s 5 11 $0.2189
✅ 112_find_pvcs_by_uuid 46.5s 6 6 $0.2377
✅ 12_job_crashing 40.4s 5 10 $0.2303
✅ 176_network_policy_blocking_traffic_no_runbooks 51.6s 6 16 $0.2774
✅ 24_misconfigured_pvc 43.9s 5 13 $0.2245
✅ 43_current_datetime_from_prompt 6.6s 1 — $0.1052
✅ 61_exact_match_counting 25.4s 4 4 $0.1617
Total 39.5s avg 4.9 avg 10.2 avg $1.9616
📜 Run @ 76da8af (#21677673666)

✅ Results of HolmesGPT evals

Automatically triggered by commit 76da8af on branch prompt-control

View workflow logs

⚠️ No eval report was generated.

📜 Run @ c4c095a (#21675502515)

✅ Results of HolmesGPT evals

Automatically triggered by commit c4c095a on branch prompt-control

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 31.8s 4 10 $0.2129
✅ 101_loki_historical_logs_pod_deleted 60.2s 7 13 $0.2893
✅ 111_pod_names_contain_service 40.7s 5 12 $0.2375
✅ 112_find_pvcs_by_uuid 40.8s 6 8 $0.2556
✅ 12_job_crashing 45.0s 6 14 $0.2654
✅ 176_network_policy_blocking_traffic_no_runbooks 53.6s 7 17 $0.2884
✅ 24_misconfigured_pvc 39.4s 5 16 $0.2404
✅ 43_current_datetime_from_prompt 5.6s 1 — $0.1049
✅ 61_exact_match_counting 17.8s 4 4 $0.1586
Total 37.2s avg 5.0 avg 11.8 avg $2.0531
📜 Run @ 501b08d (#21675429138)

✅ Results of HolmesGPT evals

Automatically triggered by commit 501b08d on branch prompt-control

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 33.2s 5 11 $0.2336
✅ 101_loki_historical_logs_pod_deleted 53.5s 7 12 $0.2881
✅ 111_pod_names_contain_service 34.5s 5 11 $0.2195
✅ 112_find_pvcs_by_uuid 31.6s 5 7 $0.2159
✅ 12_job_crashing 33.8s 5 10 $0.2222
✅ 176_network_policy_blocking_traffic_no_runbooks 41.6s 5 14 $0.2574
✅ 24_misconfigured_pvc 34.8s 5 13 $0.2317
✅ 43_current_datetime_from_prompt 5.9s 1 — $0.1060
✅ 61_exact_match_counting 19.4s 4 4 $0.1594
Total 32.0s avg 4.7 avg 10.2 avg $1.9338

✅ Results of HolmesGPT evals

Automatically triggered by commit 1bf3202 on branch prompt-control

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 38.0s 5 11 $0.2259
✅ 101_loki_historical_logs_pod_deleted 44.7s 5 9 $0.2269
✅ 111_pod_names_contain_service 36.4s 5 11 $0.2176
✅ 112_find_pvcs_by_uuid 49.9s 7 6 $0.3478
✅ 12_job_crashing 29.8s 4 7 $0.2028
✅ 176_network_policy_blocking_traffic_no_runbooks 46.9s 6 16 $0.2813
✅ 24_misconfigured_pvc 51.7s 6 14 $0.2422
✅ 43_current_datetime_from_prompt 6.8s 1 — $0.1048
✅ 61_exact_match_counting 20.2s 4 4 $0.1586
Total 36.0s avg 4.8 avg 9.8 avg $2.0079
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref prompt-control -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, easy, elasticsearch, embeds, fast, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref prompt-control -f markers=regression -f filter=

@github-actions

github-actions Bot commented Feb 4, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker image ready for d545924 (built in 3m 51s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use this tag to pull the image for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:d545924
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:d545924 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:d545924
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:d545924

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:d545924

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:d545924

@coderabbitai

coderabbitai Bot commented Feb 4, 2026 •

Copy link
Copy Markdown
Contributor

Walkthrough

Adds per-component prompt overrides: a new behavior_controls field on ChatRequest is mapped to prompt_component_overrides in the server and propagated through build_chat_messages → build_prompts → system/user prompt builders to gate individual PromptComponents.

Changes

Cohort / File(s) Summary
Model / Conversation API
holmes/core/models.py, holmes/core/conversations.py
Added behavior_controls: Optional[Dict[str, bool]] to ChatRequest. build_chat_messages(...) now accepts prompt_component_overrides: Optional[Dict[PromptComponent, bool]] and forwards it to prompt construction.
Prompt logic
holmes/core/prompt.py
Renamed env-only check to is_prompt_allowed_by_env, added is_component_enabled(component, overrides), and threaded prompt_component_overrides through build_system_prompt, build_user_prompt, build_prompts, and build_initial_ask_messages to gate prompt components.
Server integration
server.py
Maps chat_request.behavior_controls string keys to PromptComponent values, warns on unknown keys, logs application, and passes prompt_component_overrides into build_chat_messages.
CLI / interactive wiring
holmes/main.py, holmes/interactive.py
Added fast_mode CLI flag that sets specific component overrides; threaded prompt_component_overrides through run_interactive_loop, build_initial_ask_messages, and non-interactive message construction.
Tests / Exports
tests/core/test_prompt.py, holmes/core/prompt.py
Exported PromptComponent and is_component_enabled; added TestIsComponentEnabled tests validating env vs. override precedence and related behavior.

Sequence Diagram(s)

sequenceDiagram
  participant Client as Client (ChatRequest)
  participant Server as Server (server.py)
  participant Conv as Conversations.build_chat_messages
  participant Prompt as Prompt.build_prompts
  participant Builders as Prompt.system/user builders
  participant LLM as LLM

  Client->>Server: POST /chat with behavior_controls
  Server->>Server: map string keys → PromptComponent (warn on unknown)
  Server->>Conv: build_chat_messages(..., prompt_component_overrides)
  Conv->>Prompt: build_prompts(prompt_component_overrides)
  Prompt->>Builders: is_component_enabled(component, overrides)
  Builders-->>Prompt: system/user prompt fragments
  Prompt-->>Conv: composed prompts
  Conv->>LLM: send final messages
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

Suggested reviewers

  • moshemorad
  • mainred
🚥 Pre-merge checks | ✅ 2 | ❌ 1
❌ Failed checks (1 warning)
Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 77.27% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The pull request title accurately describes the main change: adding a behavior_controls field to the /api/chat endpoint for controlling prompt components, which aligns with the primary objective and the majority of code changes.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


No actionable comments were generated in the recent review. 🎉

🧹 Recent nitpick comments
holmes/main.py (1)

711-711: Consider passing None instead of empty dict for consistency.

build_system_prompt accepts Dict[PromptComponent, bool] (non-optional), so {} works correctly here. However, other call sites (e.g., build_initial_ask_messages in the ask command) pass None when no overrides are needed. If the signature can be relaxed to Optional, using None would be more consistent.

This is a minor point — no functional impact.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions

github-actions Bot commented Feb 4, 2026 •

Copy link
Copy Markdown
Contributor

🔬 CLI Performance Benchmark

🟡 Startup Time (no LLM)

Measures holmes version execution time (imports + initialization)

Metric PR Master Change
Cold Start 11.34s 9.83s +15.3%
Warm Mean 5.30s 4.64s +14.2%
Warm Min 5.27s 4.55s
Warm Max 5.33s 4.73s

🟡 Full CLI with LLM

Measures holmes ask execution time (OpenRouter + Haiku 4.5)

Metric PR Master Change
Cold Start 29.95s 31.49s -4.9%
Warm Mean 8.34s 7.40s +12.7%
Warm Min 8.14s 7.18s
Warm Max 8.85s 7.79s

PR: d5459243 | Master: 48623c7a | Iterations: 5

@Sheeproid
Sheeproid requested a review from moshemorad February 4, 2026 14:37

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 0

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
server.py (1)

362-381: ⚠️ Potential issue | 🟠 Major

Add unit coverage for behavior_controls mapping and unknown-key handling.
This is a new feature path (case-insensitive mapping, ignored unknown keys), but there’s no unit test proving it works end-to-end through /api/chat.

As per coding guidelines, All new features require unit tests; new toolsets require integration tests; complex investigations require LLM evaluation tests.

🧹 Nitpick comments (1)
tests/core/test_prompt.py (1)

520-555: Add type hints to new test methods for mypy compliance.
Annotate monkeypatch and return types (e.g., -> None) in this class to satisfy the project’s type-checking requirement.

♻️ Example update (apply similarly to the rest of this class)
-    def test_no_overrides_returns_env_var_result(self, monkeypatch):
+    def test_no_overrides_returns_env_var_result(
+        self, monkeypatch: pytest.MonkeyPatch
+    ) -> None:
As per coding guidelines, Use mypy for type checking with type hints required in all code.

Signed-off-by: Tomer Keshet <tomer@robusta.dev>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 0

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
server.py (1)

360-382: ⚠️ Potential issue | 🟠 Major

Add unit tests for behavior_controls mapping behavior.

This is a new API feature; please cover case‑insensitive key mapping and unknown‑key ignore behavior to lock the contract. As per coding guidelines: tests/**/*.py: All new features require unit tests; new toolsets require integration tests; complex investigations require LLM evaluation tests.

@Sheeproid
Sheeproid enabled auto-merge (squash) February 8, 2026 07:24
@Sheeproid
Sheeproid requested a review from RoiGlinik February 9, 2026 08:56
@Sheeproid
Sheeproid merged commit 118cb54 into master Feb 9, 2026
18 of 20 checks passed
@Sheeproid
Sheeproid deleted the prompt-control branch February 9, 2026 10:03
moshemorad pushed a commit that referenced this pull request Feb 22, 2026
…1479)

## Summary
Adds a `behavior_controls` field to the `/api/chat` endpoint that allows
clients to override prompt components at request time. The primary use
case is disabling the TodoWrite planning phase for faster responses.
  ## Changes

- **`holmes/core/models.py`**: Added `behavior_controls:
Optional[Dict[str, bool]]` to `ChatRequest`
  - **`holmes/core/prompt.py`**:
- Added `is_component_enabled()` function with precedence: env var > API
override > default
- Refactored `build_system_prompt` and `build_user_prompt` to use local
closure pattern
- **`holmes/core/conversations.py`**: Pass overrides through
`build_chat_messages`
- **`server.py`**: Convert string keys to `PromptComponent` enum with
graceful handling of unknown keys
- **`tests/core/test_prompt.py`**: Added tests for
`is_component_enabled`

  ## Usage

  ```json
  {
    "ask": "What's wrong with my pod?",
    "behavior_controls": {
      "todowrite_instructions": false,
      "todowrite_reminder": false
    }
  }

  Notes

  - Keys are case-insensitive and map 1:1 to PromptComponent enum values
  - Unknown keys are ignored with a warning log (forward compatible)
  - Env var ENABLED_PROMPTS always takes precedence over API overrides

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Per-request controls to selectively enable/disable individual prompt
components via a new request field; environment settings still take
precedence and unknown keys are ignored with a warning.
* New fast-mode CLI flag that disables specific prompt components for
quicker runs and surfaces those choices through the interactive and
non-interactive ask flows; applied overrides are logged.

* **Tests**
* Added tests covering per-component enablement, override precedence,
and env-var interactions.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Tomer Keshet <tomer@robusta.dev>
Signed-off-by: Mohse Morad <moshemorad12340@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants