Add enable_todo flag to control TodoWrite feature in evals - #2127
Conversation
Disable the TodoWrite/todos feature in eval runs by default. Both the TodoWrite tool (core_investigation toolset) and the related prompt instructions/reminder are turned off so they don't influence eval behavior or token usage. Tests that specifically need todos can opt back in by setting 'enable_todo: true' in their test_case.yaml. Signed-off-by: Claude <noreply@anthropic.com>
There was a problem hiding this comment.
Claude Code Review
This repository is configured for manual code reviews. Comment @claude review to trigger a review and subscribe this PR to future pushes, or @claude review once for a one-time review.
Tip: disable this comment in your organization's Code Review settings.
📂 Previous Runs📜 #1 · Run @ __01c438a__ (#26970039513) — Jun 4, 18:08 UTC✅ Results of HolmesGPT evalsAutomatically triggered by commit 01c438a on branch Results of HolmesGPT evals
Benchmark Comparison DetailsMaster baseline: latest master-* experiment (post-merge regression eval)
Benchmark baseline: latest ci-benchmark experiment on master
Time comparison (seconds):
Cost comparison:
Total tokens comparison:
Cached tokens comparison:
Turns comparison:
Tool calls comparison:
Comparison indicators:
✅ Results of HolmesGPT evalsAutomatically triggered by commit 4fbf2dd on branch Results of HolmesGPT evals
Benchmark Comparison DetailsMaster baseline: latest master-* experiment (post-merge regression eval)
Benchmark baseline: latest ci-benchmark experiment on master
Time comparison (seconds):
Cost comparison:
Total tokens comparison:
Cached tokens comparison:
Turns comparison:
Tool calls comparison:
Comparison indicators:
📖 Legend
🔄 Re-run evals manually
Option 1: Comment on this PR with Or with more options (one per line): Run evals on a different branch (e.g., master) for comparison:
Quick re-run: Use Option 2: Trigger via GitHub Actions UI → "Run workflow" Option 3: Add PR labels to include extra evals (applies to both automatic runs and
Examples: 🏷️ Valid tags
🤖 Valid models
Commands: CLI: |
|
✅ Docker images ready for
Use these tags to pull the images for testing. 📋 Copy commandsgcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:a0ebf462a
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:a0ebf462a me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:a0ebf462a
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:a0ebf462a
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:a0ebf462a
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:a0ebf462a me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:a0ebf462a
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:a0ebf462aPatch Helm values in one line (choose the chart you use): HolmesGPT chart: helm upgrade --install holmesgpt ./helm/holmes \
--set registry=me-west1-docker.pkg.dev/robusta-development/development \
--set image=holmes-dev:a0ebf462a \
--set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set operator.image=holmes-operator-dev:a0ebf462aRobusta wrapper chart: helm upgrade --install robusta robusta/robusta \
--reuse-values \
--set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.image=holmes-dev:a0ebf462a \
--set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.operator.image=holmes-operator-dev:a0ebf462a |
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Pro Run ID: 📒 Files selected for processing (3)
WalkthroughThis PR extends the test infrastructure to conditionally gate the TodoWrite feature. A new ChangesEnable-todo flag propagation
Estimated code review effort🎯 2 (Simple) | ⏱️ ~10 minutes Possibly related PRs
Suggested reviewers
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
✅ Deploy Preview for holmes-docs ready!
To edit notification comments on pull requests, go to your Netlify project configuration. |
Summary
Add a new
enable_todoconfiguration flag to allow tests to opt-in to the TodoWrite/todos feature, which is disabled by default in evaluations. This prevents the TodoWrite tool and related prompt instructions from being offered to the LLM unless explicitly enabled by a test case.Changes
test_case_utils.py): Addedenable_todo: bool = Falsefield toHolmesTestCaseto allow individual tests to enable the featuretest_toolset.py):enable_todoparameter toTestToolsetManager.__init__core_investigationtoolset entirely whenenable_todo=Falseto prevent the TodoWrite tool from being offered to the LLMtest_ask_holmes.py):PromptComponentfrom prompt moduleenable_todoflag from test case toTestToolsetManagerprompt_component_overridesdict to disable TodoWrite-related prompt instructions/reminders whenenable_todo=Falseprompt_component_overridesto both CLI and API investigation pathsImplementation Details
The TodoWrite feature is now controlled at three levels:
core_investigationtoolset (which contains TodoWrite) is dropped from the toolset manager unless opted inTODOWRITE_INSTRUCTIONSandTODOWRITE_REMINDER) are disabled viaprompt_component_overridesunless opted inenable_todo: Truein their test case definitionThis ensures the LLM won't be offered or reminded about the TodoWrite tool in evaluations unless a specific test explicitly enables it.
https://claude.ai/code/session_01HeCcwBCximu5VRbgno8gqr
Summary by CodeRabbit
enable_todoconfiguration option to control Todo-related functionality