Conversation
PR #1970 inlined investigation_procedure.jinja2 into generic_ask.jinja2 and, in doing so, repositioned the `# MANDATORY Task Management` block from its old slot (between `# Special cases and how to reply` and `# Tool/function calls`, near the end of the prompt) to immediately after `# INVESTIGATION PHASE TRANSITION EXAMPLES`, in the middle of the prompt. The block content is byte-identical; only its position changed. But that single positional move turns out to drive the residual ~10% eval-perf regression that survived PR #2051's revert of PR #2040. Local n=10 sweep against the docker-loki regression eval (#259) on opus-4.6: condition time turns total_tk compl_tk baseline (cf6ddb7) 78.6s 7.8 193,919 5,097 PR-merged (fix-AD only) 83.0s 8.6 202,756 4,834 <- residual Restore deleted intro lines 76.0s 9.0 208,022 4,647 THIS PATCH (just block move) 71.1s 7.6 182,068 4,421 Block move + intro lines 76.1s 7.6 185,765 4,829 Full baseline prompt 75.1s 7.2 182,008 4,914 Every z vs baseline for this patch is ≤ 0: time z=-1.91, turns z=-0.46, completion z=-2.14, total tokens z=-0.92. The block-move alone is strictly better than restoring the two deleted intro lines, and matches the full baseline-prompt restoration on aggregate metrics. No text added or removed; just position. Hypothesis on the mechanism: with the task-management rules sandwiched inside the multi-phase investigation doctrine, the model reads "If you discover additional steps during investigation, add them to your task list using TodoWrite" while it is still planning phases and acts on it more aggressively (extra TodoWrite mid-investigation, extra turns). When the same rules sit at the end of the prompt as a quiet reminder (post-Special-Cases, pre-Tool/function-calls), they don't compound with the phase-planning instructions. Signed-off-by: Claude <noreply@anthropic.com>
There was a problem hiding this comment.
Claude Code Review
This repository is configured for manual code reviews. Comment @claude review to trigger a review and subscribe this PR to future pushes, or @claude review once for a one-time review.
Tip: disable this comment in your organization's Code Review settings.
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Pro Run ID: 📒 Files selected for processing (1)
WalkthroughThis PR relocates the "MANDATORY Task Management" instruction subsection within a Jinja2 prompt template. The block is removed from its original position and re-added later under a new conditional, preserving the complete instruction set for TodoWrite task management. ChangesTask Management Instructions Relocation
Estimated code review effort🎯 1 (Trivial) | ⏱️ ~5 minutes 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
✅ Results of HolmesGPT evalsAutomatically triggered by commit f434e61 on branch Results of HolmesGPT evals
Benchmark Comparison DetailsMaster baseline: latest master-* experiment (post-merge regression eval)
Benchmark baseline: latest ci-benchmark experiment on master
Time comparison (seconds):
Cost comparison:
Total tokens comparison:
Cached tokens comparison:
Turns comparison:
Tool calls comparison:
Comparison indicators:
📖 Legend
🔄 Re-run evals manually
Option 1: Comment on this PR with Or with more options (one per line): Run evals on a different branch (e.g., master) for comparison:
Quick re-run: Use Option 2: Trigger via GitHub Actions UI → "Run workflow" Option 3: Add PR labels to include extra evals (applies to both automatic runs and
Examples: 🏷️ Valid tags
🤖 Valid models
Commands: CLI: |
|
✅ Docker images ready for
Use these tags to pull the images for testing. 📋 Copy commandsgcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:8169c6e4
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:8169c6e4 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:8169c6e4
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:8169c6e4
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:8169c6e4
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:8169c6e4 me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:8169c6e4
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:8169c6e4Patch Helm values in one line (choose the chart you use): HolmesGPT chart: helm upgrade --install holmesgpt ./helm/holmes \
--set registry=me-west1-docker.pkg.dev/robusta-development/development \
--set image=holmes-dev:8169c6e4 \
--set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set operator.image=holmes-operator-dev:8169c6e4Robusta wrapper chart: helm upgrade --install robusta robusta/robusta \
--reuse-values \
--set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.image=holmes-dev:8169c6e4 \
--set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.operator.image=holmes-operator-dev:8169c6e4 |
✅ Deploy Preview for holmes-docs ready!
To edit notification comments on pull requests, go to your Netlify project configuration. |
PR #1970 inlined investigation_procedure.jinja2 into generic_ask.jinja2 and, in doing so, repositioned the
# MANDATORY Task Managementblock from its old slot (between# Special cases and how to replyand# Tool/function calls, near the end of the prompt) to immediately after# INVESTIGATION PHASE TRANSITION EXAMPLES, in the middle of the prompt.The block content is byte-identical; only its position changed. But that single positional move turns out to drive the residual ~10% eval-perf regression that survived PR #2051's revert of PR #2040.
Local n=10 sweep against the docker-loki regression eval (#259) on opus-4.6:
condition time turns total_tk compl_tk
baseline (cf6ddb7) 78.6s 7.8 193,919 5,097
PR-merged (fix-AD only) 83.0s 8.6 202,756 4,834 <- residual
Restore deleted intro lines 76.0s 9.0 208,022 4,647
THIS PATCH (just block move) 71.1s 7.6 182,068 4,421
Block move + intro lines 76.1s 7.6 185,765 4,829
Full baseline prompt 75.1s 7.2 182,008 4,914
Every z vs baseline for this patch is ≤ 0: time z=-1.91, turns z=-0.46, completion z=-2.14, total tokens z=-0.92. The block-move alone is strictly better than restoring the two deleted intro lines, and matches the full baseline-prompt restoration on aggregate metrics. No text added or removed; just position.
Hypothesis on the mechanism: with the task-management rules sandwiched inside the multi-phase investigation doctrine, the model reads "If you discover additional steps during investigation, add them to your task list using TodoWrite" while it is still planning phases and acts on it more aggressively (extra TodoWrite mid-investigation, extra turns). When the same rules sit at the end of the prompt as a quiet reminder (post-Special-Cases, pre-Tool/function-calls), they don't compound with the phase-planning instructions.
Summary by CodeRabbit