Conversation
Wrap the detailed results table in an HTML <details> tag so the PR comment is more concise by default. The summary stats remain visible at the top, and users can expand "Detailed Results" to see the full table with per-test metrics. All information is preserved - only the presentation changes. Signed-off-by: Claude <noreply@anthropic.com>
Add a prominent status banner at the top of the PR comment: - ✅ All X/Y tests passed (when no regressions) - ❌ N regression(s) — X/Y tests passed (when regressions exist) This makes it immediately clear whether the eval passed or failed without needing to read the detailed breakdown. Signed-off-by: Claude <noreply@anthropic.com>
- Make header smaller (#### instead of ##) - Remove per-test-type breakdown (ask_holmes: X/Y, etc.) - Move historical comparison footer inside collapsible details - Keep only the clear pass/fail status banner visible Signed-off-by: Claude <noreply@anthropic.com>
…ails - Remove "Results of HolmesGPT evals" header (added by workflow) - Simplify status to just "✅ **All X/Y tests passed**" - Rename collapsible section to just "Details" - Move historical comparison info inside the single Details section - Remove redundant separate Historical Comparison Details section Signed-off-by: Claude <noreply@anthropic.com>
|
|
✅ Deploy Preview for holmes-docs ready!
To edit notification comments on pull requests, go to your Netlify project configuration. |
Results of HolmesGPT evalsAutomatically triggered by commit 64428a6 on branch 🔄 Re-run evals manually
Option 1: Comment on this PR with Or with more options (one per line): Run evals on a different branch (e.g., master) for comparison:
Quick re-run: Use Option 2: Trigger via GitHub Actions UI → "Run workflow" 🏷️ Valid markers
Commands: CLI: |
|
✅ Docker image ready for
Use this tag to pull the image for testing. 📋 Copy commandsgcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:e6d64cd
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:e6d64cd me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:e6d64cd
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:e6d64cdPatch Helm values in one line (choose the chart you use): HolmesGPT chart: helm upgrade --install holmesgpt ./helm/holmes \
--set registry=me-west1-docker.pkg.dev/robusta-development/development \
--set image=holmes-dev:e6d64cdRobusta wrapper chart: helm upgrade --install robusta robusta/robusta \
--reuse-values \
--set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.image=holmes-dev:e6d64cd |
WalkthroughRenamed and simplified historical comparison output in the test reporter; removed history-preservation helpers from eval comment scripting and inlined comment body construction in workflows; consolidated running/results comment flows and adjusted exported helper surface. Changes
Sequence Diagram(s)mermaid Estimated code review effort🎯 4 (Complex) | ⏱️ ~45 minutes Possibly related PRs
Suggested reviewers
🚥 Pre-merge checks | ✅ 3✅ Passed checks (3 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Fix all issues with AI agents
In `@tests/llm/utils/reporting/github_reporter.py`:
- Around line 200-213: The banner currently shows "All passed" when
total_regressions == 0 even if there are skipped/setup/mock failures; change the
gating in the overall status logic (the block using total_tests, total_passed,
total_regressions) to require total_passed == total_tests for the green "All X/X
tests passed" message, otherwise render a non-passing banner and include counts
of other non-passing categories (e.g., skipped_count, setup_failures,
mock_failures or their existing variables) in the summary string so the message
accurately reflects skipped/setup/mock failures as well as regressions.
Simplify PR eval comment to just two states: - Running: Simple "evals running..." message - Results: Final eval results with collapsible details Removed: - Previous Runs section (history in comment edit history) - Progress checklist updates (spammed notification history) - All intermediate comment updates during eval run The comment now only updates twice: when starting and when finished. This reduces notification spam while preserving results in the GitHub comment edit history. Signed-off-by: Claude <noreply@anthropic.com>
Signed-off-by: Claude <noreply@anthropic.com>
Signed-off-by: Claude <noreply@anthropic.com>
The buildRunningBody and buildResultsBody functions were dead code because: - Workflow loads helpers from .trusted/ (master checkout) for security - New functions wouldn't exist in master until merged - So the logic is inlined directly in the workflow YAML This removes the dead code and keeps only the helpers that are actually used. Signed-off-by: Claude <noreply@anthropic.com>
- Make "Details" summary bold to match other sections - Remove redundant footer (commands list and CLI) - Add CLI as Option 3 inside Re-run section - Add explicit /list mention for completeness Signed-off-by: Claude <noreply@anthropic.com>
- Remove emojis from section headers (Legend, Re-run, Valid markers) - Simplify warning block (remove duplicate CLI command) - Consolidate /rerun and /list into single "Other commands" line - Make historical comparison its own collapsible section - Shorten verbose descriptions Signed-off-by: Claude <noreply@anthropic.com>
|
/eval |
|
@aantn Your eval run has finished. ✅ Completed successfully 🧪 Manual Eval Results
✅ All 9/9 tests passed Details
Historical comparisonTime/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%. Compared against: github-21534141772.2746.2-425eaa10, github-21534141772.2746.2, github-21540289939.2760.1-9233df2f, +27 more 📖 Legend
🔄 Re-run evals manually
Option 1: Comment on this PR with Or with more options (one per line): Run evals on a different branch (e.g., master) for comparison:
Quick re-run: Use Option 2: Trigger via GitHub Actions UI → "Run workflow" 🏷️ Valid markers
Commands: CLI: |
Signed-off-by: Claude <noreply@anthropic.com>
- Move legend inline into Details section (compact single line) - Remove separate Legend collapsible section from helpers - Remove includeLegend option from buildRerunFooter - Remove ✅ emoji from "Results of HolmesGPT evals" header Signed-off-by: Claude <noreply@anthropic.com>
Conflict in github_reporter.py: - Master re-added per-test-type summary lines (ask_holmes: X/Y, etc.) - Kept our simplified version (just status banner, no per-type breakdown) Signed-off-by: Claude <noreply@anthropic.com>
🔬 CLI Performance Benchmark🟡 Startup Time (no LLM)Measures
🟡 Full CLI with LLMMeasures
PR: |
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Fix all issues with AI agents
In `@tests/llm/utils/reporting/github_reporter.py`:
- Around line 181-184: generate_markdown_report currently sums undefined
variables workload_health_total, workload_health_passed, and
workload_health_regressions which causes a NameError; either remove those
workload_health references from the totals calculation and compute
total_tests/total_passed/total_regressions from the existing counters
(ask_holmes_total, investigate_total, ask_holmes_passed, investigate_passed,
ask_holmes_regressions, investigate_regressions) or, if workload_health metrics
were intended, add explicit initialization and increment logic for
workload_health_total, workload_health_passed, and workload_health_regressions
in the same counting loop that handles 'ask' and 'investigate' inside
generate_markdown_report so the variables exist before being summed.
| # Calculate totals for overall status | ||
| total_tests = ask_holmes_total + investigate_total + workload_health_total | ||
| total_passed = ask_holmes_passed + investigate_passed + workload_health_passed | ||
| total_regressions = ask_holmes_regressions + investigate_regressions + workload_health_regressions |
There was a problem hiding this comment.
Critical: Undefined variables cause NameError at runtime.
Lines 182-184 reference workload_health_total, workload_health_passed, and workload_health_regressions, but these variables are never defined. The counting loop (lines 153-179) only handles ask and investigate test types—there's no workload_health counter initialization or increment logic.
This will crash generate_markdown_report() with a NameError whenever it's called.
🐛 Proposed fix: Remove undefined workload_health references
# Calculate totals for overall status
- total_tests = ask_holmes_total + investigate_total + workload_health_total
- total_passed = ask_holmes_passed + investigate_passed + workload_health_passed
- total_regressions = ask_holmes_regressions + investigate_regressions + workload_health_regressions
+ total_tests = ask_holmes_total + investigate_total
+ total_passed = ask_holmes_passed + investigate_passed
+ total_regressions = ask_holmes_regressions + investigate_regressions🧰 Tools
🪛 Ruff (0.14.14)
[error] 182-182: Undefined name workload_health_total
(F821)
[error] 183-183: Undefined name workload_health_passed
(F821)
[error] 184-184: Undefined name workload_health_regressions
(F821)
🤖 Prompt for AI Agents
In `@tests/llm/utils/reporting/github_reporter.py` around lines 181 - 184,
generate_markdown_report currently sums undefined variables
workload_health_total, workload_health_passed, and workload_health_regressions
which causes a NameError; either remove those workload_health references from
the totals calculation and compute total_tests/total_passed/total_regressions
from the existing counters (ask_holmes_total, investigate_total,
ask_holmes_passed, investigate_passed, ask_holmes_regressions,
investigate_regressions) or, if workload_health metrics were intended, add
explicit initialization and increment logic for workload_health_total,
workload_health_passed, and workload_health_regressions in the same counting
loop that handles 'ask' and 'investigate' inside generate_markdown_report so the
variables exist before being summed.
Summary by CodeRabbit
Improvements
Chores