Skip to content

More concise eval report - #1421

Closed
aantn wants to merge 15 commits into
masterfrom
claude/concise-regression-output-TjM2d
Closed

aantn wants to merge 15 commits into
masterfrom
claude/concise-regression-output-TjM2d

Conversation

@aantn

@aantn aantn commented Jan 25, 2026 •

Copy link
Copy Markdown
Collaborator

Summary by CodeRabbit

  • Improvements

    • Unified report layout with a single overall status banner and consolidated “Details” section for results.
    • Historical comparison notes streamlined and embedded inside the details view.
    • Comment content simplified: running and final-result comments display current run info and parameters clearly, with clearer manual vs automated headers.
  • Chores

    • Removed legacy in-comment history blocks and related history-preservation/progress-tracking behavior.

Wrap the detailed results table in an HTML <details> tag so the PR
comment is more concise by default. The summary stats remain visible
at the top, and users can expand "Detailed Results" to see the full
table with per-test metrics.

All information is preserved - only the presentation changes.

Signed-off-by: Claude <noreply@anthropic.com>
Add a prominent status banner at the top of the PR comment:
- ✅ All X/Y tests passed (when no regressions)
- ❌ N regression(s) — X/Y tests passed (when regressions exist)

This makes it immediately clear whether the eval passed or failed
without needing to read the detailed breakdown.

Signed-off-by: Claude <noreply@anthropic.com>
- Make header smaller (#### instead of ##)
- Remove per-test-type breakdown (ask_holmes: X/Y, etc.)
- Move historical comparison footer inside collapsible details
- Keep only the clear pass/fail status banner visible

Signed-off-by: Claude <noreply@anthropic.com>
…ails

- Remove "Results of HolmesGPT evals" header (added by workflow)
- Simplify status to just "✅ **All X/Y tests passed**"
- Rename collapsible section to just "Details"
- Move historical comparison info inside the single Details section
- Remove redundant separate Historical Comparison Details section

Signed-off-by: Claude <noreply@anthropic.com>
@linux-foundation-easycla

linux-foundation-easycla Bot commented Jan 25, 2026 •

Copy link
Copy Markdown

CLA Signed

The committers listed above are authorized under a signed CLA.

  • ✅ login: aantn / name: Natan Yellin (64428a6)

@netlify

netlify Bot commented Jan 25, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit 64428a6
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/6980b531f5cf920008381180
😎 Deploy Preview https://deploy-preview-1421--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@github-actions

github-actions Bot commented Jan 25, 2026 •

Copy link
Copy Markdown
Contributor

Results of HolmesGPT evals

Automatically triggered by commit 64428a6 on branch claude/concise-regression-output-TjM2d. View workflow logs

⚠️ No eval report was generated.

🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/concise-regression-output-TjM2d -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, easy, elasticsearch, embeds, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/concise-regression-output-TjM2d -f markers=regression -f filter=

@github-actions

github-actions Bot commented Jan 25, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker image ready for e6d64cd (built in 3m 32s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use this tag to pull the image for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:e6d64cd
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:e6d64cd me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:e6d64cd
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:e6d64cd

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:e6d64cd

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:e6d64cd

@coderabbitai

coderabbitai Bot commented Jan 25, 2026 •

Copy link
Copy Markdown
Contributor

Walkthrough

Renamed and simplified historical comparison output in the test reporter; removed history-preservation helpers from eval comment scripting and inlined comment body construction in workflows; consolidated running/results comment flows and adjusted exported helper surface.

Changes

Cohort / File(s) Summary
GitHub report generation
tests/llm/utils/reporting/github_reporter.py
Renamed _generate_historical_details_section() → _generate_historical_comparison_info() and changed its return to a concise info string. Rewrote generate_markdown_report() to remove branch headers, compute consolidated totals, emit an overall banner, wrap detailed results in one <details> block, and inline historical comparison content when present.
Eval comment helpers (script)
.github/scripts/eval-comment-helpers.js
Removed history-related constants and functions (HISTORY_RUN_END_MARKER, MAX_COMMENT_SIZE, parseRunHistory, extractCurrentRun, buildAutoCommentWithHistory, renderProgress, buildBody). Reduced exported surface to AUTO_EVAL_COMMENT_IDENTIFIER, buildParams, renderParamsTable, and buildRerunFooter; simplified buildRerunFooter signature and footer content.
Workflow usage
.github/workflows/eval-regression.yaml
Replaced history-aware comment assembly with inline body construction using AUTO_EVAL_COMMENT_IDENTIFIER, renderParamsTable, buildParams, and buildRerunFooter. Removed history-parsing branches and updated create/update comment flows to use the new inline bodies.

Sequence Diagram(s)

mermaid
sequenceDiagram
participant Workflow as Workflow
participant Helpers as eval-comment-helpers.js
participant Reporter as github_reporter.py
participant GitHub as GitHub API

Workflow->>Helpers: renderParamsTable(params)
Helpers-->>Workflow: params table (running body)
Workflow->>GitHub: create/update comment (running)
Workflow->>Reporter: generate_markdown_report(sorted_results, include_historical)
Reporter-->>Workflow: report content (table + historical comparison info)
Workflow->>Helpers: assemble final body (AUTO_EVAL_COMMENT_IDENTIFIER + params + report + buildRerunFooter)
Helpers-->>Workflow: final results body
Workflow->>GitHub: create/update comment (final results)

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly related PRs

Suggested reviewers

  • Sheeproid
  • arikalon1
🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately summarizes the primary objective of the changeset: making the evaluation report output more concise by consolidating sections and removing redundant elements.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Fix all issues with AI agents
In `@tests/llm/utils/reporting/github_reporter.py`:
- Around line 200-213: The banner currently shows "All passed" when
total_regressions == 0 even if there are skipped/setup/mock failures; change the
gating in the overall status logic (the block using total_tests, total_passed,
total_regressions) to require total_passed == total_tests for the green "All X/X
tests passed" message, otherwise render a non-passing banner and include counts
of other non-passing categories (e.g., skipped_count, setup_failures,
mock_failures or their existing variables) in the summary string so the message
accurately reflects skipped/setup/mock failures as well as regressions.

Comment thread tests/llm/utils/reporting/github_reporter.py
Simplify PR eval comment to just two states:
- Running: Simple "evals running..." message
- Results: Final eval results with collapsible details

Removed:
- Previous Runs section (history in comment edit history)
- Progress checklist updates (spammed notification history)
- All intermediate comment updates during eval run

The comment now only updates twice: when starting and when finished.
This reduces notification spam while preserving results in the
GitHub comment edit history.

Signed-off-by: Claude <noreply@anthropic.com>
Signed-off-by: Claude <noreply@anthropic.com>
Signed-off-by: Claude <noreply@anthropic.com>
The buildRunningBody and buildResultsBody functions were dead code because:
- Workflow loads helpers from .trusted/ (master checkout) for security
- New functions wouldn't exist in master until merged
- So the logic is inlined directly in the workflow YAML

This removes the dead code and keeps only the helpers that are actually used.

Signed-off-by: Claude <noreply@anthropic.com>
- Make "Details" summary bold to match other sections
- Remove redundant footer (commands list and CLI)
- Add CLI as Option 3 inside Re-run section
- Add explicit /list mention for completeness

Signed-off-by: Claude <noreply@anthropic.com>
- Remove emojis from section headers (Legend, Re-run, Valid markers)
- Simplify warning block (remove duplicate CLI command)
- Consolidate /rerun and /list into single "Other commands" line
- Make historical comparison its own collapsible section
- Shorten verbose descriptions

Signed-off-by: Claude <noreply@anthropic.com>
@aantn

aantn commented Jan 31, 2026

Copy link
Copy Markdown
Collaborator Author

/eval
markers: regression

@github-actions

Copy link
Copy Markdown
Contributor

@aantn Your eval run has finished. ✅ Completed successfully


🧪 Manual Eval Results

Parameter Value
Triggered via /eval comment
Branch claude/concise-regression-output-TjM2d
Model opus-4.5
Markers regression
Iterations 1
Duration 5m 12s
Workflow View logs | Rerun

✅ All 9/9 tests passed

Details
Status Test case Time Turns Tools Cost
✅ 09_crashpod 34.6s ±0% 6 11 $0.2423
✅ 101_loki_historical_logs_pod_deleted 55.4s ↑28% 8 14 $0.3143
✅ 111_pod_names_contain_service 32.2s ±0% 5 11 $0.2244
✅ 12_job_crashing 31.8s ±0% 5 10 $0.2309
✅ 162_get_runbooks 40.8s ±0% 6 11 $0.2719
✅ 176_network_policy_blocking_traffic_no_runbooks 45.1s ±0% 6 17 $0.2762
✅ 24_misconfigured_pvc 35.6s ±0% 6 14 $0.2447
✅ 43_current_datetime_from_prompt 4.9s ±0% 1 — $0.1049
✅ 61_exact_match_counting 13.3s ↓11% 3 2 $0.1425
Total 32.6s avg 5.1 avg 11.2 avg $2.0522
Historical comparison

Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%.

Compared against: github-21534141772.2746.2-425eaa10, github-21534141772.2746.2, github-21540289939.2760.1-9233df2f, +27 more

📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/concise-regression-output-TjM2d -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, easy, elasticsearch, embeds, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/concise-regression-output-TjM2d -f markers=regression -f filter=

claude and others added 4 commits January 31, 2026 07:17
Signed-off-by: Claude <noreply@anthropic.com>
- Move legend inline into Details section (compact single line)
- Remove separate Legend collapsible section from helpers
- Remove includeLegend option from buildRerunFooter
- Remove ✅ emoji from "Results of HolmesGPT evals" header

Signed-off-by: Claude <noreply@anthropic.com>
Conflict in github_reporter.py:
- Master re-added per-test-type summary lines (ask_holmes: X/Y, etc.)
- Kept our simplified version (just status banner, no per-type breakdown)

Signed-off-by: Claude <noreply@anthropic.com>
@github-actions

github-actions Bot commented Feb 2, 2026

Copy link
Copy Markdown
Contributor

🔬 CLI Performance Benchmark

🟡 Startup Time (no LLM)

Measures holmes version execution time (imports + initialization)

Metric PR Master Change
Cold Start 9.71s 10.13s -4.1%
Warm Mean 4.56s 4.88s -6.4%
Warm Min 4.53s 4.78s
Warm Max 4.63s 5.04s

🟡 Full CLI with LLM

Measures holmes ask execution time (OpenRouter + Haiku 4.5)

Metric PR Master Change
Cold Start 27.75s 29.30s -5.3%
Warm Mean 7.75s 8.24s -5.9%
Warm Min 7.68s 8.14s
Warm Max 7.79s 8.37s

PR: e6d64cd3 | Master: 73ecc9ef | Iterations: 5

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Fix all issues with AI agents
In `@tests/llm/utils/reporting/github_reporter.py`:
- Around line 181-184: generate_markdown_report currently sums undefined
variables workload_health_total, workload_health_passed, and
workload_health_regressions which causes a NameError; either remove those
workload_health references from the totals calculation and compute
total_tests/total_passed/total_regressions from the existing counters
(ask_holmes_total, investigate_total, ask_holmes_passed, investigate_passed,
ask_holmes_regressions, investigate_regressions) or, if workload_health metrics
were intended, add explicit initialization and increment logic for
workload_health_total, workload_health_passed, and workload_health_regressions
in the same counting loop that handles 'ask' and 'investigate' inside
generate_markdown_report so the variables exist before being summed.

Comment on lines +181 to +184
# Calculate totals for overall status
total_tests = ask_holmes_total + investigate_total + workload_health_total
total_passed = ask_holmes_passed + investigate_passed + workload_health_passed
total_regressions = ask_holmes_regressions + investigate_regressions + workload_health_regressions

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🔴 Critical

Critical: Undefined variables cause NameError at runtime.

Lines 182-184 reference workload_health_total, workload_health_passed, and workload_health_regressions, but these variables are never defined. The counting loop (lines 153-179) only handles ask and investigate test types—there's no workload_health counter initialization or increment logic.

This will crash generate_markdown_report() with a NameError whenever it's called.

🐛 Proposed fix: Remove undefined workload_health references
     # Calculate totals for overall status
-    total_tests = ask_holmes_total + investigate_total + workload_health_total
-    total_passed = ask_holmes_passed + investigate_passed + workload_health_passed
-    total_regressions = ask_holmes_regressions + investigate_regressions + workload_health_regressions
+    total_tests = ask_holmes_total + investigate_total
+    total_passed = ask_holmes_passed + investigate_passed
+    total_regressions = ask_holmes_regressions + investigate_regressions
🧰 Tools
🪛 Ruff (0.14.14)

[error] 182-182: Undefined name workload_health_total

(F821)


[error] 183-183: Undefined name workload_health_passed

(F821)


[error] 184-184: Undefined name workload_health_regressions

(F821)

🤖 Prompt for AI Agents
In `@tests/llm/utils/reporting/github_reporter.py` around lines 181 - 184,
generate_markdown_report currently sums undefined variables
workload_health_total, workload_health_passed, and workload_health_regressions
which causes a NameError; either remove those workload_health references from
the totals calculation and compute total_tests/total_passed/total_regressions
from the existing counters (ask_holmes_total, investigate_total,
ask_holmes_passed, investigate_passed, ask_holmes_regressions,
investigate_regressions) or, if workload_health metrics were intended, add
explicit initialization and increment logic for workload_health_total,
workload_health_passed, and workload_health_regressions in the same counting
loop that handles 'ask' and 'investigate' inside generate_markdown_report so the
variables exist before being summed.

@aantn aantn closed this Feb 22, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants