Skip to content

Add PR comment slash command for manual eval runs - #1230

Closed
aantn wants to merge 1 commit into
masterfrom
codex/linear-mention-rob-52-manually-trigger-evals-on-github-3zegrt
Closed

aantn wants to merge 1 commit into
masterfrom
codex/linear-mention-rob-52-manually-trigger-evals-on-github-3zegrt

Conversation

@aantn

@aantn aantn commented Dec 23, 2025 •

Copy link
Copy Markdown
Collaborator

Summary

  • add an issue_comment workflow that parses /run-evals requests from authorized collaborators and runs the requested eval selection
  • post a fresh comment for every manual run with parameters, duration, results, and a link to the uploaded log
  • document the automated PR regression run and how to manually trigger custom evals via slash command (with defaults and simplified examples)
  • update the regression CI comment to include inline instructions for re-running evals from a PR comment, highlighting defaults and removing test-path usage
  • fix manual eval workflow outputs to write to GITHUB_OUTPUT and escape the regression comment separator, and ensure regression workflow YAML is valid
  • clarify docs that a regression eval slice runs automatically on every PR
  • announce manual eval kickoff immediately via PR comment when a /run-evals request is accepted

Testing

  • Not run (CI workflow change only)

Codex Task

Summary by CodeRabbit

  • New Features

    • Manual evaluation runs can now be triggered via PR comments with customizable parameters and defaults.
  • Improvements

    • Evaluation reports now include manual rerun instructions with command examples.
  • Documentation

    • Added comprehensive guidance on triggering evaluations from PR comments, including command formats and usage examples.

✏️ Tip: You can customize this high-level summary in your review settings.

Signed-off-by: Codex <codex@openai.com>
@coderabbitai

coderabbitai Bot commented Dec 23, 2025 •

Copy link
Copy Markdown
Contributor

Walkthrough

The PR introduces a new GitHub Actions workflow enabling manual HolmesGPT evaluations triggered via PR comments, updates an existing evaluation workflow to include rerun instructions in its report, and adds documentation on how to invoke evaluations from PR comments.

Changes

Cohort / File(s) Summary
GitHub Actions Workflow — Manual Evaluation Trigger
\.github/workflows/manual-evals.yaml
New workflow that parses /run-evals comments on PRs, validates invoker permissions, sets up HolmesGPT and KIND cluster, executes pytest with dynamically constructed parameters, captures logs and metrics, uploads artifacts, generates a Markdown report, and posts results back to the PR.
GitHub Actions Workflow — Report Enhancement
\.github/workflows/eval-regression.yaml
Adds conditional step to append manual rerun instructions and command examples to the evaluation report (evals_report.md) after tests complete.
Documentation — Evaluation Execution Guide
docs/development/evaluations/running-evals.md
Adds two documentation sections explaining how to trigger evaluations from PR comments, including permission requirements, command syntax with defaults, usage examples, and description of posted results.

Sequence Diagram(s)

sequenceDiagram
    actor User
    participant GH as GitHub
    participant WF as Workflow
    participant Parser
    participant Setup
    participant Pytest
    participant Report
    participant PR as PR Comments

    User->>GH: Comment with /run-evals command
    GH->>WF: Trigger workflow on comment
    
    WF->>WF: Check user permissions<br/>(collaborator/member/owner)
    alt Unauthorized
        WF->>PR: Post rejection message
    else Authorized
        WF->>Parser: Parse /run-evals comment
        Parser->>WF: Extract test path, keywords,<br/>markers, workers, defaults
        
        WF->>PR: Announce eval start
        
        WF->>Setup: Setup HolmesGPT environment<br/>and KIND cluster
        Setup->>WF: Environment ready
        
        WF->>Pytest: Execute pytest with<br/>constructed command
        Pytest->>Pytest: Run evaluations
        Pytest->>WF: Capture logs & metrics
        
        WF->>Report: Generate Markdown report<br/>(status, timing, command)
        WF->>GH: Upload eval log artifact
        
        Report->>PR: Post final result comment<br/>(report + invocation details)
    end
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Possibly related PRs

  • More doc improvements #1094: Modifies the same documentation file (docs/development/evaluations/running-evals.md) for evaluation execution guidance, indicating potential coordination or overlap in evaluation documentation changes.

Suggested reviewers

  • moshemorad
  • arikalon1

Pre-merge checks

✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The pull request title directly and clearly summarizes the main change: adding a PR comment slash command feature for manually triggering evaluation runs. It accurately reflects the core objective without being vague or off-topic.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (3)
.github/workflows/eval-regression.yaml (1)

65-85: The backslash escape for the separator is unnecessary in Markdown context.

Line 70 uses \--- with a comment about escaping to avoid YAML document start. However, this heredoc writes to a Markdown file (evals_report.md), not a YAML file. The backslash will appear literally in the output. If the intent is a horizontal rule, use --- without the backslash; if it's meant to be a visual separator, consider using --- or *** directly.

Suggested fix
          cat <<'EOF' >> evals_report.md
-          \---  # separator (escaped to avoid YAML document start)
+
+          ---

           Want to rerun evals with custom filters or models? Comment on this PR with:
docs/development/evaluations/running-evals.md (2)

202-212: Add a language specifier to the fenced code block.

The code block starting at line 204 lacks a language identifier, which triggers MD040. Since this is a command template, consider using text or bash as the language.

Suggested fix
 **Command format (defaults in parentheses)**

-```
+```text
 /run-evals models=<comma-separated models> \        # default: gpt-4o
           markers="<pytest -m expression>" \        # default: "llm and easy"

216-224: Consider using proper headings instead of bold text for section labels.

Lines 202 and 216 use bold emphasis (**Command format...**, **Examples**) as pseudo-headings. Converting these to ### headings would improve document structure, accessibility, and resolve the MD036 linting warnings.

Suggested fix
-**Command format (defaults in parentheses)**
+### Command format (defaults in parentheses)
-**Examples**
+### Examples
📜 Review details

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between d8ec062 and 85a5f47.

📒 Files selected for processing (3)
  • .github/workflows/eval-regression.yaml
  • .github/workflows/manual-evals.yaml
  • docs/development/evaluations/running-evals.md
🧰 Additional context used
📓 Path-based instructions (1)
docs/**/*.md

📄 CodeRabbit inference engine (CLAUDE.md)

When writing documentation in the docs/ directory, always add a blank line between headers/bold text and lists for proper MkDocs rendering

Files:

  • docs/development/evaluations/running-evals.md
🧠 Learnings (4)
📚 Learning: 2025-12-21T13:17:48.366Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:48.366Z
Learning: LLM evaluation tests run automatically in CI

Applied to files:

  • .github/workflows/eval-regression.yaml
  • docs/development/evaluations/running-evals.md
📚 Learning: 2025-12-21T13:17:48.366Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:48.366Z
Learning: Applies to tests/llm/** : Implement full architecture even if complex in evals (e.g., use Loki for log aggregation properly, not simplified alternatives)

Applied to files:

  • docs/development/evaluations/running-evals.md
📚 Learning: 2025-12-21T13:17:48.366Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:48.366Z
Learning: Applies to tests/llm/** : No fake/obvious logs in eval scenarios - avoid logs like 'Memory usage stabilized at 800MB'

Applied to files:

  • docs/development/evaluations/running-evals.md
📚 Learning: 2025-12-21T13:17:48.366Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:48.366Z
Learning: Applies to tests/llm/**/*.yaml : Never use resource names that hint at the problem or expected behavior in evals - use neutral names that don't give away what the LLM should discover

Applied to files:

  • docs/development/evaluations/running-evals.md
🪛 LanguageTool
docs/development/evaluations/running-evals.md

[uncategorized] ~196-~196: The official name of this software platform is spelled with a capital “H”.
Context: ...regression slice of the eval suite (see .github/workflows/eval-regression.yaml). Use t...

(GITHUB)

🪛 markdownlint-cli2 (0.18.1)
docs/development/evaluations/running-evals.md

202-202: Emphasis used instead of a heading

(MD036, no-emphasis-as-heading)


204-204: Fenced code blocks should have a language specified

(MD040, fenced-code-language)


216-216: Emphasis used instead of a heading

(MD036, no-emphasis-as-heading)

⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (5)
  • GitHub Check: llm_evals
  • GitHub Check: build (3.12)
  • GitHub Check: build (3.10)
  • GitHub Check: build (3.11)
  • GitHub Check: build
🔇 Additional comments (7)
docs/development/evaluations/running-evals.md (1)

194-233: Documentation aligns well with the workflow implementation.

The documented parameters, defaults, and command format match the implementation in .github/workflows/manual-evals.yaml. The examples are practical and demonstrate common use cases effectively.

.github/workflows/manual-evals.yaml (6)

3-15: Workflow trigger and job condition are correctly configured.

The issue_comment trigger with the PR check (github.event.issue.pull_request != '') is the correct pattern for handling slash commands on pull requests. Permissions are appropriately scoped.


93-106: Authorization feedback mechanism is well-designed.

The condition ensures a response is only posted when there's a meaningful message (unauthorized attempt), avoiding noise on non-trigger comments.


223-252: Report generation is comprehensive and informative.

The generated report includes all relevant parameters, timing information, and a clear success/failure indicator. The download link correctly points to the workflow run for log access.


254-261: Fresh comment per manual run is correctly configured.

Setting delete-previous: 'false' ensures each manual evaluation run posts a new comment, preserving the history of manual runs as intended by the PR objectives.


62-62: The test_path parameter is user-controllable but trusted.

The test_path parameter defaults to tests/llm/ but could be overridden by authorized users. Since only collaborators/members/owners can trigger this workflow (checked at line 48), this is acceptable. However, note that test_path is not documented in the docs or workflow instructions—users may not know they can customize it.

Is test_path intentionally undocumented to discourage its use, or should it be added to the documentation?


185-185: This hardcoding is intentional and appropriate—no refactoring needed.

The --strict-setup-exceptions=22_high_latency_dbi_down targets a specific test case fixture (matching tests/llm/fixtures/test_ask_holmes/22_high_latency_dbi_down/test_case.yaml), not an environment-specific setting. The value is documented in evaluation results and is not intended to change. Parameterizing this would add unnecessary complexity without benefit.

set_output("response-message", response_message)
exit(0)

tokens = shlex.split(body)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Wrap shlex.split() in a try-except to handle malformed input gracefully.

If a user posts a comment with unclosed quotes (e.g., /run-evals models="gpt-4o), shlex.split() will raise a ValueError, causing the step to fail with an unclear error. Consider catching this and posting a helpful response.

Suggested fix
-          tokens = shlex.split(body)
+          try:
+              tokens = shlex.split(body)
+          except ValueError as e:
+              response_message = f"⚠️ Could not parse command: {e}. Check for unclosed quotes."
+              set_output("should-run", "false")
+              set_output("response-message", response_message)
+              exit(0)
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
tokens = shlex.split(body)
try:
tokens = shlex.split(body)
except ValueError as e:
response_message = f"⚠️ Could not parse command: {e}. Check for unclosed quotes."
set_output("should-run", "false")
set_output("response-message", response_message)
exit(0)
🤖 Prompt for AI Agents
In .github/workflows/manual-evals.yaml around line 56, the call tokens =
shlex.split(body) can raise ValueError for malformed input (e.g., unclosed
quotes); wrap this call in a try/except that catches ValueError, handle it by
posting a clear user-facing response (or logging) explaining the parse error and
how to format the command, and ensure the workflow step exits gracefully (do not
let the exception propagate and fail the job).

@github-actions

Copy link
Copy Markdown
Contributor

Results of HolmesGPT evals

  • ask_holmes: 31/37 test cases were successful, 1 regressions, 2 setup failures, 3 mock failures
Test suite Test case Status
ask 01_how_many_pods ✅
ask 02_what_is_wrong_with_pod ✅
ask 04_related_k8s_events ✅
ask 05_image_version ✅
ask 09_crashpod ✅
ask 10_image_pull_backoff ✅
ask 110_k8s_events_image_pull ✅
ask 11_init_containers ✅
ask 13a_pending_node_selector_basic ✅
ask 14_pending_resources ✅
ask 15_failed_readiness_probe ✅
ask 163_compaction_follow_up ✅
ask 17_oom_kill 🚧
ask 18_oom_kill_from_issues_history ✅
ask 19_detect_missing_app_details ✅
ask 20_long_log_file_search ❌
ask 24_misconfigured_pvc ✅
ask 24a_misconfigured_pvc_basic ✅
ask 28_permissions_error 🚧
ask 39_failed_toolset ✅
ask 41_setup_argo ✅
ask 42_dns_issues_steps_new_tools ✅
ask 43_current_datetime_from_prompt ✅
ask 45_fetch_deployment_logs_simple ✅
ask 51_logs_summarize_errors ✅
ask 53_logs_find_term ✅
ask 54_not_truncated_when_getting_pods ✅
ask 59_label_based_counting ✅
ask 60_count_less_than ✅
ask 61_exact_match_counting ✅
ask 63_fetch_error_logs_no_errors ✅
ask 79_configmap_mount_issue ✅
ask 83_secret_not_found ✅
ask 86_configmap_like_but_secret ✅
ask 93_calling_datadog[0] 🔧
ask 93_calling_datadog[1] 🔧
ask 93_calling_datadog[2] 🔧

Legend

  • ✅ the test was successful
  • :minus: the test was skipped
  • ⚠️ the test failed but is known to be flaky or known to fail
  • 🚧 the test had a setup failure (not a code regression)
  • 🔧 the test failed due to mock data issues (not a code regression)
  • 🚫 the test was throttled by API rate limits/overload
  • ❌ the test failed and should be fixed before merging the PR--- # separator (escaped to avoid YAML document start)

Want to rerun evals with custom filters or models? Comment on this PR with:

/run-evals models=gpt-4o markers="llm and easy" keyword=""

Defaults: models=gpt-4o, markers="llm and easy", keyword="", iterations=1, workers=6, classifier_model=gpt-4o.

Examples:

  • Focus a test: /run-evals keyword=80_pvc_storage_class_mismatch iterations=2
  • Change models/markers: /run-evals models="gpt-4o,anthropic/claude-sonnet-4-20250514" markers="llm and medium" workers=4

See docs/development/evaluations/running-evals.md for details.

@github-actions

Copy link
Copy Markdown
Contributor

@github-actions

Copy link
Copy Markdown
Contributor

@aantn

aantn commented Dec 23, 2025

Copy link
Copy Markdown
Collaborator Author

/run-evals keyword=80_pvc_storage_class_mismatch iterations=2

@aantn

aantn commented Dec 24, 2025

Copy link
Copy Markdown
Collaborator Author

/run-evals models=gpt-4o markers="llm and easy" keyword=""

@aantn
aantn marked this pull request as draft December 25, 2025 05:33
@aantn aantn closed this Dec 28, 2025
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant