Skip to content

Improvements to evals framework - #702

Merged
aantn merged 19 commits into
masterfrom
session-scoped-setup-v3
Jul 24, 2025
Merged

aantn merged 19 commits into
masterfrom
session-scoped-setup-v3

Conversation

@aantn

@aantn aantn commented Jul 24, 2025

Copy link
Copy Markdown
Collaborator

Many improvements to evals framework, primarily:

  • Perform each eval's before_test and after_test only once even if the eval is run multiple times (i.e. ITERATIONS=100)
  • Allow skipping setup or cleanup so you can iterate on an eval faster (use --skip-setup and --skip-cleanup)
  • Lots of refactoring of evals framework and remove redundancies + reorganize code more logically

Note:

  • Evals no longer report timing on before_test and after_test to braintrust, as this happens globally before any tests run, not as part of the test itself. However, you can still track long setups and teardowns easily - they are reported at the end of every test run.

@aantn
aantn requested a review from Sheeproid July 24, 2025 08:53
@coderabbitai

coderabbitai Bot commented Jul 24, 2025 •

Copy link
Copy Markdown
Contributor

Warning

Rate limit exceeded

@aantn has exceeded the limit for the number of commits or files that can be reviewed per hour. Please wait 15 minutes and 29 seconds before requesting another review.

⌛ How to resolve this issue?

After the wait time has elapsed, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout.

Please see our FAQ for further information.

📥 Commits

Reviewing files that changed from the base of the PR and between 45b66aa and 3b14f09.

📒 Files selected for processing (1)
  • tests/llm/utils/commands.py (3 hunks)

Walkthrough

This update introduces a new top-level conftest.py to centralize pytest configuration, including custom CLI options for mock generation and setup/cleanup control, and logging suppression for noisy libraries. It delegates the test summary reporting hook to an external function. Supporting modules and documentation are updated to reflect these changes, and related test, utility, and reporting modules are refactored for modularity and improved output handling.

Changes

Files/Paths Change Summary
conftest.py New file: Adds pytest CLI options for mock generation and setup/cleanup, configures logging suppression, assigns terminal summary hook to external function.
pyproject.toml Adds pytest-shared-session-scope dependency, updates pytest logging config, increases slowest test count.
tests/llm/conftest.py Major refactor: Removes legacy setup/teardown/reporting logic, introduces session-scoped fixtures for infrastructure coordination, updates/renames fixtures, delegates reporting to external modules.
tests/llm/test_ask_holmes.py Refactors test: Early property setup, removes inline setup/teardown, improves output/logging via helpers, integrates with new fixtures and reporting utilities.
tests/llm/test_investigate.py, tests/llm/test_workload_health.py Updates tests to use new property management utilities for initializing/updating test metadata and results.
tests/llm/utils/braintrust.py Removes data-fetching utilities, adds URL generation function for Braintrust experiment links.
tests/llm/utils/commands.py Refactors command execution: Adds CommandResult class, centralizes command running logic in run_commands, removes separate before/after functions.
tests/llm/utils/langfuse.py Removes unused function for dataset item resolution.
tests/llm/utils/mock_toolset.py Adds session-level mock clearing/reporting utilities, introduces MockGenerationConfig, refactors mock clearing logic, improves error handling and reporting.
tests/llm/utils/property_manager.py New module: Provides utilities for managing test property initialization and result updates in pytest.
tests/llm/utils/setup_cleanup.py New module: Adds parallelized setup/cleanup command execution and test case extraction utilities for session-scoped fixtures.
tests/llm/utils/system.py Updates branch name detection to support Git worktree setups and adds error handling.
tests/llm/utils/test_case_utils.py Lowers logging verbosity for test case loading to debug level.
tests/llm/utils/test_helpers.py New module: Adds output truncation, tool call printing, correctness evaluation, and span logging helpers for tests.
tests/llm/utils/test_mock_toolset.py Updates test to use renamed mock clearing method.
tests/llm/utils/test_results.py New module: Adds TestResult and TestStatus classes for test outcome encapsulation and status reporting.
tests/llm/utils/reporting/github_reporter.py New module: Adds GitHub Actions report generation and markdown summary utilities.
tests/llm/utils/reporting/terminal_reporter.py New module: Adds Rich-based terminal reporting and LLM-powered failure analysis for test results.
tests/llm/utils/reporting/__init__.py Adds module docstring.
CLAUDE.md, docs/development/evals/index.md, docs/development/evals/writing.md Documentation: Adds CLI flag references, environment variable usage, marker lists, and usage patterns for new pytest options and workflows.

Sequence Diagram(s)

sequenceDiagram
    participant User
    participant Pytest
    participant conftest.py
    participant ReportingModules
    participant TestModules

    User->>Pytest: Run pytest with CLI flags (e.g., --generate-mocks)
    Pytest->>conftest.py: Parse CLI options, configure logging
    Pytest->>TestModules: Execute tests (setup/teardown via fixtures)
    TestModules->>ReportingModules: Update test properties/results
    Pytest->>ReportingModules: On terminal summary, call show_llm_summary_report
    ReportingModules->>ReportingModules: Aggregate results, generate reports
    ReportingModules->>User: Output summary (terminal, GitHub, mock report)
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~40 minutes

Possibly related PRs

  • #620: Refactors and modularizes pytest terminal summary and test result aggregation logic, directly related to the movement and delegation of the summary hook in this PR.
  • #687: Improves evals mock generation, adds detailed mock generation flags and reporting, closely related to the new CLI options and mock reporting in this PR.
  • #619: Modifies the internal logic of the pytest terminal summary function, related to this PR's reassignment and modularization of summary reporting.

Suggested reviewers

  • moshemorad
✨ Finishing Touches
  • 📝 Generate Docstrings
🧪 Generate unit tests
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch session-scoped-setup-v3

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share
🪧 Tips

Chat

There are 3 ways to chat with CodeRabbit:

  • Review comments: Directly reply to a review comment made by CodeRabbit. Example:
    • I pushed a fix in commit <commit_id>, please review it.
    • Explain this complex logic.
    • Open a follow-up GitHub issue for this discussion.
  • Files and specific lines of code (under the "Files changed" tab): Tag @coderabbitai in a new review comment at the desired location with your query. Examples:
    • @coderabbitai explain this code block.
    • @coderabbitai modularize this function.
  • PR comments: Tag @coderabbitai in a new PR comment to ask questions about the PR branch. For the best results, please provide a very specific query, as very limited context is provided in this mode. Examples:
    • @coderabbitai gather interesting stats about this repository and render them as a table. Additionally, render a pie chart showing the language distribution in the codebase.
    • @coderabbitai read src/utils.ts and explain its main purpose.
    • @coderabbitai read the files in the src/scheduler package and generate a class diagram using mermaid and a README in the markdown format.
    • @coderabbitai help me debug CodeRabbit configuration file.

Support

Need help? Create a ticket on our support page for assistance with any issues or questions.

Note: Be mindful of the bot's finite context window. It's strongly recommended to break down tasks such as reading entire modules into smaller chunks. For a focused discussion, use review comments to chat about specific files and their changes, instead of using the PR comments.

CodeRabbit Commands (Invoked using PR comments)

  • @coderabbitai pause to pause the reviews on a PR.
  • @coderabbitai resume to resume the paused reviews.
  • @coderabbitai review to trigger an incremental review. This is useful when automatic reviews are disabled for the repository.
  • @coderabbitai full review to do a full review from scratch and review all the files again.
  • @coderabbitai summary to regenerate the summary of the PR.
  • @coderabbitai generate docstrings to generate docstrings for this PR.
  • @coderabbitai generate sequence diagram to generate a sequence diagram of the changes in this PR.
  • @coderabbitai generate unit tests to generate unit tests for this PR.
  • @coderabbitai resolve resolve all the CodeRabbit review comments.
  • @coderabbitai configuration to show the current CodeRabbit configuration for the repository.
  • @coderabbitai help to get help.

Other keywords and placeholders

  • Add @coderabbitai ignore anywhere in the PR description to prevent this PR from being reviewed.
  • Add @coderabbitai summary to generate the high-level summary at a specific location in the PR description.
  • Add @coderabbitai anywhere in the PR title to generate the title automatically.

CodeRabbit Configuration File (.coderabbit.yaml)

  • You can programmatically configure CodeRabbit by adding a .coderabbit.yaml file to the root of your repository.
  • Please see the configuration documentation for more information.
  • If your editor has YAML language server enabled, you can add the path at the top of this file to enable auto-completion and validation: # yaml-language-server: $schema=https://coderabbit.ai/integrations/schema.v2.json

Documentation and Community

  • Visit our Documentation for detailed information on how to use CodeRabbit.
  • Join our Discord Community to get help, request features, and share feedback.
  • Follow us on X/Twitter for updates and announcements.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 4

🧹 Nitpick comments (12)
conftest.py (1)

35-69: Consider removing or documenting commented code.

The large block of commented code for pytest-xdist worker-specific logging should either be removed if no longer needed or documented with a clear explanation of why it's preserved for future use.

-    # Configure worker-specific log files for xdist compatibility
-    # worker_id = getattr(config, "workerinput", {}).get("workerid", "master")
-    # if worker_id != "master":
-    #    # Set worker-specific log file to avoid conflicts
-    #    config.option.log_file = f"tests-{worker_id}.log"
-
-    # Determine worker id
-    # Also see: https://pytest-xdist.readthedocs.io/en/latest/how-to.html#creating-one-log-file-for-each-worker
-    # worker_id = os.environ.get("PYTEST_XDIST_WORKER", default="gw0")
-
-    # # Create logs folder
-    # logs_folder = os.environ.get("LOGS_FOLDER", default="logs_folder")
-    # os.makedirs(logs_folder, exist_ok=True)
-
-    # # Create file handler to output logs into corresponding worker file
-    # file_handler = logging.FileHandler(f"{logs_folder}/logs_worker_{worker_id}.log", mode="w")
-    # file_handler.setFormatter(
-    #     logging.Formatter(
-    #         fmt="{asctime} {levelname}:{name}:{lineno}:{message}",
-    #         style="{",
-    #     )
-    # )
-    # # Create stream handler to output logs on console
-    # # This is a workaround for a known limitation:
-    # # https://pytest-xdist.readthedocs.io/en/latest/known-limitations.html
-    # console_handler = logging.StreamHandler(sys.stderr)  # pytest only prints error logs
-    # console_handler.setFormatter(
-    #     logging.Formatter(
-    #         # Include worker id in log messages, \r is needed to separate lines in console
-    #         fmt="\r{asctime} " + worker_id + ":{levelname}:{name}:{lineno}:{message}",
-    #         style="{",
-    #     )
-    # )
-    # # Configure logging
-    # logging.basicConfig(level=logging.INFO, force=True, handlers=[console_handler, file_handler])
+    # TODO: Worker-specific logging for pytest-xdist can be implemented here if needed
+    # See: https://pytest-xdist.readthedocs.io/en/latest/how-to.html#creating-one-log-file-for-each-worker
tests/llm/utils/test_results.py (3)

22-32: Consider edge cases in test ID extraction.

The current logic assumes test cases follow the pattern test_name[number_description] and extracts the number prefix. However, if the test case doesn't contain underscores or has a different format, it might not work as expected.

Consider adding more robust parsing:

 @property
 def test_id(self) -> str:
     """Extract test ID from pytest nodeid.
 
     Example: 'test_ask_holmes[01_how_many_pods]' -> '01'
     """
     if "[" in self.nodeid and "]" in self.nodeid:
         test_case = self.nodeid.split("[")[1].split("]")[0]
-        # Extract number from start of test case name
-        return test_case.split("_")[0] if "_" in test_case else test_case
+        # Extract number from start of test case name
+        parts = test_case.split("_")
+        if parts and parts[0].isdigit():
+            return parts[0]
+        return test_case
     return "unknown"

66-72: Simplify the regression check logic.

The static analysis tool correctly identifies that this can be simplified by returning the negated condition directly.

 @property
 def is_regression(self) -> bool:
     if self.passed or self.is_mock_failure:
         return False
     # Known failure (expected to fail)
-    if self.actual_score == 0 and self.expected_score == 0:
-        return False
-    return True
+    return not (self.actual_score == 0 and self.expected_score == 0)

54-63: Consider the TODO comment about mock failures.

There's a TODO comment suggesting that mock_failures should potentially affect the passed property. This might impact the overall test reporting logic.

Should mock failures be considered as non-passing tests? The current logic treats them separately, but the TODO suggests they might need to be integrated into the pass/fail determination. Would you like me to help clarify this logic or open an issue to track this decision?

tests/llm/utils/property_manager.py (1)

36-43: Fix unused loop variable.

The static analysis tool correctly identifies that prop_value is not used in the loop body.

-    for i, (prop_key, prop_value) in enumerate(request.node.user_properties):
+    for i, (prop_key, _prop_value) in enumerate(request.node.user_properties):
         if prop_key == key:
             request.node.user_properties[i] = (key, value)
             return
tests/llm/utils/test_helpers.py (1)

52-61: Consider edge case in correctness evaluation printing.

The function assumes correctness_eval.metadata exists and contains a "rationale" key. Consider adding defensive checks.

 def print_correctness_evaluation(correctness_eval: Any) -> None:
     """Print correctness evaluation results."""
     print("\n⚖️ CORRECTNESS EVALUATION:")
     print(f"   Score: {correctness_eval.score}")
     print("   Rationale: ")
-    rationale = correctness_eval.metadata.get("rationale", "")
+    rationale = getattr(correctness_eval, 'metadata', {}).get("rationale", "")
     for line in rationale.split("\n"):
         if line.strip():
             print(f"      {line}")
tests/llm/reporting/terminal_reporter.py (1)

110-157: Consider reliability and cost implications of LLM analysis.

The LLM-powered failure analysis is innovative, but there are several considerations:

  1. Network dependency: Analysis will fail if the LLM API is unavailable
  2. Cost implications: Running GPT-4o analysis for every failed test could be expensive
  3. Rate limiting: Multiple concurrent requests might hit API limits

Consider adding configuration options to control when analysis runs:

# Add environment variable check
import os

def _get_llm_analysis(result: TestResult) -> str:
    if not os.getenv("ENABLE_LLM_ANALYSIS", "false").lower() == "true":
        return "LLM analysis disabled (set ENABLE_LLM_ANALYSIS=true to enable)"
    
    # Existing implementation...

Would you like me to help implement caching for analysis results or add configuration options to control when LLM analysis is performed?

tests/llm/utils/commands.py (1)

17-20: Consider using Optional[int] for exit_code parameter

The exit_code parameter defaults to None but is typed as int. This could lead to type checking issues.

-        exit_code: int = None,
+        exit_code: Optional[int] = None,
tests/llm/reporting/github_reporter.py (2)

18-19: Use context manager for file operations

While the current implementation works, using a context manager is more Pythonic and ensures proper file closure even if an exception occurs during writing.

-        with open("evals_report.txt", "w", encoding="utf-8") as file:
-            file.write(markdown)
+        Path("evals_report.txt").write_text(markdown, encoding="utf-8")

Note: You'll need to import Path from pathlib at the top of the file:

from pathlib import Path

32-40: Consider using a more structured approach for tracking test counts

The current variable initialization with multiple assignments on single lines can be hard to read and maintain. Consider using a dictionary or dataclass to organize these counts.

-    ask_holmes_total = ask_holmes_passed = ask_holmes_regressions = (
-        ask_holmes_mock_failures
-    ) = 0
-    investigate_total = investigate_passed = investigate_regressions = (
-        investigate_mock_failures
-    ) = 0
-    workload_health_total = workload_health_passed = workload_health_regressions = (
-        workload_health_mock_failures
-    ) = 0
+    test_counts = {
+        "ask": {"total": 0, "passed": 0, "regressions": 0, "mock_failures": 0},
+        "investigate": {"total": 0, "passed": 0, "regressions": 0, "mock_failures": 0},
+        "workload_health": {"total": 0, "passed": 0, "regressions": 0, "mock_failures": 0},
+    }

Then update the counting logic to use the dictionary structure. This would make the code more maintainable and easier to extend.

tests/llm/utils/mock_toolset.py (1)

723-734: Consider using contextlib.suppress for cleaner exception handling

The nested try-except blocks could be simplified, though the current implementation with fallback behavior might be intentional for robustness.

 def _safe_print(terminalreporter, message: str = "") -> None:
     """Safely print to terminal reporter to avoid I/O errors"""
-    try:
-        terminalreporter.write_line(message)
-    except Exception:
-        # If write_line fails, try direct write
-        try:
-            terminalreporter._tw.write(message + "\n")
-        except Exception:
-            # Last resort - ignore if all writing fails
-            pass
+    with contextlib.suppress(Exception):
+        terminalreporter.write_line(message)
+        return
+    
+    # If write_line failed, try direct write as fallback
+    with contextlib.suppress(Exception):
+        terminalreporter._tw.write(message + "\n")

Note: You'll need to import contextlib at the top of the file.

tests/llm/conftest.py (1)

323-323: Remove personal comment

The comment "NATAN - this link is correct" appears to be a personal note that should be removed.

-    # NATAN - this link is correct
     braintrust_url = get_braintrust_url(
📜 Review details

Configuration used: CodeRabbit UI
Review profile: CHILL
Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between a5fa1b7 and 28b1f24.

⛔ Files ignored due to path filters (1)
  • poetry.lock is excluded by !**/*.lock
📒 Files selected for processing (21)
  • conftest.py (1 hunks)
  • pyproject.toml (2 hunks)
  • tests/llm/conftest.py (7 hunks)
  • tests/llm/reporting/__init__.py (1 hunks)
  • tests/llm/reporting/github_reporter.py (1 hunks)
  • tests/llm/reporting/property_manager.py (1 hunks)
  • tests/llm/reporting/terminal_reporter.py (1 hunks)
  • tests/llm/test_ask_holmes.py (10 hunks)
  • tests/llm/test_investigate.py (3 hunks)
  • tests/llm/test_workload_health.py (3 hunks)
  • tests/llm/utils/braintrust.py (2 hunks)
  • tests/llm/utils/commands.py (2 hunks)
  • tests/llm/utils/langfuse.py (0 hunks)
  • tests/llm/utils/mock_toolset.py (7 hunks)
  • tests/llm/utils/property_manager.py (1 hunks)
  • tests/llm/utils/setup_cleanup.py (1 hunks)
  • tests/llm/utils/system.py (1 hunks)
  • tests/llm/utils/test_case_utils.py (2 hunks)
  • tests/llm/utils/test_helpers.py (1 hunks)
  • tests/llm/utils/test_mock_toolset.py (1 hunks)
  • tests/llm/utils/test_results.py (1 hunks)
💤 Files with no reviewable changes (1)
  • tests/llm/utils/langfuse.py
🧰 Additional context used
🧠 Learnings (2)
tests/llm/test_ask_holmes.py (2)

Learnt from: nherment
PR: #610
File: .github/workflows/llm-evaluation.yaml:39-42
Timestamp: 2025-07-08T08:45:41.069Z
Learning: The robusta-dev/holmesgpt codebase has comprehensive existing validation for Azure environment variables (AZURE_API_BASE, AZURE_API_KEY, AZURE_API_VERSION) and MODEL in tests/llm/utils/classifiers.py, tests/llm/conftest.py, and holmes/core/llm.py. Don't suggest adding redundant validation logic.

Learnt from: Sheeproid
PR: #586
File: tests/llm/fixtures/test_ask_holmes/03_what_is_the_command_to_port_forward/test_case.yaml:4-4
Timestamp: 2025-07-02T10:27:17.231Z
Learning: In LLM-as-judge test cases for HolmesGPT, expected outputs should be descriptive rather than prescriptive when testing for flexible responses like port numbers. Using specific values in expected outputs can cause unnecessary test failures when the AI generates different but equally valid responses.

tests/llm/utils/mock_toolset.py (1)

Learnt from: nherment
PR: #535
File: holmes/plugins/toolsets/bash/bash_toolset.py:207-209
Timestamp: 2025-06-24T05:51:04.543Z
Learning: The init_config method in toolsets should be idempotent - safely callable multiple times without errors. self.config should maintain consistent typing (not alternate between dict and config object types) throughout the object lifecycle.

🧬 Code Graph Analysis (5)
tests/llm/utils/test_mock_toolset.py (1)
tests/llm/utils/mock_toolset.py (1)
  • clear_mocks_for_test (265-293)
tests/llm/test_workload_health.py (1)
tests/llm/utils/property_manager.py (2)
  • set_initial_properties (7-33)
  • update_test_results (46-56)
tests/llm/test_investigate.py (2)
tests/llm/utils/test_case_utils.py (2)
  • InvestigateTestCase (64-68)
  • MockHelper (78-155)
tests/llm/utils/property_manager.py (2)
  • set_initial_properties (7-33)
  • update_test_results (46-56)
tests/llm/utils/property_manager.py (1)
tests/llm/utils/test_case_utils.py (2)
  • Evaluation (27-29)
  • HolmesTestCase (43-56)
tests/llm/utils/braintrust.py (1)
tests/llm/utils/test_results.py (2)
  • test_id (23-32)
  • test_name (35-48)
🪛 Ruff (0.12.2)
tests/llm/utils/test_results.py

70-72: Return the negated condition directly

Inline condition

(SIM103)

tests/llm/utils/property_manager.py

38-38: Loop control variable prop_value not used within loop body

Rename unused prop_value to _prop_value

(B007)

tests/llm/utils/mock_toolset.py

729-733: Use contextlib.suppress(Exception) instead of try-except-pass

(SIM105)

🔇 Additional comments (33)
tests/llm/reporting/__init__.py (1)

1-1: LGTM - Clean package initialization.

The docstring clearly describes the purpose of this new reporting package namespace.

tests/llm/utils/test_mock_toolset.py (1)

206-206: LGTM - Consistent with method rename.

The test correctly uses the renamed clear_mocks_for_test() method, which better clarifies that it clears mocks for a single test case folder rather than all mocks globally.

tests/llm/utils/test_case_utils.py (1)

99-99: LGTM - Appropriate logging level adjustment.

Changing test case loading messages from info to debug level reduces log noise during normal test runs while maintaining visibility when debug logging is enabled. This aligns well with the broader test infrastructure improvements.

Also applies to: 145-145, 147-149, 153-153

tests/llm/utils/system.py (1)

10-35: LGTM - Robust Git worktree support.

The enhanced get_active_branch_name() function properly handles both standard Git directories and Git worktree setups. The implementation correctly:

  • Detects when .git is a file (worktree case) vs directory
  • Parses the gitdir: format to find the actual Git directory
  • Falls back gracefully with proper error handling
  • Uses modern pathlib for better path operations

This will improve branch detection reliability across different Git configurations.

tests/llm/test_investigate.py (3)

24-25: LGTM - Clean integration of property management utilities.

The new imports bring in the standardized property management utilities that replace manual property handling throughout the test.


113-114: LGTM - Early property initialization.

Calling set_initial_properties() early ensures test metadata is available even if the test fails before completion, improving debugging and reporting capabilities.


214-215: LGTM - Consolidated result handling.

The update_test_results() call replaces the previous manual appending of multiple user properties, centralizing test result data management and improving maintainability.

tests/llm/test_workload_health.py (3)

26-26: Good integration with new property management utilities.

The import of property management utilities aligns well with the framework refactoring goals to centralize and standardize test metadata handling.


105-106: Excellent defensive programming approach.

Setting initial properties early ensures test metadata is available even if the test fails prematurely, which improves debugging and reporting reliability.


178-179: Clean consolidation of property updates.

Replacing manual property appends with the update_test_results utility function reduces code duplication and improves maintainability across the test suite.

conftest.py (2)

5-30: Well-designed CLI options support the framework improvements.

The new pytest options (--generate-mocks, --regenerate-all-mocks, --skip-setup, --skip-cleanup) directly support the PR objectives of enabling faster iteration on evals by avoiding repeated initialization or cleanup steps.


71-84: Good logging noise reduction.

Suppressing verbose logs from LiteLLM, httpx, and related libraries will significantly improve test output readability and focus attention on relevant test information.

pyproject.toml (3)

50-50: Good addition for session-scoped fixture support.

The pytest-shared-session-scope dependency directly supports the framework improvements for coordinated setup and cleanup operations mentioned in the PR objectives.


107-107: Improved test performance visibility.

Increasing the number of slowest tests reported from 5 to 10 will provide better insights into test performance, which is valuable when optimizing the evals framework.


110-120: Comprehensive logging configuration enhances debugging.

The detailed pytest logging configuration provides good visibility into test execution while maintaining reasonable log levels (INFO instead of DEBUG) to avoid excessive noise.

tests/llm/utils/braintrust.py (1)

178-213: Well-designed URL generation utility.

The get_braintrust_url function provides a clean, focused approach to Braintrust integration by generating URLs for test linking rather than fetching experiment data. The implementation correctly handles optional parameters, validates API key availability, and provides clear documentation.

tests/llm/reporting/property_manager.py (3)

6-21: Clean encapsulation of property management.

The TestPropertyManager class provides a well-structured approach to managing test metadata, with proper defensive programming (checking for user_properties attribute) and clear method organization.


22-41: Excellent standardization of test result properties.

The add_test_result method consolidates common test metadata patterns into a single, well-documented interface. This will significantly improve consistency across the test suite.


65-75: Creative pytest fixture registration approach.

While the pytest_plugin() function approach for fixture registration is less common than using conftest.py, it's a valid pattern that keeps the fixture close to its implementation and maintains good encapsulation.

tests/llm/utils/property_manager.py (1)

14-18: Good handling of different correctness evaluation types.

The code properly handles both Evaluation objects and direct values for the correctness score, which provides good flexibility for different test configurations.

tests/llm/utils/test_helpers.py (2)

18-19: Good backward compatibility approach.

The backward compatibility alias ensures existing code continues to work while providing a more descriptive function name for new usage.


63-77: Robust error handling in span logging.

The function properly handles cases where tool results might not have a data attribute and falls back to string representation. Good defensive programming.

tests/llm/utils/setup_cleanup.py (3)

44-49: Good resource management with ThreadPoolExecutor.

The code properly limits the number of workers to avoid overwhelming the system while ensuring efficient parallel execution.


80-85: Excellent error visibility with warnings.

Using warnings.warn() ensures that timeout and error information is visible in pytest output, which is crucial for debugging test infrastructure issues.


136-156: Efficient deduplication logic.

The function properly handles deduplication of test cases based on ID while filtering for cases that need setup. The logic is clear and efficient.

tests/llm/reporting/terminal_reporter.py (2)

126-146: Well-structured prompt engineering.

The prompt provides comprehensive context including test details, expected vs actual output, tools called, and error messages. The categorization system helps users understand different types of failures.


17-33: Good table design with fixed widths.

The table columns are appropriately sized for terminal display, and the use of Rich styling makes the output readable and professional.

tests/llm/utils/commands.py (2)

61-133: Well-structured command execution with comprehensive error handling

The refactored _run_commands function effectively consolidates the command execution logic with proper error handling for different failure scenarios. The use of CommandResult to encapsulate outcomes provides a clean interface for callers.


135-148: Clean refactoring of setup/cleanup functions

The simplified before_test and after_test functions effectively delegate to the shared _run_commands implementation, reducing code duplication while maintaining a clear interface.

tests/llm/test_ask_holmes.py (2)

85-86: Good practice: Early property initialization

Setting initial properties early in the test ensures that metadata is available even if the test fails during setup. This improves observability and debugging capabilities.


166-167: Excellent integration with centralized infrastructure

The use of log_tool_calls_to_spans helper function provides consistent tool call logging across all tests, improving traceability in Braintrust.

tests/llm/utils/mock_toolset.py (1)

76-135: Well-implemented session-level mock clearing

The clear_all_mocks function provides a robust mechanism for clearing mock files across all test cases. The error handling and progress reporting are well done.

tests/llm/conftest.py (1)

67-152: Excellent implementation of shared test infrastructure coordination

The shared_test_infrastructure fixture effectively uses pytest_shared_session_scope to coordinate setup and cleanup across parallel test workers. The implementation properly handles skip options and ensures operations are performed only once even with multiple workers.

Comment thread tests/llm/conftest.py Outdated
Comment thread tests/llm/conftest.py Outdated
Comment thread tests/llm/utils/mock_toolset.py
Comment thread tests/llm/utils/property_manager.py

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

♻️ Duplicate comments (1)
tests/llm/conftest.py (1)

202-202: Duplicate comment: Fix typo in fixture name

This was already identified in previous reviews.

🧹 Nitpick comments (5)
tests/llm/utils/reporting/terminal_reporter.py (1)

40-53: Consider using dataclass or factory method for TestResult construction

The TestResult construction is verbose with many parameters. Consider creating a factory method or using keyword arguments more systematically to improve readability.

-        test_result = TestResult(
-            nodeid=result.get("nodeid", ""),
-            expected=result["expected"],
-            actual=result["actual"],
-            pass_fail=pass_fail,
-            tools_called=result["tools_called"],
-            logs="",  # We don't have logs in this context
-            test_type=result["test_type"],
-            error_message=None,
-            execution_time=result.get("execution_time"),
-            expected_correctness_score=result["expected_correctness_score"],
-            actual_correctness_score=result["actual_correctness_score"],
-            mock_data_failure=result.get("mock_data_failure", False),
-        )
+        test_result = TestResult.from_result_dict(result, pass_fail=pass_fail)
tests/llm/utils/reporting/github_reporter.py (1)

32-40: Simplify variable initialization with tuple unpacking

The variable initialization can be simplified and made more readable.

-    ask_holmes_total = ask_holmes_passed = ask_holmes_regressions = (
-        ask_holmes_mock_failures
-    ) = 0
-    investigate_total = investigate_passed = investigate_regressions = (
-        investigate_mock_failures
-    ) = 0
-    workload_health_total = workload_health_passed = workload_health_regressions = (
-        workload_health_mock_failures
-    ) = 0
+    # Initialize counters for each test type
+    counters = {
+        "ask": {"total": 0, "passed": 0, "regressions": 0, "mock_failures": 0},
+        "investigate": {"total": 0, "passed": 0, "regressions": 0, "mock_failures": 0},
+        "workload_health": {"total": 0, "passed": 0, "regressions": 0, "mock_failures": 0},
+    }
tests/llm/utils/setup_cleanup.py (1)

78-83: Potential race condition in remaining cases calculation

The calculation of remaining cases could be inaccurate due to race conditions when multiple threads complete simultaneously.

-                remaining_cases = (
-                    len(test_cases)
-                    - successful_test_cases
-                    - failed_test_cases
-                    - timed_out_test_cases
-                )
+                completed_cases = successful_test_cases + failed_test_cases + timed_out_test_cases
+                remaining_cases = len(test_cases) - completed_cases - 1  # -1 for current case
tests/llm/utils/commands.py (1)

11-36: Consider using dataclass for CommandResult

The CommandResult class would benefit from using a dataclass to reduce boilerplate and improve type hints.

+from dataclasses import dataclass
+
-class CommandResult:
-    def __init__(
-        self,
-        command: str,
-        test_case_id: str,
-        success: bool,
-        exit_code: int = None,
-        elapsed_time: float = 0,
-        error_type: str = None,
-        error_details: str = None,
-    ):
-        self.command = command
-        self.test_case_id = test_case_id
-        self.success = success
-        self.exit_code = exit_code
-        self.elapsed_time = elapsed_time
-        self.error_type = error_type  # 'timeout', 'failure', or None
-        self.error_details = error_details
+@dataclass
+class CommandResult:
+    command: str
+    test_case_id: str
+    success: bool
+    exit_code: Optional[int] = None
+    elapsed_time: float = 0
+    error_type: Optional[str] = None  # 'timeout', 'failure', or None
+    error_details: Optional[str] = None
tests/llm/conftest.py (1)

130-147: Potential performance issue with test case reconstruction

The cleanup logic iterates through all session items to reconstruct test cases, which could be inefficient for large test suites.

-            # Reconstruct test cases from IDs
-            from tests.llm.utils.test_case_utils import HolmesTestCase  # type: ignore[attr-defined]  # type: ignore[attr-defined]
-
-            cleanup_test_cases = []
-
-            for item in request.session.items:
-                if (
-                    item.get_closest_marker("llm")
-                    and hasattr(item, "callspec")
-                    and "test_case" in item.callspec.params
-                ):
-                    test_case = item.callspec.params["test_case"]
-                    if (
-                        isinstance(test_case, HolmesTestCase)
-                        and test_case.id in test_case_ids
-                        and test_case not in cleanup_test_cases
-                    ):
-                        cleanup_test_cases.append(test_case)
+            # Store full test case objects in setup phase instead of just IDs
+            # This avoids expensive reconstruction during cleanup
📜 Review details

Configuration used: CodeRabbit UI
Review profile: CHILL
Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 28b1f24 and c4bde64.

⛔ Files ignored due to path filters (1)
  • poetry.lock is excluded by !**/*.lock
📒 Files selected for processing (10)
  • CLAUDE.md (1 hunks)
  • docs/development/evals/index.md (3 hunks)
  • docs/development/evals/writing.md (1 hunks)
  • pyproject.toml (2 hunks)
  • tests/llm/conftest.py (6 hunks)
  • tests/llm/utils/commands.py (2 hunks)
  • tests/llm/utils/reporting/__init__.py (1 hunks)
  • tests/llm/utils/reporting/github_reporter.py (1 hunks)
  • tests/llm/utils/reporting/terminal_reporter.py (1 hunks)
  • tests/llm/utils/setup_cleanup.py (1 hunks)
✅ Files skipped from review due to trivial changes (4)
  • tests/llm/utils/reporting/init.py
  • docs/development/evals/index.md
  • docs/development/evals/writing.md
  • CLAUDE.md
🚧 Files skipped from review as they are similar to previous changes (1)
  • pyproject.toml
🧰 Additional context used
🧬 Code Graph Analysis (1)
tests/llm/utils/setup_cleanup.py (3)
tests/llm/utils/commands.py (2)
  • run_commands (61-132)
  • exit_info (31-35)
tests/llm/utils/test_case_utils.py (1)
  • HolmesTestCase (43-56)
tests/llm/utils/test_helpers.py (1)
  • truncate_output (7-15)
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (3)
  • GitHub Check: Pre-commit checks
  • GitHub Check: llm_evals
  • GitHub Check: Pre-commit checks
🔇 Additional comments (12)
tests/llm/utils/reporting/terminal_reporter.py (3)

12-15: LGTM: Good early return pattern

The function correctly handles the empty results case with an early return, avoiding unnecessary processing.


148-157: Robust error handling for LLM API calls

Good practice to catch all exceptions and return a meaningful error message instead of letting the API failure crash the reporting.


119-119: litellm dependency is already declared in pyproject.toml

The litellm = "1.74.7" entry appears on line 43 of pyproject.toml, so the test-only import inside _get_llm_analysis is covered by your project’s dependencies. No further action required.

tests/llm/utils/setup_cleanup.py (3)

37-71: Excellent concurrent execution design

The function properly uses ThreadPoolExecutor with appropriate worker limits and handles both setup and cleanup operations cleanly. The future-to-test-case mapping is well-implemented.


95-107: Good timeout handling with comprehensive logging

The timeout handling is well-implemented with proper logging, error formatting, and warning emission to ensure visibility in pytest output.


156-176: Efficient deduplication logic

The function properly extracts unique test cases using a set for deduplication while maintaining the original order. Good use of early filtering based on required conditions.

tests/llm/utils/commands.py (2)

76-84: Good command processing logic

The command processing correctly handles empty lines and accumulates output properly. The timing measurement is accurate.


93-106: Comprehensive error handling for CalledProcessError

The error handling properly captures all relevant information including exit codes, stdout, stderr, and timing. The error details formatting is well-structured.

tests/llm/conftest.py (4)

44-63: Robust mock configuration handling

The fixture properly handles CLI options with safe defaults and correctly determines the mock mode based on environment and options.


68-84: Good early exit handling for shared fixture

The fixture correctly handles collect-only mode and mock mode by yielding twice as required by pytest-shared-session-scope, even when skipping setup.


333-347: Good xdist worker detection

The function properly detects xdist worker processes to avoid duplicate output. The worker ID detection logic is robust.


358-364: Excellent modularization of reporting

The reporting has been properly modularized with clear separation of concerns between GitHub output, console output, and mock operations reporting.

Comment thread tests/llm/utils/commands.py
Comment thread tests/llm/utils/reporting/github_reporter.py
Comment thread tests/llm/utils/reporting/github_reporter.py
@aantn
aantn enabled auto-merge (squash) July 24, 2025 10:07

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 0

♻️ Duplicate comments (1)
tests/llm/conftest.py (1)

335-367: Verify debug print statements have been removed

Based on previous review comments, debug print statements were flagged for removal throughout this file. Please confirm that all debug print statements have been cleaned up before merging, as they shouldn't be present in production code.

🧹 Nitpick comments (2)
holmes/core/tracing.py (1)

130-130: Consider investigating the necessity of type ignore comments

The # type: ignore comments on these return statements suggest typing issues. Consider whether these can be resolved through:

  1. Better type annotations for the Braintrust spans
  2. Refining the Union return type
  3. Using more specific type hints

If the comments are necessary due to third-party library limitations, consider adding brief explanatory comments.

Also applies to: 135-135

tests/llm/conftest.py (1)

35-41: Consider minor optimization for is_llm_test function

The function works correctly but could be slightly optimized:

def is_llm_test(nodeid: str) -> bool:
    """Check if a test nodeid is for an LLM test."""
-    return any(
-        [
-            "test_ask_holmes" in nodeid,
-            "test_investigate" in nodeid,
-            "test_workload_health" in nodeid,
-        ]
-    )
+    return any(
+        test_type in nodeid 
+        for test_type in ["test_ask_holmes", "test_investigate", "test_workload_health"]
+    )

Using a generator expression avoids creating an intermediate list.

📜 Review details

Configuration used: CodeRabbit UI
Review profile: CHILL
Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 0b9282d and 45b66aa.

📒 Files selected for processing (8)
  • holmes/core/tool_calling_llm.py (3 hunks)
  • holmes/core/tracing.py (2 hunks)
  • tests/llm/conftest.py (6 hunks)
  • tests/llm/test_ask_holmes.py (8 hunks)
  • tests/llm/test_investigate.py (4 hunks)
  • tests/llm/test_workload_health.py (5 hunks)
  • tests/llm/utils/setup_cleanup.py (1 hunks)
  • tests/llm/utils/test_helpers.py (1 hunks)
🚧 Files skipped from review as they are similar to previous changes (4)
  • tests/llm/test_investigate.py
  • tests/llm/utils/test_helpers.py
  • tests/llm/utils/setup_cleanup.py
  • tests/llm/test_ask_holmes.py
🧰 Additional context used
🧬 Code Graph Analysis (2)
holmes/core/tool_calling_llm.py (7)
holmes/core/tracing.py (2)
  • DummySpan (38-54)
  • start_span (41-42)
holmes/core/tools.py (2)
  • get_parameterized_one_liner (168-169)
  • get_parameterized_one_liner (195-202)
holmes/plugins/toolsets/logging_utils/logging_api.py (1)
  • get_parameterized_one_liner (105-132)
holmes/plugins/toolsets/internet/internet.py (1)
  • get_parameterized_one_liner (216-218)
holmes/plugins/toolsets/opensearch/opensearch.py (4)
  • get_parameterized_one_liner (104-105)
  • get_parameterized_one_liner (134-135)
  • get_parameterized_one_liner (162-163)
  • get_parameterized_one_liner (183-184)
holmes/plugins/toolsets/prometheus/prometheus.py (4)
  • get_parameterized_one_liner (355-356)
  • get_parameterized_one_liner (458-459)
  • get_parameterized_one_liner (567-570)
  • get_parameterized_one_liner (713-719)
holmes/plugins/toolsets/robusta/robusta.py (3)
  • get_parameterized_one_liner (76-77)
  • get_parameterized_one_liner (140-141)
  • get_parameterized_one_liner (198-199)
tests/llm/test_workload_health.py (2)
holmes/core/tracing.py (1)
  • SpanType (28-35)
tests/llm/utils/property_manager.py (2)
  • set_initial_properties (7-33)
  • update_test_results (46-56)
🪛 Ruff (0.12.2)
tests/llm/test_workload_health.py

138-139: Use a single with statement with multiple contexts instead of nested with statements

(SIM117)

⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (3)
  • GitHub Check: Pre-commit checks
  • GitHub Check: llm_evals
  • GitHub Check: Pre-commit checks
🔇 Additional comments (17)
holmes/core/tracing.py (3)

35-35: LGTM: Clean addition of EVAL span type

The new EVAL = "eval" enum member follows the existing pattern and provides appropriate categorization for evaluation contexts.


41-41: LGTM: Simplified dummy span method signature

Removing the explicit type annotation for span_type in the dummy implementation makes sense since this no-op method ignores all parameters anyway.


125-125: LGTM: Simplified span type handling

Using span_type.value directly instead of getattr(SpanTypeAttribute, span_type.name) is cleaner and more straightforward.

holmes/core/tool_calling_llm.py (3)

42-42: LGTM: Import cleanup aligns with span type refactor

Removing the unused SpanType import is appropriate since the code now uses string literals for span types.


425-425: LGTM: Consistent span type refactor

Using the string literal "tool" instead of SpanType.TOOL aligns with the broader tracing refactor and maintains the same functionality.


454-455: LGTM: Enhanced tool call tracing

The new metadata fields improve observability:

  • "description" provides human-readable tool call summaries via get_parameterized_one_liner
  • "structured_tool_result" captures the complete tool response for detailed analysis

These enhancements will be valuable for debugging and monitoring tool execution.

tests/llm/test_workload_health.py (5)

9-9: LGTM: Appropriate import for span typing

The SpanType import is correctly used later in the test for proper span categorization.


27-27: LGTM: Property manager integration

The imports support the refactor to centralize test metadata handling through the property manager utilities.

Also applies to: 29-29


106-107: LGTM: Early property initialization

Setting initial properties early ensures test metadata is available even if the test fails during execution, improving error reporting and debugging.


139-139: LGTM: Appropriate span type categorization

Using SpanType.LLM for the "Holmes Run" span correctly categorizes this as an LLM operation.


179-180: LGTM: Consolidated test result updates

Using update_test_results consolidates property updates into a single function call, improving consistency and maintainability across test files.

tests/llm/conftest.py (6)

5-26: LGTM: Comprehensive import refactor supports modularity

The updated imports reflect the successful extraction of functionality into dedicated utility modules:

  • pytest_shared_session_scope for coordinating setup/cleanup across workers
  • Dedicated reporting modules for GitHub and terminal output
  • Centralized mock management utilities
  • Test result handling utilities

This modularization improves maintainability and code organization.


47-62: LGTM: Well-structured mock configuration fixture

The fixture correctly handles option retrieval with safe defaults and implements clear logic for determining mock mode based on environment variables and CLI options. The hierarchical decision making (regenerate-all implies generate, environment overrides) is intuitive.


68-152: LGTM: Sophisticated infrastructure coordination fixture

This fixture effectively solves the complex problem of coordinating setup/cleanup across pytest-xdist workers. Key strengths:

  • Proper handling of different execution modes (collect-only, mock, live)
  • Respects CLI flags for skipping setup/cleanup phases
  • Uses pytest_shared_session_scope correctly with the two-yield pattern
  • Comprehensive logic for first-worker setup and last-worker cleanup
  • Clear logging for debugging coordination issues

This addresses a real challenge in distributed test execution and follows the established patterns for the shared session scope library.


201-201: LGTM: Typo fix addressed

The fixture name has been corrected from "llm_availablity_check" to "llm_availability_check", addressing the previous review comment.


328-332: LGTM: Excellent UX improvement with clickable links

The ANSI escape code implementation for clickable terminal links is a thoughtful enhancement. The escape sequence format \033]8;;URL\033\\TEXT\033]8;;\033\\ is correctly implemented and will provide better user experience in modern terminals that support it.


335-367: LGTM: Clean refactor with proper xdist handling

The refactor successfully:

  • Delegates reporting responsibilities to dedicated modules (separation of concerns)
  • Handles xdist worker coordination to prevent duplicate output
  • Simplifies the main function by removing complex inline logic

This design is much more maintainable and testable than having all reporting logic inline.

@github-actions

Copy link
Copy Markdown
Contributor

Results of HolmesGPT evals

  • ask_holmes: 21/72 test cases were successful, 0 regressions, 44 mock failures
  • investigate: 3/16 test cases were successful, 0 regressions, 13 mock failures
Test suite Test case Status
ask 01_how_many_pods ✅
ask 02_what_is_wrong_with_pod 🔧
ask 03_what_is_the_command_to_port_forward 🔧
ask 04_related_k8s_events 🔧
ask 05_image_version 🔧
ask 06_explain_issue 🔧
ask 07_high_latency 🔧
ask 08_sock_shop_frontend ⚠️
ask 09_crashpod 🔧
ask 10_image_pull_backoff 🔧
ask 11_init_containers 🔧
ask 12_job_crashing 🔧
ask 13_pending_node_selector 🔧
ask 14_pending_resources 🔧
ask 15_failed_readiness_probe 🔧
ask 16_failed_no_toolset_found ✅
ask 17_oom_kill 🔧
ask 18_crash_looping_v2 🔧
ask 19_detect_missing_app_details 🔧
ask 20_long_log_file_search ✅
ask 21_job_fail_curl_no_svc_account ⚠️
ask 22_high_latency_dbi_down 🔧
ask 23_app_error_in_current_logs 🔧
ask 24_misconfigured_pvc ✅
ask 25_misconfigured_ingress_class 🔧
ask 26_multi_container_logs 🔧
ask 27_permissions_error_no_helm_tools 🔧
ask 28_permissions_error_helm_tools_enabled 🔧
ask 29_events_from_alert_manager 🔧
ask 30_basic_promql_graph_cluster_memory 🔧
ask 31_basic_promql_graph_pod_memory 🔧
ask 32_basic_promql_graph_pod_cpu 🔧
ask 33_http_latency_graph 🔧
ask 34_memory_graph 🔧
ask 35_tempo 🔧
ask 36_argocd_find_resource 🔧
ask 37_argocd_wrong_namespace 🔧
ask 38_rabbitmq_split_head 🔧
ask 39_failed_toolset ✅
ask 40_disabled_toolset ✅
ask 41_setup_argo ✅
ask 42_dns_issues_result_all_tools 🔧
ask 42_dns_issues_result_new_tools 🔧
ask 42_dns_issues_result_new_tools_no_runbook 🔧
ask 42_dns_issues_result_old_tools 🔧
ask 42_dns_issues_steps_new_all_tools 🔧
ask 42_dns_issues_steps_new_tools 🔧
ask 42_dns_issues_steps_old_tools 🔧
ask 43_current_datetime_from_prompt ✅
ask 43_slack_deployment_logs 🔧
ask 44_slack_statefulset_logs 🔧
ask 45_fetch_deployment_logs_simple ✅
ask 46_job_crashing_no_longer_exists 🔧
ask 47_truncated_logs_context_window ⚠️
ask 48_logs_since_thursday 🔧
ask 49_logs_since_last_week 🔧
ask 50_logs_since_specific_date 🔧
ask 51_logs_summarize_errors ✅
ask 52_logs_login_issues ⚠️
ask 53_logs_find_term ✅
ask 54_not_truncated_when_getting_pods ✅
ask 55_kafka_runbook 🔧
ask 56_kafka_runbook_no_tool ✅
ask 57_wrong_namespace ⚠️
ask 58_counting_pods_by_status ⚠️
ask 59_label_based_counting ✅
ask 60_time_based_filtering 🔧
ask 61_exact_match_counting ✅
ask 62_fetch_error_logs_with_errors ✅
ask 63_fetch_error_logs_no_errors ✅
ask 64_keda_vs_hpa_confusion ⚠️
ask 65_health_check_followup 🔧
investigate 01_oom_kill 🔧
investigate 02_crashloop_backoff 🔧
investigate 03_cpu_throttling ✅
investigate 04_image_pull_backoff 🔧
investigate 05_crashpod ✅
investigate 06_job_failure 🔧
investigate 07_job_syntax_error ✅
investigate 08_memory_pressure 🔧
investigate 09_high_latency 🔧
investigate 10_KubeDeploymentReplicasMismatch 🔧
investigate 11_KubePodCrashLooping 🔧
investigate 12_KubePodNotReady 🔧
investigate 13_Watchdog 🔧
investigate 14_tempo 🔧
investigate 15_dns_resolution 🔧
investigate 16_dns_resolution_no_tool 🔧

Legend

  • ✅ the test was successful
  • ⚠️ the test failed but is known to be flaky or known to fail
  • 🔧 the test failed due to mock data issues (not a code regression)
  • ❌ the test failed and should be fixed before merging the PR

@aantn
aantn merged commit 283b605 into master Jul 24, 2025
@aantn
aantn deleted the session-scoped-setup-v3 branch July 24, 2025 11:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants