ROB-1303 unified logging api loki - #430
Conversation
WalkthroughThis update introduces a unified, strongly-typed logging API for Kubernetes and Grafana Loki log retrieval, replacing legacy YAML-based and loosely-typed tools. It implements new toolset classes, prompt templates, and test suites for log fetching, with improved time filtering and error handling. Numerous test fixtures and cases are updated or removed to align with the new logging interface and tool naming conventions. Changes
Sequence Diagram(s)sequenceDiagram
participant User
participant PromptEngine
participant UnifiedLogToolset
participant KubernetesAPI
participant LokiAPI
User->>PromptEngine: Requests logs for a pod
PromptEngine->>UnifiedLogToolset: fetch_pod_logs(params)
alt Kubernetes logs
UnifiedLogToolset->>KubernetesAPI: Fetch pod logs (with filtering, time window)
KubernetesAPI-->>UnifiedLogToolset: Log lines or error
else Loki logs
UnifiedLogToolset->>LokiAPI: Query logs by label (with match, time window)
LokiAPI-->>UnifiedLogToolset: Log lines or error
end
UnifiedLogToolset-->>PromptEngine: StructuredToolResult (logs, status)
PromptEngine-->>User: Rendered logs or error message
sequenceDiagram
participant TestRunner
participant MockToolsets
participant UnifiedLogToolset
participant Braintrust
TestRunner->>Braintrust: Start evaluation span
TestRunner->>MockToolsets: Initialize with parent_span
TestRunner->>UnifiedLogToolset: fetch_pod_logs(params, parent_span)
UnifiedLogToolset-->>MockToolsets: Save/return logs (mocked or real)
TestRunner->>Braintrust: Evaluate correctness/context/sections (with parent_span)
Braintrust-->>TestRunner: Evaluation results
TestRunner->>Braintrust: End evaluation span
Note ⚡️ AI Code Reviews for VS Code, Cursor, WindsurfCodeRabbit now has a plugin for VS Code, Cursor and Windsurf. This brings AI code reviews directly in the code editor. Each commit is reviewed immediately, finding bugs before the PR is raised. Seamless context handoff to your AI code agent ensures that you can easily incorporate review feedback. Note ⚡️ Faster reviews with cachingCodeRabbit now supports caching for code and dependencies, helping speed up reviews. This means quicker feedback, reduced wait times, and a smoother review experience overall. Cached data is encrypted and stored securely. This feature will be automatically enabled for all accounts on May 16th. To opt out, configure ✨ Finishing Touches
🪧 TipsChatThere are 3 ways to chat with CodeRabbit:
SupportNeed help? Create a ticket on our support page for assistance with any issues or questions. Note: Be mindful of the bot's finite context window. It's strongly recommended to break down tasks such as reading entire modules into smaller chunks. For a focused discussion, use review comments to chat about specific files and their changes, instead of using the PR comments. CodeRabbit Commands (Invoked using PR comments)
Other keywords and placeholders
CodeRabbit Configuration File (
|
There was a problem hiding this comment.
Caution
Inline review comments failed to post. This is likely due to GitHub's limits when posting large numbers of comments. If you are seeing this consistently it is likely a permissions issue. Please check "Moderation" -> "Code review limits" under your organization settings.
Actionable comments posted: 17
🧹 Nitpick comments (23)
tests/llm/fixtures/test_ask_holmes/35_tempo/kubectl_logs_backend.txt (1)
1-1: Ensure consistency of the new tool name across fixtures and prompts.The
tool_namehas been updated tofetch_pod_logsto match the new unified logging API. Please verify that:
- All other test fixtures use
fetch_pod_logsinstead of legacy names (kubectl_logs,kubectl_logs_all_containers, etc.).- Prompt templates and examples in documentation/reference files also reference
fetch_pod_logs.- There are no remaining references to the old tool names in CI configurations or code comments.
tests/llm/fixtures/test_ask_holmes/25_misconfigured_ingress_class/test_case.yaml (1)
3-4: Ensure consistent CLI flag formatting
Consider adding a space between-fand the file path for clarity and to match commonkubectlusage (-f ./ingress_with_class.yaml).tests/llm/fixtures/test_ask_holmes/26_multi_container_logs/kubectl_logs_filter.txt (1)
2-5: Define expectedstdoutfixture content
Currently,stdout:andstderr:are empty, which may not fully exercise the log filtering logic. Consider adding representative log entries containing the"render time"keyword (and non-matching entries) to validate that filtering works as intended.holmes/plugins/prompts/_fetch_logs.jinja2 (1)
1-43: Add file documentation for better maintainability.While the template is well-structured, consider adding a comment at the top of the file explaining its purpose and how it works with different logging toolsets for future maintainers.
+{# + This template provides logging instructions based on available toolsets. + It checks for various logging providers (Loki, Coralogix, Kubernetes, OpenSearch) + and includes appropriate instructions for the first available/enabled one. + + For the unified logging API (k8s_base_ts), it includes the _default_log_prompt.jinja2 + template which provides standardized instructions across logging providers. +#} {%- set loki_ts = toolsets | selectattr("name", "equalto", "grafana/loki") | first -%}holmes/plugins/prompts/_default_log_prompt.jinja2 (1)
1-12: Improve bullet formatting and clarify time parameters.
- Use consistent single‐asterisk (
*) bullets for top‐level and two‐space indents for nested items.- Clarify that
start_timeshould be computed as the issue’sstarts_atminus 300 seconds, andend_timeset to the issue’sstarts_at.Proposed diff:
* Use the tool `fetch_pod_logs` to access an application's logs * Prior to fetching logs, ensure the pod exists using kubectl tools * If you find no logs, double check that the namespace and pod names are exact. Use kubectl tools to find the right resource names and pod name. * If you are not given the pod's namespace, look for existing pods using kubectl tools and infer the namespace that way * If you are not given the pod's exact name, or only have an application or deployment name, look for related pods using kubectl commands. Ask the user if you can't infer the pod logs. * Do fetch application logs yourself and do not ask users to do so * If you have an issue id or finding id, use `fetch_finding_by_id` as it contains time metadata about the issue (`starts_at`, `updated_at`, and `ends_at`). * Then, set `start_time` to `issue.starts_at - 300` (five minutes before the issue’s start time) and `end_time` to the issue’s `starts_at` when calling `fetch_pod_logs`. * If there are too many logs or not enough, narrow or widen the timestamps. * If looking for a specific keyword, use the `filter` argument. * If you are not provided with time information, ignore the `start_time` and `end_time`; `fetch_pod_logs` will default to the latest logs.</blockquote></details> <details> <summary>tests/llm/test_investigate.py (3)</summary><blockquote> `32-37`: **Unused stored attribute – possible dead code** `MockConfig.__init__` stores `self._parent_span` but the attribute is never referenced inside the class. If the only consumer is `MockToolsets`, pass the span directly and drop the field to prevent stale state. ```diff - def __init__(self, test_case: InvestigateTestCase, parent_span: Span): + def __init__(self, test_case: InvestigateTestCase, parent_span: Span): super().__init__() self._test_case = test_case - self._parent_span = parent_span + # Pass through; no need to persist
38-43: Forwardingparent_spanbut not documenting the new constructor
MockToolsetsnow relies on a newparent_spankw-arg, yet its constructor’s docstring or type signature (in the implementation file) was not updated in this PR. Future maintainers will be confused by the silent dependency.
Please updatetests/llm/utils/mock_toolset.pywith an explicitparent_span: Span | None = Noneparameter and docstring.
125-149: Evaluation helpers swallow classifier exceptions
All three evaluation helpers (evaluate_correctness,evaluate_sections,evaluate_context_usage) are called without error handling. If any of them throws, you lose diagnostic information from later assertions and the Braintrust span ends without scores. Consider catchingExceptionaround each helper, logging inside the span, and recording a score of0.0to avoid hard failures unrelated to the LLM output.holmes/plugins/toolsets/__init__.py (2)
80-82: Lazy instantiation suggestion
Even after deferring the import, the toolset object is still constructed at startup. If the cluster has no kube-config,KubernetesLogsToolset()may attempt to connect and raise. Consider wrapping the append intry/except Exception as exc:and logging a warning so the rest of Holmes continues to load.
90-97: YAML/py duplication guard looks correct – minor readability nit
Good catch skippingkubernetes_logs.yamlwhen the new Python toolset is active. A tiny readability tweak:- if filename == "kubernetes_logs.yaml" and not USE_LEGACY_KUBERNETES_LOGS: + if not USE_LEGACY_KUBERNETES_LOGS and filename == "kubernetes_logs.yaml":Placing the inexpensive boolean first avoids the string comparison in the common case.
tests/plugins/prompt/test_fetch_logs.py (2)
8-12: Avoid brittle relative path constructionUsing
os.path.join(THIS_DIR, "../../../...")assumes the test file will always live exactly three directories below the YAML file.
A small package-layout change will break the tests.-THIS_DIR = os.path.abspath(os.path.dirname(__file__)) -KUBERNETES_YAML_TOOLSET_PATH = os.path.join( - THIS_DIR, "../../../holmes/plugins/toolsets/kubernetes_logs.yaml" -) +from importlib.resources import files +# Locate the YAML file relative to the installed package instead of the test file. +KUBERNETES_YAML_TOOLSET_PATH = ( + files("holmes.plugins.toolsets") + / "kubernetes_logs.yaml" +)
20-28: Remove noisyprint()calls in testsThe three
print(f"** PROMPT: ...")statements clutter the test output and slow down CI runners.
If the output is really needed for debugging, use the-spytest flag orcapsysfixture instead.- print(f"** PROMPT:\n{prompt}")Also applies to: 30-38, 41-49
tests/plugins/toolsets/kubernetes/test_kubernetes_logs.py (1)
14-18: Patch the class method instead of the instance
patch.object(self.toolset, "_initialize_client")stubs only the instance copy created insetUp.
If the test ever reinstantiates the toolset, or other tests import the class directly, the real client
initialisation could leak through.-patcher = patch.object(self.toolset, "_initialize_client") +patcher = patch.object(KubernetesLogsToolset, "_initialize_client")tests/plugins/toolsets/kubernetes/test_kubernetes_logs_filter_by_timestamp.py (1)
127-129: Suppress gratuitous debug printsThe
print()calls inside parameterised tests clutter CI logs and slow down execution.
Rely on pytest’s assertion introspection or usecapsyswhen the diff is really needed.- print(f"EXPECTED:\n{expected_output}") - print(f"ACTUAL:\n{result}")Also applies to: 154-156
tests/llm/test_ask_holmes.py (2)
73-84: Avoid shadowing the built-ininputAssigning to a variable named
inputhides Python’s built-in function within the remainder of the scope and can lead to confusing errors during debugging.- input = test_case.user_prompt + user_input = test_case.user_prompt
82-93: Missing score for mandatory correctness evaluationIf
evaluate_correctnessraises or returnsNone, the subsequent access tocorrectness_eval.scorewill raise.
Wrap the call with a try/except or validate the return type to ensure the test always records a numeric score.tests/plugins/toolsets/grafana/test_grafana_loki.py (2)
108-112: Fragile header-skipping logicThe test discards the first two lines assuming they are always “link” and “query”.
Any change in Loki’s output format (extra blank line, order swap, etc.) will break the test.Prefer an explicit filter:
content_lines = [l for l in result.data.splitlines() if l and not l.startswith(("link:", "query:"))]This keeps the intent clear and resilient.
125-129: Brittle assertion on exact line countThe test enforces
len(result.data.split("\n")) == 1, which can fail if Loki adds a trailing newline or extra metadata.
A safer check is to assert that at least one log line is returned and that every line contains the expected term.lines = [l for l in result.data.splitlines() if l] assert lines, "No log lines returned" for line in lines: assert TEST_SEARCH_TERM in lineholmes/plugins/toolsets/kubernetes_logs.py (2)
153-165: Single-container pods skip explicit container queryingWhen
containershas length 1, the branch guarded bylen(containers) > 1is skipped, and the fallback fetches logs without the
containerparameter.
If the pod previously had a different container name, the returned logs may be incomplete.You can simplify and unify the path:
for container_name in containers or [None]: ...and always prefix only when
container_nameis notNoneand there are multiple containers.
222-229: Ruff hint – collapsible error branchesThe two early-return branches on 400/404 status codes can be merged for brevity:
- except ApiException as e: - if e.status == 400 and "previous terminated container" in str(e).lower(): - return [] - elif e.status == 404: - return [] + except ApiException as e: + if e.status == 404 or ( + e.status == 400 and "previous terminated container" in str(e).lower() + ): + return []Not critical, but trims five lines.
🧰 Tools
🪛 Ruff (0.11.9)
222-225: Combine
ifbranches using logicaloroperatorCombine
ifbranches(SIM114)
holmes/plugins/toolsets/logging_api.py (1)
146-153: Handling of negativestart_timevalues can mis-detect strings with whitespace or “s” suffix
start_time.startswith("-") and start_time[1:].isdigit()returnsFalsefor valid inputs like"-3600 "or"-3600s".
A more robust check would attemptint(start_time)inside atry/except.- if isinstance(start_time, int) or ( - isinstance(start_time, str) - and start_time.startswith("-") - and start_time[1:].isdigit() - ): + if isinstance(start_time, (int, str)): + try: + seconds_before = abs(int(start_time)) + is_relative = True + except ValueError: + is_relative = False + + if is_relative:This keeps relative time parsing flexible without sacrificing validation.
holmes/plugins/toolsets/grafana/toolset_grafana_loki.py (2)
50-53: Clarify error message – toolset is Loki, not OpenSearchThe prerequisite failure message still refers to “OpenSearch”, which will mislead users configuring the Grafana / Loki backend and make troubleshooting harder.
- if not config: - return False, "Missing OpenSearch configuration. Check your config." + if not config: + return False, "Missing Grafana Loki configuration. Check your config."
91-98: Return path is fine butlogs.sortmutates caller-visible listBecause
query_loki_logs_by_labelalready returns a list that might be reused by the caller/tests, sorting it in-place couples your implementation to the function’s return value. A defensive copy keeps side-effects local and avoids surprises.- logs.sort(key=lambda x: x["timestamp"]) + logs = sorted(logs, key=lambda x: x["timestamp"])
🛑 Comments failed to post (17)
tests/llm/fixtures/test_investigate/03_cpu_throttling/kubectl_previous_logs.txt (1)
1-1:
⚠️ Potential issueMissing
previousflag in fetch_pod_logs invocationThe fixture is for previous logs, but the
fetch_pod_logsinvocation lacks apreviousparameter. Without it, only current logs will be fetched, breaking this scenario.Suggested diff:
-{"toolset_name":"kubernetes/logs","tool_name":"fetch_pod_logs","match_params":{"pod_name":"frontend-service","namespace":"default"}} +{"toolset_name":"kubernetes/logs","tool_name":"fetch_pod_logs","match_params":{"pod_name":"frontend-service","namespace":"default","previous": true}}📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.{"toolset_name":"kubernetes/logs","tool_name":"fetch_pod_logs","match_params":{"pod_name":"frontend-service","namespace":"default","previous": true}}🤖 Prompt for AI Agents
In tests/llm/fixtures/test_investigate/03_cpu_throttling/kubectl_previous_logs.txt at line 1, the fetch_pod_logs invocation is missing the 'previous' flag, which is necessary to fetch previous logs as intended by this fixture. Add "previous": true to the match_params object to ensure the logs fetched are from the previous container instance.tests/llm/fixtures/test_investigate/07_job_syntax_error/kubectl_logs_all_containers.txt (1)
1-1:
⚠️ Potential issueMissing all_containers flag for multi-container logs
This fixture targets logs from all containers, but the
fetch_pod_logsinvocation lacks anall_containersparameter (or equivalent). The new toolset likely requires explicit instruction to fetch logs across every container.Suggested diff:
-{"toolset_name":"kubernetes/logs","tool_name":"fetch_pod_logs","match_params":{"pod_name":"*","namespace":"default"}} +{"toolset_name":"kubernetes/logs","tool_name":"fetch_pod_logs","match_params":{"pod_name":"*","namespace":"default","all_containers": true}}📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.{"toolset_name":"kubernetes/logs","tool_name":"fetch_pod_logs","match_params":{"pod_name":"*","namespace":"default","all_containers": true}}🤖 Prompt for AI Agents
In tests/llm/fixtures/test_investigate/07_job_syntax_error/kubectl_logs_all_containers.txt at line 1, the JSON object for fetch_pod_logs is missing the all_containers flag needed to fetch logs from all containers in a pod. Add an "all_containers": true key-value pair inside the match_params object to explicitly instruct the toolset to retrieve logs from every container.tests/llm/fixtures/test_investigate/04_image_pull_backoff/kubectl_previous_logs.txt (1)
1-1:
⚠️ Potential issueMissing
previousflag in fetch_pod_logs invocationThis scenario fetches previous logs, but the invocation is missing the
previousparameter. It must be added to retrieve terminated container logs correctly.Suggested diff:
-{"toolset_name":"kubernetes/logs","tool_name":"fetch_pod_logs","match_params":{"pod_name":"customer-relations-webapp-5d98ffcfd-tz4nc","namespace":"default"}} +{"toolset_name":"kubernetes/logs","tool_name":"fetch_pod_logs","match_params":{"pod_name":"customer-relations-webapp-5d98ffcfd-tz4nc","namespace":"default","previous": true}}📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.{"toolset_name":"kubernetes/logs","tool_name":"fetch_pod_logs","match_params":{"pod_name":"customer-relations-webapp-5d98ffcfd-tz4nc","namespace":"default","previous": true}}🤖 Prompt for AI Agents
In tests/llm/fixtures/test_investigate/04_image_pull_backoff/kubectl_previous_logs.txt at line 1, the fetch_pod_logs invocation is missing the 'previous' flag needed to retrieve logs from terminated containers. Add the 'previous' parameter set to true within the match_params object to ensure the scenario fetches the previous logs correctly.tests/llm/fixtures/test_ask_holmes/08_sock_shop_frontend/test_case.yaml (1)
20-20:
⚠️ Potential issueFix YAML indentation for
generate_mocks
generate_mocks: Falseis not indented under theevaluationblock, so it will be treated as a top-level key. It should be nested alongsidecorrectness:evaluation: correctness: 1 - generate_mocks: False + generate_mocks: False📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.evaluation: correctness: 1 generate_mocks: False🤖 Prompt for AI Agents
In tests/llm/fixtures/test_ask_holmes/08_sock_shop_frontend/test_case.yaml at line 20, the key `generate_mocks: False` is incorrectly placed at the top level. Move it to be indented under the `evaluation` block, aligned with the `correctness` key, to ensure it is properly nested within the evaluation configuration.tests/llm/fixtures/test_ask_holmes/20_long_log_file_search/test_case.yaml (1)
8-8:
⚠️ Potential issueCorrect indentation for
generate_mocks
generate_mocks: Trueis currently at the root level; it must be nested underevaluationto be recognized:evaluation: correctness: 1 - generate_mocks: True + generate_mocks: True🤖 Prompt for AI Agents
In tests/llm/fixtures/test_ask_holmes/20_long_log_file_search/test_case.yaml at line 8, the key `generate_mocks: True` is incorrectly placed at the root level. Move this line so that it is indented under the `evaluation` section to ensure it is properly recognized as part of the evaluation configuration..github/workflows/llm-evaluation.yaml (1)
20-24: 💡 Verification agent
🧩 Analysis chain
Pin Python version consistently
Set up Python to use version3.12—ensure this version is supported across all workflows and aligns with your.github/workflows/build-binaries-and-brew.yaml.Run this script to verify Python versions in other workflows:
🏁 Script executed:
#!/bin/bash # Verify Python versions across workflows rg --no-heading -n 'python-version:' .github/workflowsLength of output: 400
Align Python versions across workflows
Thellm-evaluation.yamlworkflow pins Python 3.12, butbuild-binaries-and-brew.yamlis still on 3.11. To avoid surprises, choose one version and use it everywhere. Since your matrix inbuild-and-test.yamlalready includes 3.12, update the brew workflow to match:• In
.github/workflows/build-binaries-and-brew.yaml(around line 24):- python-version: '3.11' + python-version: 3.12• Verify all other workflows (e.g. CI matrix) support the chosen version.
Committable suggestion skipped: line range outside the PR's diff.
🤖 Prompt for AI Agents
In .github/workflows/llm-evaluation.yaml at lines 20 to 24, Python version is pinned to 3.12, but other workflows like build-binaries-and-brew.yaml use 3.11. To maintain consistency and avoid version conflicts, update the Python version in .github/workflows/build-binaries-and-brew.yaml around line 24 to 3.12, and verify all other workflows also support Python 3.12 to ensure uniformity across the project.tests/llm/test_investigate.py (1)
83-90:
⚠️ Potential issueMissed
finally– evaluation span may leak on early failure
Ifinvestigate_issues, an assertion, or any evaluation raises,bt_helper.end_evaluation()is never executed, leaving the span open in Braintrust. Wrap the body intry/finally:- eval_span = bt_helper.start_evaluation(experiment_name, name=test_case.id) - ... - if bt_helper and eval_span: - bt_helper.end_evaluation(...) + eval_span = bt_helper.start_evaluation(experiment_name, name=test_case.id) + try: + ... + finally: + if bt_helper and eval_span: + bt_helper.end_evaluation(...)Committable suggestion skipped: line range outside the PR's diff.
🤖 Prompt for AI Agents
In tests/llm/test_investigate.py around lines 83 to 90, the evaluation span started by bt_helper.start_evaluation is not properly closed if an exception occurs, potentially leaking the span. To fix this, wrap the code following the start_evaluation call in a try block and ensure bt_helper.end_evaluation() is called in a finally block to guarantee the evaluation span is always ended regardless of errors.holmes/plugins/toolsets/__init__.py (1)
6-13:
⚠️ Potential issueEager import defeats optional-dependency flag
KubernetesLogsToolset(and its heavykubernetesPyPI dependency) is imported at module import time even whenUSE_LEGACY_KUBERNETES_LOGSis true. This means:
- Environments that wish to stay on the legacy path must still install the Kubernetes client.
- Any import error inside
kubernetes_logs.pycrashes Holmes startup before the flag is checked.Move the import inside the
if not USE_LEGACY_KUBERNETES_LOGS:block:if not USE_LEGACY_KUBERNETES_LOGS: from holmes.plugins.toolsets.kubernetes_logs import KubernetesLogsToolsetand adjust the subsequent references.
🤖 Prompt for AI Agents
In holmes/plugins/toolsets/__init__.py around lines 6 to 13, the KubernetesLogsToolset is imported eagerly regardless of the USE_LEGACY_KUBERNETES_LOGS flag, causing unnecessary dependency loading and potential import errors. Move the import of KubernetesLogsToolset inside an if not USE_LEGACY_KUBERNETES_LOGS: block to defer the import until the flag is checked. Then update any references to KubernetesLogsToolset to ensure they only occur when the import is valid.tests/plugins/prompt/test_fetch_logs.py (1)
20-23: 🛠️ Refactor suggestion
Do not rely on private attributes for state manipulation
The tests mutate the “private”
_statusattribute of each toolset:toolset._status = ToolsetStatusEnum.ENABLEDProvide an official helper (e.g.
toolset.set_status(ToolsetStatusEnum.ENABLED)) or expose a public property to avoid tight coupling to the implementation.Also applies to: 31-34, 42-45
🤖 Prompt for AI Agents
In tests/plugins/prompt/test_fetch_logs.py around lines 20 to 23 (and similarly at lines 31-34 and 42-45), the test code directly modifies the private attribute _status of toolset instances, which creates tight coupling to the implementation. To fix this, replace direct assignments to _status with calls to a public method like set_status or use a public property setter designed for changing the toolset's status. If such a method or property does not exist, add one to the Toolset class to safely update the status without accessing private attributes.tests/plugins/toolsets/kubernetes/test_kubernetes_logs.py (1)
32-40:
⚠️ Potential issueDecorator-parameter order mismatch breaks all patched tests
unittest.mock.patchinjects mocks in the same order the decorators are applied (top → bottom).
Here@patch("kubernetes.client.CoreV1Api")is applied first, so the first function argument should
bemock_core_v1_api, followed bymock_config. Currently they are reversed, causing the “config” mock to
be treated as an API instance and vice-versa, which will explode on.return_value.-@patch("kubernetes.client.CoreV1Api") -@patch("kubernetes.config") -def test_default_log_formatting(self, mock_config, mock_api): +@patch("kubernetes.client.CoreV1Api") # top decorator -> first arg +@patch("kubernetes.config") # bottom decorator -> second arg +def test_default_log_formatting(self, mock_api, mock_config):Apply the same fix to all test functions patched in this module.
Also applies to: 66-74, 112-120, 162-170, 194-202
🤖 Prompt for AI Agents
In tests/plugins/toolsets/kubernetes/test_kubernetes_logs.py at lines 32 to 40 and similarly at lines 66-74, 112-120, 162-170, and 194-202, the order of parameters in the test functions does not match the order of the @patch decorators, causing mocks to be assigned incorrectly. Fix this by reversing the order of the function parameters so that the first parameter corresponds to the first (top) @patch decorator and the second parameter corresponds to the second decorator, ensuring mocks are injected in the correct order.tests/plugins/toolsets/grafana/test_grafana_loki.py (2)
88-95: 🛠️ Refactor suggestion
The “basic” query asserts on a hard-coded search term that is never requested
TEST_SEARCH_TERMis not supplied inFetchPodLogsParams, yet the assertion requires it to appear in the response.
If the pod stops emitting the word “WARNING”, this test will fail for reasons unrelated to correctness.Options:
• Passmatch=TEST_SEARCH_TERMso the API guarantees the presence.
• Or drop the assertion and only checkresult.status/error.🤖 Prompt for AI Agents
In tests/plugins/toolsets/grafana/test_grafana_loki.py around lines 88 to 95, the test asserts that TEST_SEARCH_TERM appears in the logs without requesting it in FetchPodLogsParams, which can cause flaky failures. Fix this by adding match=TEST_SEARCH_TERM to the FetchPodLogsParams call to ensure the logs contain the search term, or alternatively remove the assertion checking for TEST_SEARCH_TERM and only verify result.status and result.error.
60-65: 🛠️ Refactor suggestion
Avoid bypassing the toolset’s public configuration API and fix misleading doc-string
- The doc-string still mentions
OpenSearchLogsToolset, which can confuse future maintainers.- Assigning
toolset.config = loki_config.model_dump()mutates a (likely) public attribute directly.
IfGrafanaLokiToolsetlater adds validation or side-effects inside a setter /configure()helper, this test will silently skip them.-"""Create an OpenSearchLogsToolset with the test configuration""" +"""Create a GrafanaLokiToolset instance with the test configuration""" - -toolset = GrafanaLokiToolset() -toolset.config = loki_config.model_dump() +# Prefer dedicated ctor/initialiser if available +# e.g. toolset = GrafanaLokiToolset(config=loki_config) +# or, if not, expose a helper: +# toolset.configure(loki_config) +toolset = GrafanaLokiToolset() +toolset.config = loki_config.model_dump() # TODO: replace direct mutationConsider changing the production class to expose an explicit
configure()or constructor parameter and update the test accordingly.📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.@pytest.fixture def loki_toolset(loki_config) -> GrafanaLokiToolset: """Create a GrafanaLokiToolset instance with the test configuration""" # Prefer a dedicated ctor/initializer if available: # e.g. toolset = GrafanaLokiToolset(config=loki_config) # or, if not, expose a helper: # toolset.configure(loki_config) toolset = GrafanaLokiToolset() toolset.config = loki_config.model_dump() # TODO: replace direct mutation toolset.check_prerequisites()🤖 Prompt for AI Agents
In tests/plugins/toolsets/grafana/test_grafana_loki.py around lines 60 to 65, update the doc-string to correctly reference GrafanaLokiToolset instead of OpenSearchLogsToolset to avoid confusion. Replace the direct assignment to toolset.config with a call to a proper configuration method like configure() or pass the configuration via the constructor if supported, ensuring any validation or side-effects in the class are executed. If such a method does not exist, add one to the production class and use it in the test to set the configuration properly.holmes/plugins/toolsets/kubernetes_logs.py (2)
86-98: 🛠️ Refactor suggestion
Potential duplication of log lines when combining
previousand current logs
fetch_pod_logs()concatenatesprevious=Trueandprevious=Falseresults without de-duplication.
When a container hasn’t rolled over, the same lines may appear in both calls, doubling the output and breakinglimitlogic.Consider deduplicating while preserving order:
all_logs = list(dict.fromkeys(all_logs)) # preserves first-seen orderor compare timestamps if available.
🤖 Prompt for AI Agents
In holmes/plugins/toolsets/kubernetes_logs.py around lines 86 to 98, the code concatenates logs fetched with previous=True and previous=False without removing duplicates, which can cause duplicated log lines and affect the limit logic. To fix this, after combining the two log lists, deduplicate the entries while preserving their order by converting the list to a dict and back to a list, or implement a timestamp-based comparison if timestamps are available, ensuring no duplicate log lines appear in the final all_logs list.
70-78:
⚠️ Potential issue
config.ConfigExceptionmay not exist – import the correct exception class
kubernetes.configdoes not exposeConfigExceptionat the module root in some client versions.
Accessingconfig.ConfigExceptioncan raiseAttributeError, causing_initialize_client()to mask the real configuration issue.-from kubernetes import client, config -from kubernetes.client.exceptions import ApiException +from kubernetes import client, config +from kubernetes.client.exceptions import ApiException +from kubernetes.config.config_exception import ConfigException ... - except config.ConfigException: + except ConfigException:This keeps the code compatible with both in-cluster and out-of-cluster setups.
📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.from kubernetes import client, config from kubernetes.client.exceptions import ApiException +from kubernetes.config.config_exception import ConfigException ... try: try: config.load_incluster_config() - except config.ConfigException: + except ConfigException: logging.debug( f"_initialize_client for {self.name} toolset falling back to loading kube_config" ) config.load_kube_config()🤖 Prompt for AI Agents
In holmes/plugins/toolsets/kubernetes_logs.py around lines 70 to 78, the code catches config.ConfigException which may not exist in some kubernetes client versions, causing an AttributeError. To fix this, import the correct exception class explicitly from kubernetes.config or its submodules and catch that instead of config.ConfigException, ensuring compatibility with different client versions and avoiding masking real configuration errors.tests/llm/utils/classifiers.py (1)
101-103:
⚠️ Potential issueBug: passing Python built-in
inputinstead of the real input string
evaluate_correctness()has noinputparameter, yet you passinput=inputto the classifier.
At runtime this resolves to the built-ininput()function, leading to meaningless prompts and unpredictable behaviour.-def evaluate_correctness( - expected_elements: list[str], output: Optional[str], parent_span: Span -): +def evaluate_correctness( + expected_elements: list[str], + output: Optional[str], + *, + input: Optional[str] = None, # new, mirrors other helpers + parent_span: Span, +): ... - correctness_eval = classifier( - input=input, output=output, expected=expected_elements_str - ) + correctness_eval = classifier( + input=input, + output=output, + expected=expected_elements_str, + )📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.def evaluate_correctness( expected_elements: list[str], output: Optional[str], *, input: Optional[str] = None, # new, mirrors other helpers parent_span: Span, ): # … existing implementation … correctness_eval = classifier( input=input, output=output, expected=expected_elements_str, ) # … rest of implementation …🤖 Prompt for AI Agents
In tests/llm/utils/classifiers.py around lines 101 to 103, the code incorrectly passes the built-in Python function `input` as a parameter to the classifier because `evaluate_correctness()` does not have an `input` argument. To fix this, identify the correct variable holding the input string intended for the classifier and pass that variable instead of `input=input`. Remove or replace the `input=input` argument with the actual input string variable to avoid passing the built-in function.tests/llm/utils/mock_toolset.py (2)
142-176: 🛠️ Refactor suggestion
Same exception-swallowing issue as above in
MockToolWrapper._invokeApply the same
else/finallyrestructuring here to ensure errors propagate correctly.🧰 Tools
🪛 Ruff (0.11.9)
152-155: Use ternary operator
result = mock.return_value if mock else self._unmocked_tool.invoke(params)instead ofif-else-blockReplace
if-else-block withresult = mock.return_value if mock else self._unmocked_tool.invoke(params)(SIM108)
🤖 Prompt for AI Agents
In tests/llm/utils/mock_toolset.py between lines 142 and 176, the _invoke method currently catches exceptions and logs them but then re-raises, which can obscure error propagation. Refactor the try-except-finally block by moving the span.end() call into a finally block and restructuring the else block so that the span.log call for successful results happens only if no exception occurs. This ensures exceptions propagate correctly without being swallowed while maintaining proper span lifecycle management.
118-119:
⚠️ Potential issueShared mutable default list leaks state between tests
mocks: List[ToolMock] = []is evaluated once at import time, so instances ofMockToolWrappershare the same list, causing cross-test pollution.-from typing import Any, Dict, List, Optional +from typing import Any, Dict, List, Optional +from pydantic import Field ... -class MockToolWrapper(Tool, BaseModel): - mocks: List[ToolMock] = [] +class MockToolWrapper(Tool, BaseModel): + mocks: List[ToolMock] = Field(default_factory=list)📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.from typing import Any, Dict, List, Optional from pydantic import Field class MockToolWrapper(Tool, BaseModel): mocks: List[ToolMock] = Field(default_factory=list)🤖 Prompt for AI Agents
In tests/llm/utils/mock_toolset.py at line 118, the declaration of mocks as a mutable default list causes shared state across test instances. To fix this, change mocks to be initialized inside the constructor or a method so that each instance gets its own new list, avoiding cross-test pollution.
Summary by CodeRabbit
New Features
Bug Fixes
Refactor
Tests
Chores