Skip to content
Merged
6 changes: 3 additions & 3 deletions holmes/plugins/toolsets/coralogix/toolset_coralogix_logs.py
Original file line number Diff line number Diff line change
Expand Up @@ -41,11 +41,11 @@ def __init__(self):
docs_url="https://docs.robusta.dev/master/configuration/holmesgpt/toolsets/coralogix_logs.html",
icon_url="https://avatars.githubusercontent.com/u/35295744?s=200&v=4",
prerequisites=[CallablePrerequisite(callable=self.prerequisites_callable)],
tools=[
PodLoggingTool(self),
],
tools=[], # Initialize with empty tools first
tags=[ToolsetTag.CORE],
)
# Now that parent is initialized and self.name exists, create the tool
self.tools = [PodLoggingTool(self)]

def get_example_config(self):
example_config = CoralogixConfig(
Expand Down
43 changes: 43 additions & 0 deletions holmes/plugins/toolsets/datadog/datadog_logs_instructions.jinja2
Original file line number Diff line number Diff line change
@@ -0,0 +1,43 @@
## Datadog Logs Tools Usage Guide

Before running logs queries:

** You are often (but not always) running in a kubernetes environment. So users might ask you questions about kubernetes workloads without explicitly stating their type.
** When getting ambiguous questions, use kubectl_find_resource to find the resource you are being asked about!
** Find the involved resource name and kind
** If you can't figure out what is the type of the resource, ask the user for more information and don't guess


### General guideline
- This toolset is used to read pod logs.
- Assume the pod should have logs. If logs not found, try to adjust the query

### CRITICAL: Pod Name Resolution Workflow

**When user provides an exact pod name** (e.g., `my-workload-5f9d8b7c4d-x2km9`):
- FIRST query Datadog directly with that pod name using appropriate tags
- Do NOT try to verify if the pod exists in Kubernetes first
- This allows querying historical pods that have been deleted/replaced

**When user provides a generic workload name** (e.g., "my-workload", "nginx", "telemetry-processor"):
- First use `kubectl_find_resource` to find actual pod names
- Example: `kubectl_find_resource` with "my-workload" → finds pods like "my-workload-8f8cdfxyz-c7zdr"
- Then use those specific pod names in Datadog queries
- Alternative: Use deployment-level tags when appropriate

**Why this matters:**
- Pod names in Datadog are the actual Kubernetes pod names (with random suffixes)
- Historical pods that no longer exist in the cluster can still have logs in Datadog
- Deployment/service names alone are NOT pod names (they need the suffix)

### Time Parameters
- Use RFC3339 format: `2023-03-01T10:30:00Z`
- Or relative seconds: `-3600` for 1 hour ago
- Defaults to 1 hour window if not specified

### Common Investigation Patterns

**For Pod/Container Metrics (MOST COMMON):**
1. User asks: "Show logs for my-workload"
2. Use `kubectl_find_resource` → find pod "my-workload-abc123-xyz"
3. Query Datadog for pod "my-workload-abc123-xyz" logs
Original file line number Diff line number Diff line change
Expand Up @@ -32,19 +32,22 @@ When investigating metrics-related issues:
- IMPORTANT: This toolset DOES NOT support promql queries.

### CRITICAL: Pod Name Resolution Workflow
When users ask for metrics about a deployment, service, or workload (e.g., "my-workload", "nginx-deployment"):

**ALWAYS follow this two-step process:**
1. **First**: Use `kubectl_find_resource` to find the actual pod names
- Example: `kubectl_find_resource` with "my-workload" → finds pods like "my-workload-8f8cdfxyz-c7zdr"
2. **Then**: Use those specific pod names in Datadog queries
- Correct: `container.cpu.usage{pod_name:my-workload-8f8cdfxyz-c7zdr}`
- WRONG: `container.cpu.usage{pod_name:my-workload}` ← This will return no data!
**When user provides an exact pod name** (e.g., `my-workload-5f9d8b7c4d-x2km9`):
- Query Datadog directly with that pod name using appropriate metrics and tags
- Do NOT try to verify if the pod exists in Kubernetes first
- This allows querying historical pods that have been deleted/replaced

**When user provides a generic workload name** (e.g., "my-workload", "nginx", "telemetry-processor"):
- First use `kubectl_find_resource` to find actual pod names
- Example: `kubectl_find_resource` with "my-workload" → finds pods like "my-workload-8f8cdfxyz-c7zdr"
- Then use those specific pod names in Datadog queries
- Alternative: Use deployment-level tags when appropriate

**Why this matters:**
- Pod names in Datadog are the actual Kubernetes pod names (with random suffixes)
- Deployment/service names are NOT pod names
- Using deployment names as pod_name filters will always return empty results
- Historical pods that no longer exist in the cluster can still have metrics in Datadog
- Deployment/service names alone are NOT pod names (they need the suffix)

### Time Parameters
- Use RFC3339 format: `2023-03-01T10:30:00Z`
Expand Down
23 changes: 17 additions & 6 deletions holmes/plugins/toolsets/datadog/toolset_datadog_logs.py
Original file line number Diff line number Diff line change
@@ -1,3 +1,4 @@
import os
from enum import Enum
import json
import logging
Expand Down Expand Up @@ -141,22 +142,25 @@ class DatadogLogsToolset(BasePodLoggingToolset):

@property
def supported_capabilities(self) -> Set[LoggingCapability]:
"""Datadog logs API only supports substring matching, no exclude filter"""
return set() # No regex support, no exclude filter
"""Datadog logs API supports historical data and substring matching"""
return {
LoggingCapability.HISTORICAL_DATA
} # No regex support, no exclude filter, but supports historical data

def __init__(self):
super().__init__(
name="datadog/logs",
description="Toolset for interacting with Datadog to fetch logs",
description="Toolset for fetching logs from Datadog, including historical data for pods no longer in the cluster",
docs_url="https://docs.datadoghq.com/api/latest/logs/",
icon_url="https://imgix.datadoghq.com//img/about/presskit/DDlogo.jpg",
prerequisites=[CallablePrerequisite(callable=self.prerequisites_callable)],
tools=[
PodLoggingTool(self),
],
tools=[], # Initialize with empty tools first
experimental=True,
tags=[ToolsetTag.CORE],
)
# Now that parent is initialized and self.name exists, create the tool
self.tools = [PodLoggingTool(self)]
self._reload_instructions()

Comment thread
arikalon1 marked this conversation as resolved.
def logger_name(self) -> str:
return "DataDog"
Expand Down Expand Up @@ -272,3 +276,10 @@ def get_example_config(self) -> Dict[str, Any]:
"dd_app_key": "your-datadog-application-key",
"site_api_url": "https://api.datadoghq.com",
}

def _reload_instructions(self):
"""Load Datadog logs specific troubleshooting instructions."""
template_file_path = os.path.abspath(
os.path.join(os.path.dirname(__file__), "datadog_logs_instructions.jinja2")
)
self._load_llm_instructions(jinja_template=f"file://{template_file_path}")
6 changes: 3 additions & 3 deletions holmes/plugins/toolsets/datadog/toolset_datadog_metrics.py
Original file line number Diff line number Diff line change
Expand Up @@ -54,7 +54,7 @@ class ListActiveMetrics(BaseDatadogMetricsTool):
def __init__(self, toolset: "DatadogMetricsToolset"):
super().__init__(
name="list_active_datadog_metrics",
description=f"List active metrics from the last {ACTIVE_METRICS_DEFAULT_LOOK_BACK_HOURS} hours. This includes metrics that have actively reported data points.",
description=f"List active metrics from Datadog for the last {ACTIVE_METRICS_DEFAULT_LOOK_BACK_HOURS} hours. This includes metrics that have actively reported data points, including from pods no longer in the cluster.",
parameters={
"from_time": ToolParameter(
description=f"Start time for listing metrics. Can be an RFC3339 formatted datetime (e.g. '2023-03-01T10:30:00Z') or a negative integer for relative seconds from now (e.g. -86400 for 24 hours ago). Defaults to {ACTIVE_METRICS_DEFAULT_LOOK_BACK_HOURS} hours ago",
Expand Down Expand Up @@ -182,7 +182,7 @@ class QueryMetrics(BaseDatadogMetricsTool):
def __init__(self, toolset: "DatadogMetricsToolset"):
super().__init__(
name="query_datadog_metrics",
description="Query timeseries data for a specific metric",
description="Query timeseries data from Datadog for a specific metric, including historical data for pods no longer in the cluster",
parameters={
"query": ToolParameter(
description="The metric query string (e.g., 'system.cpu.user{host:myhost}')",
Expand Down Expand Up @@ -562,7 +562,7 @@ class DatadogMetricsToolset(Toolset):
def __init__(self):
super().__init__(
name="datadog/metrics",
description="Toolset for interacting with Datadog to fetch metrics and metadata",
description="Toolset for fetching metrics and metadata from Datadog, including historical data for pods no longer in the cluster",
docs_url="https://docs.datadoghq.com/api/latest/metrics/",
icon_url="https://imgix.datadoghq.com//img/about/presskit/DDlogo.jpg",
prerequisites=[CallablePrerequisite(callable=self.prerequisites_callable)],
Expand Down
6 changes: 3 additions & 3 deletions holmes/plugins/toolsets/grafana/toolset_grafana_loki.py
Original file line number Diff line number Diff line change
Expand Up @@ -47,10 +47,10 @@ def __init__(self):
icon_url="https://grafana.com/media/docs/loki/logo-grafana-loki.png",
docs_url="https://docs.robusta.dev/master/configuration/holmesgpt/toolsets/grafanaloki.html",
prerequisites=[CallablePrerequisite(callable=self.prerequisites_callable)],
tools=[
PodLoggingTool(self),
],
tools=[], # Initialize with empty tools first
)
# Now that parent is initialized and self.name exists, create the tool
self.tools = [PodLoggingTool(self)]

def prerequisites_callable(self, config: dict[str, Any]) -> tuple[bool, str]:
if not config:
Expand Down
6 changes: 3 additions & 3 deletions holmes/plugins/toolsets/kubernetes_logs.py
Original file line number Diff line number Diff line change
Expand Up @@ -67,11 +67,11 @@ def __init__(self):
icon_url="https://encrypted-tbn0.gstatic.com/images?q=tbn:ANd9GcRPKA-U9m5BxYQDF1O7atMfj9EMMXEoGu4t0Q&s",
prerequisites=[prerequisite],
is_default=True,
tools=[
PodLoggingTool(self),
],
tools=[], # Initialize with empty tools first
tags=[ToolsetTag.CORE],
)
# Now that parent is initialized and self.name exists, create the tool
self.tools = [PodLoggingTool(self)]
enabled, disabled_reason = self.health_check()
prerequisite.enabled = enabled
prerequisite.disabled_reason = disabled_reason
Expand Down
12 changes: 11 additions & 1 deletion holmes/plugins/toolsets/logging_utils/logging_api.py
Original file line number Diff line number Diff line change
Expand Up @@ -27,6 +27,9 @@ class LoggingCapability(str, Enum):

REGEX_FILTER = "regex_filter" # If not supported, falls back to substring matching
EXCLUDE_FILTER = "exclude_filter" # If not supported, parameter is not shown at all
HISTORICAL_DATA = (
"historical_data" # Can fetch logs for pods no longer in the cluster
)


class LoggingConfig(BaseModel):
Expand Down Expand Up @@ -78,9 +81,16 @@ def __init__(self, toolset: BasePodLoggingToolset):
parameters = self._get_tool_parameters(toolset)

# Build description based on capabilities
description = "Fetch logs for a Kubernetes pod"
# Include the toolset name in the description
toolset_name = toolset.name if toolset.name else "logging backend"
description = f"Fetch logs for a Kubernetes pod from {toolset_name}"
capabilities = toolset.supported_capabilities

if LoggingCapability.HISTORICAL_DATA in capabilities:
description += (
" (including historical data for pods no longer in the cluster)"
)

if (
LoggingCapability.REGEX_FILTER in capabilities
and LoggingCapability.EXCLUDE_FILTER in capabilities
Expand Down
6 changes: 3 additions & 3 deletions holmes/plugins/toolsets/opensearch/opensearch_logs.py
Original file line number Diff line number Diff line change
Expand Up @@ -45,13 +45,13 @@ def __init__(self):
docs_url="https://docs.robusta.dev/master/configuration/holmesgpt/toolsets/opensearch_logs.html",
icon_url="https://opensearch.org/wp-content/uploads/2025/01/opensearch_mark_default.png",
prerequisites=[CallablePrerequisite(callable=self.prerequisites_callable)],
tools=[
PodLoggingTool(self),
],
tools=[], # Initialize with empty tools first
tags=[
ToolsetTag.CORE,
],
)
# Now that parent is initialized and self.name exists, create the tool
self.tools = [PodLoggingTool(self)]

def get_example_config(self) -> Dict[str, Any]:
example_config = OpenSearchLoggingConfig(
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,9 @@
# Enables Datadog logs toolset to query historical logs

toolsets:
kubernetes/logs:
enabled: False

datadog/logs:
enabled: true
config:
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,58 @@
apiVersion: v1
kind: Namespace
metadata:
name: orbit-91g
---
apiVersion: v1
kind: Secret
metadata:
name: orbiter-script
namespace: orbit-91g
type: Opaque
stringData:
Comment thread
arikalon1 marked this conversation as resolved.
script.sh: |
#!/bin/sh
echo "Starting orbiter-monitor pod..."
# Simulate some CPU and memory usage
while true; do
# Consume some CPU with a calculation
echo "scale=5000; 4*a(1)" | busybox awk 'BEGIN{for(i=0;i<100;i++){}}' 2>/dev/null
# Sleep to create varying pattern
sleep 5
done
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: orbiter-monitor
namespace: orbit-91g
spec:
Comment thread
arikalon1 marked this conversation as resolved.
replicas: 1
selector:
matchLabels:
app: orbiter-monitor
template:
metadata:
labels:
app: orbiter-monitor
spec:
volumes:
- name: script-volume
secret:
secretName: orbiter-script
defaultMode: 0755
containers:
- name: app
image: busybox:latest
command: ["/bin/sh", "-c"]
args: ["/etc/scripts/script.sh"]
volumeMounts:
- name: script-volume
mountPath: /etc/scripts
resources:
requests:
memory: "256Mi"
cpu: "100m"
limits:
memory: "512Mi"
cpu: "500m"
Original file line number Diff line number Diff line change
@@ -0,0 +1,37 @@
description: |
This test verifies Holmes behavior when Datadog has metrics for a specific pod name
that doesn't match any existing pods in the cluster. The test sends metrics to Datadog
for a hardcoded pod name (orbiter-monitor-77aac74f7c-hl4cd) that doesn't exist in the cluster,
while deploying a real deployment named orbiter-monitor whose pods have different random suffixes.
Holmes should still be able to query and display the metrics from Datadog even though
the exact pod name doesn't exist in Kubernetes.

Comment thread
arikalon1 marked this conversation as resolved.
before_test: |
# Deploy the deployment to the cluster
kubectl apply -f manifest.yaml

# Wait for the pod to be running
kubectl wait --for=condition=ready pod -l app=orbiter-monitor -n orbit-91g --timeout=60s

# Send CPU and memory metrics to Datadog using a hardcoded pod name that won't exist
# This simulates historical data for a pod that may have been deleted/recreated
bash ../../shared/send_datadog_metrics.sh orbit-91g orbiter-monitor-77aac74f7c-hl4cd orbiter-monitor

user_prompt:
- "Show me a cpu graph of the orbiter-monitor-77aac74f7c-hl4cd pod in namespace orbit-91g?"
- "Show me memory graph for orbiter-monitor-77aac74f7c-hl4cd in orbit-91g"

expected_output: |
Output must contain 1 or more embeds in the following format
<<{"type": "datadogql", "tool_name": "query_datadog_metrics", "random_key": "iD8G"}>>

random_key may be different than the above example so long as its a random looking key, but all other parameters (type and tool_name) must be as described

Output must NOT tell the user it doesn't have access to metrics or that they should use another tool

tags:
- datadog

after_test: |
# Cleanup the deployed resources
kubectl delete -f manifest.yaml --ignore-not-found=true
Original file line number Diff line number Diff line change
@@ -0,0 +1,9 @@
toolsets:
kubernetes/core:
enabled: true
datadog/metrics:
enabled: true
config:
dd_api_key: "{{env.DATADOG_API_KEY}}"
dd_app_key: "{{env.DATADOG_APP_KEY}}"
site_api_url: "https://api.datadoghq.eu"
Original file line number Diff line number Diff line change
@@ -1,12 +1,12 @@
[
{
"role": "user",
"content": "Can you show the memory usage of the robusta-runner pod?",
"content": "Can you show the memory usage of the robusta-runner-78599b764d-f847h pod?",
"token_count": 510
},
{
"role": "assistant",
"content": "The `robusta-runner` pod in the `default` namespace is running without any OOM (Out of Memory) issues. Memory limits and requests are set to 1Gi. No OOM-related logs were found.",
"content": "The `robusta-runner-78599b764d-f847h` pod in the `default` namespace is running without any OOM (Out of Memory) issues. Memory limits and requests are set to 1Gi. No OOM-related logs were found.",
Comment thread
arikalon1 marked this conversation as resolved.
"annotations": []
}
]
Original file line number Diff line number Diff line change
@@ -1,3 +1,7 @@
before_test: |
# Send CPU and memory metrics to Datadog for a specific robusta-runner mentioned in conversation history
bash ../../shared/send_datadog_metrics.sh default robusta-runner-78599b764d-f847h robusta-runner
Comment thread
arikalon1 marked this conversation as resolved.

user_prompt:
- "Can you show this as a graph?"
- "Could you visualize this data in a chart?"
Expand Down
Loading
Loading