Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
30 commits
Select commit Hold shift + click to select a range
d60516f
feat: datadog toolset
nherment Jul 15, 2025
7912a2a
feat: improve rate limit logic to also apply to first API call
nherment Jul 15, 2025
660d5df
fix: fetch logs from latest to earliest, use datetimes instead of tim…
nherment Jul 15, 2025
4c08985
chore: linting
nherment Jul 15, 2025
4a57df0
test: add prereq tests for datadog/logs toolset
nherment Jul 15, 2025
5d97c8c
test: add integration tests for datadog/logs toolset
nherment Jul 15, 2025
0e01b78
chore: linting
nherment Jul 15, 2025
7b2cf58
chore: address PR comments
nherment Jul 16, 2025
3ff03a8
Merge branch 'master' into rob-1723_datadog_logs_toolset
nherment Jul 16, 2025
fa98a7e
Merge branch 'master' into rob-1723_datadog_logs_toolset
nherment Jul 16, 2025
f4179a2
fix: tests:
nherment Jul 16, 2025
507032b
chore: address PR comments
nherment Jul 16, 2025
66e0e3c
feat: add datadog metrics toolset
nherment Jul 17, 2025
a08ff49
feat: add datadog metrics toolset
nherment Jul 17, 2025
df6433a
chore: linting
nherment Jul 17, 2025
78d7283
fix: test
nherment Jul 17, 2025
1e799d8
chore: linting
nherment Jul 17, 2025
78be66a
chore: linting
nherment Jul 17, 2025
1761dad
Merge branch 'master' into rob-1740_datadog_metrics_toolset
nherment Jul 17, 2025
8b3186b
feat: datadog/traces toolset
nherment Jul 17, 2025
259771c
Merge branch 'master' into rob-1740_datadog_metrics_toolset
nherment Jul 18, 2025
607efb5
feat: datadog/traces toolset
nherment Jul 18, 2025
43e7ee1
Merge branch 'rob-1740_datadog_metrics_toolset' into rob-1739_datadog…
nherment Jul 18, 2025
00750bf
feat: datadog/rds toolset
nherment Jul 18, 2025
6c33c70
Merge branch 'rob-1739_datadog_traces_toolset' into rob-1741_datadog_…
nherment Jul 18, 2025
026e4bf
Merge branch 'master' into rob-1741_datadog_db_analysis_toolset
nherment Aug 4, 2025
a46b64e
chore: linting
nherment Aug 4, 2025
f0fc685
Merge branch 'master' into rob-1741_datadog_db_analysis_toolset
nherment Aug 5, 2025
9c6df5f
Merge branch 'master' into rob-1741_datadog_db_analysis_toolset
nherment Aug 5, 2025
36aa1c9
Merge branch 'master' into rob-1741_datadog_db_analysis_toolset
nherment Aug 5, 2025
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 7 additions & 1 deletion holmes/plugins/toolsets/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -20,7 +20,12 @@
from holmes.plugins.toolsets.datadog.toolset_datadog_metrics import (
DatadogMetricsToolset,
)
from holmes.plugins.toolsets.datadog.toolset_datadog_traces import DatadogTracesToolset
from holmes.plugins.toolsets.datadog.toolset_datadog_traces import (
DatadogTracesToolset,
)
from holmes.plugins.toolsets.datadog.toolset_datadog_rds import (
DatadogRDSToolset,
)
from holmes.plugins.toolsets.git import GitToolset
from holmes.plugins.toolsets.grafana.toolset_grafana import GrafanaToolset
from holmes.plugins.toolsets.grafana.toolset_grafana_loki import GrafanaLokiToolset
Expand Down Expand Up @@ -75,6 +80,7 @@ def load_python_toolsets(dal: Optional[SupabaseDal]) -> List[Toolset]:
DatadogLogsToolset(),
DatadogMetricsToolset(),
DatadogTracesToolset(),
DatadogRDSToolset(),
PrometheusToolset(),
OpenSearchLogsToolset(),
OpenSearchTracesToolset(),
Expand Down
82 changes: 82 additions & 0 deletions holmes/plugins/toolsets/datadog/datadog_rds_instructions.jinja2
Original file line number Diff line number Diff line change
@@ -0,0 +1,82 @@
## Datadog RDS Performance Analysis Instructions

You have access to tools for analyzing RDS database performance and identifying problematic instances using Datadog metrics.

### Available Tools:

1. **datadog_rds_performance_report** - Generate comprehensive performance report for a specific RDS instance
- Analyzes latency, resource utilization, and storage metrics
- Identifies performance issues and bottlenecks
- Provides actionable recommendations
- Returns formatted report with executive summary

2. **datadog_rds_top_worst_performing** - Get summary of worst performing RDS instances
- Analyzes all RDS instances in the environment
- Ranks by latency, CPU, or composite performance score
- Shows top N worst performers with their key metrics
- Helps prioritize optimization efforts

### Usage Guidelines:

**For investigating a specific RDS instance:**
```
Use datadog_rds_performance_report with:
- db_instance_identifier: "instance-name"
- start_time: "-3600" (last hour)
```

**For finding problematic instances across the fleet:**
```
Use datadog_rds_top_worst_performing with:
- top_n: 10 (show top 10 worst)
- sort_by: "latency" (or "cpu", "composite")
- start_time: "-3600"
```

### Key Performance Thresholds:

The tools automatically flag issues based on these thresholds:
- **Latency**: >10ms average (warning), >50ms peak (critical)
- **CPU**: >70% average (warning), >90% peak (critical)
- **Memory**: <100MB freeable memory (warning)
- **Burst Balance**: <30% (warning, indicates I/O constraints)
- **Disk Queue Depth**: >5 average (indicates I/O bottleneck)

### Common Scenarios:

1. **Application experiencing slow database queries:**
- Generate performance report for the specific RDS instance
- Look for latency spikes and resource constraints
- Follow recommendations for optimization

2. **Proactive performance monitoring:**
- Use top worst performing to identify problem instances
- Generate detailed reports for the worst performers
- Plan capacity upgrades based on findings

3. **Capacity planning:**
- Analyze resource utilization trends
- Identify instances approaching limits
- Plan upgrades before performance degradation

### Interpreting Results:

**Performance Report Sections:**
- **Executive Summary**: High-level assessment and severity
- **Metrics Tables**: Statistical analysis of each metric
- **Issues**: Specific problems detected with thresholds exceeded
- **Recommendations**: Prioritized actions to resolve issues

**Top Worst Performing Report:**
- **Rankings**: Instances sorted by selected metric
- **Key Metrics**: Latency, CPU, burst balance for each instance
- **Summary**: Overall patterns across the fleet

### Example Responses:

When asked about database performance issues:
1. First use `datadog_rds_top_worst_performing` to identify problem instances
2. Then use `datadog_rds_performance_report` on the worst performers
3. Summarize findings and provide prioritized recommendations

Always consider the time range - recent data (last hour) for current issues, longer ranges (last 24 hours) for trends.
Loading