From b54d185b92c5062ae512908accecd5957f32e3ae Mon Sep 17 00:00:00 2001 From: Arik Alon Date: Wed, 29 Oct 2025 10:22:32 +0200 Subject: [PATCH] allow Holmes to read CRDs Will help to generate the CRDs feature required cluster roles fix aws mcp instructions --- .../templates/holmesgpt-service-account.yaml | 8 + .../templates/mcp-servers/aws/_helpers.tpl | 142 +++++++++++++++++- 2 files changed, 147 insertions(+), 3 deletions(-) diff --git a/helm/holmes/templates/holmesgpt-service-account.yaml b/helm/holmes/templates/holmesgpt-service-account.yaml index 22480915cd..99a05bb1e8 100644 --- a/helm/holmes/templates/holmesgpt-service-account.yaml +++ b/helm/holmes/templates/holmesgpt-service-account.yaml @@ -129,6 +129,14 @@ rules: - get - list + - apiGroups: + - "apiextensions.k8s.io" + resources: + - "customresourcedefinitions" + verbs: + - "list" + - "get" + - apiGroups: - networking.k8s.io resources: diff --git a/helm/holmes/templates/mcp-servers/aws/_helpers.tpl b/helm/holmes/templates/mcp-servers/aws/_helpers.tpl index b5581e19a4..cb28b87090 100644 --- a/helm/holmes/templates/mcp-servers/aws/_helpers.tpl +++ b/helm/holmes/templates/mcp-servers/aws/_helpers.tpl @@ -33,7 +33,7 @@ IMPORTANT: When investigating issues related to AWS resources or Kubernetes work - Check CloudTrail for security group modifications: `aws cloudtrail lookup-events --lookup-attributes AttributeKey=ResourceName,AttributeValue=SG_ID` ### When investigating configuration changes: -- Query CloudTrail for recent API calls: `aws cloudtrail lookup-events --start-time TIME --max-items 50` +- Query CloudTrail for recent API calls: `aws cloudtrail lookup-events --start-time TIME --max-items 100` - Filter by specific resources: `aws cloudtrail lookup-events --lookup-attributes AttributeKey=ResourceName,AttributeValue=RESOURCE_ID` - Look for security-related changes: Search for events like RevokeSecurityGroupIngress, AuthorizeSecurityGroupIngress, ModifyDBInstance, etc. - Identify who made changes: Check UserName and SourceIPAddress in CloudTrail events @@ -42,6 +42,12 @@ IMPORTANT: When investigating issues related to AWS resources or Kubernetes work - Check EKS cluster status: `aws eks describe-cluster --name CLUSTER_NAME` - Examine node groups: `aws eks describe-nodegroup --cluster-name CLUSTER_NAME --nodegroup-name NODEGROUP` - Search CloudWatch Container Insights logs if available +```bash +aws logs filter-log-events \ + --log-group-name /aws/containerinsights/CLUSTER_NAME/application \ + --start-time $(date -d '1 hour ago' +%s)000 \ + --max-items 500 +``` - Check EC2 instances hosting the nodes: `aws ec2 describe-instances --instance-ids INSTANCE_ID` ### For networking issues: @@ -60,7 +66,7 @@ IMPORTANT: When investigating issues related to AWS resources or Kubernetes work ### CloudTrail Investigation (ALWAYS use when troubleshooting): Find all recent changes in the last hour: ``` -aws cloudtrail lookup-events --start-time $(date -u -d '1 hour ago' +%Y-%m-%dT%H:%M:%S) --max-items 50 +aws cloudtrail lookup-events --start-time $(date -u -d '1 hour ago' +%Y-%m-%dT%H:%M:%S) --max-items 100 ``` Find changes to a specific security group: @@ -70,7 +76,7 @@ aws cloudtrail lookup-events --lookup-attributes AttributeKey=ResourceName,Attri Find who made changes: ``` -aws cloudtrail lookup-events --lookup-attributes AttributeKey=EventName,AttributeValue=RevokeSecurityGroupIngress --max-items 10 +aws cloudtrail lookup-events --lookup-attributes AttributeKey=EventName,AttributeValue=RevokeSecurityGroupIngress --max-items 20 ``` ### Resource State Verification: @@ -105,5 +111,135 @@ Example commands: - aws rds describe-db-instances - aws cloudwatch get-metric-statistics - aws ce get-cost-and-usage + +## ⚠️ MEMORY OPTIMIZATION GUIDELINES ⚠️ + +**The AWS MCP server can experience memory pressure with very large queries. Follow these guidelines to balance data retrieval with stability:** + +### Query Limits by Service Type + +| Service | Safe Limit | Max Limit | Notes | +|---------|-----------|-----------|--------| +| **CloudWatch Logs** | 500 items | 1000 items | ALWAYS use time constraints (1-hour window max initially) | +| **CloudTrail** | 200 items | 500 items | Heavy JSON payloads, use 2-hour windows initially | +| **EC2 Describe** | 500 items | 1000 items | Use filters when possible | +| **RDS/ELB** | 100 items | 200 items | Usually fewer resources | +| **S3 List** | 1000 items | 5000 items | Metadata only, not object contents | +| **Cost & Usage** | 30 days | 90 days | Use DAILY granularity, not HOURLY | + +### Critical Rules +1. **NEVER download S3 object contents** (can be GBs) +2. **ALWAYS use time constraints for logs** (CloudWatch, VPC Flow Logs) +3. **ALWAYS use RECENT time constraints for logs** - Even with --max-items, AWS scans ALL data in the time range! +4. **ALWAYS use small time windows. 1-2 hours. It's ok to do multiple queries. +5. **START with recommended limits**, increase if needed +6. **If query times out**, reduce time window FIRST (not just --max-items)3. **START with recommended limits**, increase if needed + +## Investigation Principles + +## Memory-Optimized Query Examples + +### CloudTrail Investigation (Medium Memory Risk): +```bash +# Good - Reasonable time window with sufficient data +aws cloudtrail lookup-events --start-time $(date -u -d '2 hours ago' +%Y-%m-%dT%H:%M:%S) --max-items 200 + +# Better - Targeted search for specific events +aws cloudtrail lookup-events \ +--lookup-attributes AttributeKey=EventName,AttributeValue=RevokeSecurityGroupIngress \ +--start-time $(date -u -d '4 hours ago' +%Y-%m-%dT%H:%M:%S) \ +--max-items 100 + +# If you need more history, paginate +aws cloudtrail lookup-events --start-time $(date -u -d '6 hours ago' +%Y-%m-%dT%H:%M:%S) --max-items 200 +# Then use --starting-token if needed for next page +``` + +### CloudWatch Logs (High Memory Risk): +```bash +# Safe - Time-bounded with reasonable limits +aws logs filter-log-events \ +--log-group-name /aws/eks/cluster/cluster \ +--start-time $(date -d '1 hour ago' +%s)000 \ +--filter-pattern "ERROR" \ +--max-items 500 + +# For debugging specific pods/containers +aws logs filter-log-events \ +--log-group-name /aws/containerinsights/CLUSTER/application \ +--filter-pattern "{ $.kubernetes.pod_name = \"POD_NAME\" }" \ +--start-time $(date -d '2 hours ago' +%s)000 \ +--max-items 300 +``` + +### EC2 Operations: +```bash +# Can handle more items for instance metadata +aws ec2 describe-instances --max-results 500 + +# With filters for large environments +aws ec2 describe-instances \ +--filters "Name=tag:Environment,Values=production" \ +--max-results 200 +``` + + +## Progressive Investigation Strategy + +### Start Conservative, Then Expand: +1. **Initial Query**: Use recommended limits (see table above) +2. **If Insufficient**: Double the limit or time window +3. **If Times Out**: Halve the limit and add more filters +4. **For Historical Analysis**: Use pagination with --starting-token + + +## What NOT to Query: + +- ❌ **S3 Object Contents**: Use presigned URLs or tell user to download +- ❌ **CloudWatch Logs without time bounds**: Always specify --start-time +- ❌ **VPC Flow Logs for busy networks**: Use specific filters or sample +- ❌ **Full Config History**: Use --limit and specific resource types +- ❌ **X-Ray Traces in bulk**: Query specific trace IDs + +### 🔄 PAGINATION BEST PRACTICES + +**Use pagination to prevent OOM while getting comprehensive data:** + +#### How to Paginate: +```bash +# Step 1: Initial query with --max-items +aws cloudtrail lookup-events --start-time $(date -u -d '2 hours ago' +%Y-%m-%dT%H:%M:%S) --max-items 100 > page1.json + +# Step 2: Check if NextToken exists in output +# If NextToken exists, there's more data available + +# Step 3: Get next page using --starting-token +aws cloudtrail lookup-events --start-time $(date -u -d '2 hours ago' +%Y-%m-%dT%H:%M:%S) --max-items 100 --starting-token > page2.json + +# Continue until no NextToken is returned + +Services Supporting Pagination: + +- CloudTrail: Use --max-items and --starting-token +- CloudWatch Logs: Use --max-items and --starting-token +- EC2: Use --max-results and --next-token +- RDS: Use --max-records and --marker +- S3: Use --max-items and --starting-token +- Cost Explorer: Results are paginated by default with NextPageToken + +When to Use Pagination: + +- Always for CloudTrail when investigating beyond 1 hour +- Always for CloudWatch Logs when searching broad patterns +- For EC2 when describing >200 instances +- For any query that returns a NextToken/Marker + +Pagination Strategy: + +1. Start with smaller pages (100-200 items) +2. Process each page before fetching next +3. Stop when you find what you need +4. If investigating trends, sample pages instead of fetching all + {{- end -}} {{- end -}}