Skip to content

GCP MCP integration - #1310

Merged
arikalon1 merged 8 commits into
masterfrom
gcloud-mcp
Jan 5, 2026
Merged

arikalon1 merged 8 commits into
masterfrom
gcloud-mcp

Conversation

@arikalon1

@arikalon1 arikalon1 commented Jan 2, 2026 •

Copy link
Copy Markdown
Collaborator

Add support for GCP mcp servers
include:

  • gcloud (gcloud cli)
  • observability
  • storage

Summary by CodeRabbit

  • New Features

    • Added Google Cloud Platform (GCP) MCP toolset with gcloud, observability, and storage components; per-component enablement, images, ports, resources, deployment options, optional NetworkPolicy, and workload-identity-friendly defaults.
  • Enhancements

    • Improved MCP command one-liner generation and handling for GCP commands.
    • Toolset configuration now merges GCP MCP servers into the overall MCP registry when enabled.
  • Documentation

    • Added comprehensive GCP MCP documentation and updated built-in toolsets index to include GCP and MariaDB.

✏️ Tip: You can customize this high-level summary in your review settings.

@linux-foundation-easycla

linux-foundation-easycla Bot commented Jan 2, 2026 •

Copy link
Copy Markdown

CLA Not Signed

@coderabbitai

coderabbitai Bot commented Jan 2, 2026 •

Copy link
Copy Markdown
Contributor

Walkthrough

Adds Google Cloud Platform (GCP) MCP support: Helm templates (helpers, deployment, networkpolicy), values and toolset-config integration, new GCP docs/nav/index entries, and a small MCP plugin change for CLI/gcloud argument formatting.

Changes

Cohort / File(s) Summary
Docs
docs/data-sources/builtin-toolsets/.nav.yml, docs/data-sources/builtin-toolsets/index.md, docs/data-sources/builtin-toolsets/gcp.md
Add GCP (MCP) documentation, nav and index entries; add MariaDB (MCP) index entry; adjust nav indentation and provide detailed GCP MCP usage/deployment guidance.
Helm — GCP helpers
helm/holmes/templates/mcp-servers/gcp/_helpers.tpl
New helper templates defining overridable llmInstructions for gcloud, observability, and storage.
Helm — GCP deployment
helm/holmes/templates/mcp-servers/gcp/deployment.yaml
New conditional ConfigMap, optional ServiceAccount, multi-container Deployment (gcloud/observability/storage) with volumes, probes, resources, and Service; gated by .Values.mcpAddons.gcp.*.
Helm — GCP network policy
helm/holmes/templates/mcp-servers/gcp/networkpolicy.yaml
New NetworkPolicy guarded by .Values.mcpAddons.gcp.networkPolicy.enabled; per-service ingress toggles; egress allows DNS and external traffic while blocking GCP metadata IP.
Helm — Toolset config & values
helm/holmes/templates/toolset-config.yaml, helm/holmes/values.yaml
Add mcpAddons.gcp values subtree and toolset-config logic to construct/merge gcp_gcloud, gcp_observability, gcp_storage into mcp_servers; per-service toggles, images, ports, resources, and llmInstructions placeholders.
MCP plugin
holmes/plugins/toolsets/mcp/toolset_mcp.py
get_parameterized_one_liner now prefers params.cli_command; special-cases run_gcloud_command with list args to format gcloud <args>; removes prior MCPConfig/StdioMCPConfig-specific string output.

Sequence Diagram(s)

sequenceDiagram
    autonumber
    participant Dev as Developer (values.yaml)
    participant Helm as Helm templates
    participant K8s as Kubernetes API
    participant MCP as GCP MCP Pods
    participant Holmes as Holmes app

    Dev->>Helm: provide .Values.mcpAddons.gcp + llmInstructions
    Helm->>K8s: render & apply ConfigMap, ServiceAccount, Deployment, Service, NetworkPolicy
    K8s->>MCP: schedule Pods (gcloud / observability / storage)
    Helm->>Holmes: render toolset-config merging gcp_* into mcp_servers
    Holmes->>MCP: call MCP endpoints (commands formatted by toolset_mcp.get_parameterized_one_liner)
    MCP-->>Holmes: responses / outputs
    Note right of Helm: llmInstructions included in generated toolset config
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly related PRs

Suggested reviewers

  • moshemorad
  • nherment
  • Avi-Robusta

Pre-merge checks

❌ Failed checks (1 warning)
Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. You can run @coderabbitai generate docstrings to improve docstring coverage.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title 'GCP MCP integration' accurately reflects the main change—adding support for GCP MCP servers with gcloud, observability, and storage components across documentation, Helm charts, and code.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions

github-actions Bot commented Jan 2, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker image ready for 51ea980 (built in 47s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use this tag to pull the image for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:51ea980
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:51ea980 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:51ea980
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:51ea980

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:51ea980

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:51ea980

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 7

🧹 Nitpick comments (2)
holmes/plugins/toolsets/mcp/toolset_mcp.py (1)

188-190: Minor redundancy in f-string.

The params.get('cli_command') is already checked in the condition, so you can use it directly.

🔎 Suggested simplification
         # AWS MCP cli_command
         if params and params.get("cli_command"):
-            return f"{params.get('cli_command')}"
+            return params["cli_command"]
helm/holmes/values.yaml (1)

222-227: Resource limits are inconsistent across GCP services.

The gcloud service has a CPU limit defined in requests but is missing it in limits, while observability and storage have only memory limits. Consider adding CPU limits for consistency and better resource management:

🔎 Suggested resource consistency
     gcloud:
       enabled: true
       image: "gcloud-cli-mcp:1.0.7"
       port: 8000
       resources:
         requests:
           memory: "256Mi"
           cpu: "100m"
         limits:
           memory: "1Gi"
+          cpu: "500m"

Apply similar CPU limits to observability and storage services for consistency.

📜 Review details

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 0eced48 and 1d25e71229460a48f22e4434ec8ad30f1809d667.

📒 Files selected for processing (8)
  • docs/data-sources/builtin-toolsets/.nav.yml
  • docs/data-sources/builtin-toolsets/index.md
  • helm/holmes/templates/mcp-servers/gcp/_helpers.tpl
  • helm/holmes/templates/mcp-servers/gcp/deployment.yaml
  • helm/holmes/templates/mcp-servers/gcp/networkpolicy.yaml
  • helm/holmes/templates/toolset-config.yaml
  • helm/holmes/values.yaml
  • holmes/plugins/toolsets/mcp/toolset_mcp.py
🧰 Additional context used
📓 Path-based instructions (3)
docs/**/*.md

📄 CodeRabbit inference engine (CLAUDE.md)

When writing documentation in the docs/ directory, always add a blank line between a header/bold text and a list, otherwise MkDocs won't render the list properly

Files:

  • docs/data-sources/builtin-toolsets/index.md
**/*.py

📄 CodeRabbit inference engine (CLAUDE.md)

**/*.py: Use Ruff for formatting and linting (configured in pyproject.toml)
Type hints required (mypy configuration in pyproject.toml)
ALWAYS place Python imports at the top of the file, not inside functions or methods

Files:

  • holmes/plugins/toolsets/mcp/toolset_mcp.py
holmes/plugins/toolsets/**

📄 CodeRabbit inference engine (CLAUDE.md)

Toolsets: organize as holmes/plugins/toolsets/{name}.yaml or {name}/ directories

Files:

  • holmes/plugins/toolsets/mcp/toolset_mcp.py
🧠 Learnings (1)
📚 Learning: 2025-12-29T08:35:37.678Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-12-29T08:35:37.678Z
Learning: New toolsets require integration tests

Applied to files:

  • docs/data-sources/builtin-toolsets/index.md
🪛 YAMLlint (1.37.1)
helm/holmes/templates/mcp-servers/gcp/networkpolicy.yaml

[error] 1-1: syntax error: expected the node content, but found '-'

(syntax)

helm/holmes/templates/toolset-config.yaml

[error] 1-1: syntax error: expected the node content, but found '-'

(syntax)

helm/holmes/templates/mcp-servers/gcp/deployment.yaml

[error] 1-1: syntax error: expected the node content, but found '-'

(syntax)

🔇 Additional comments (12)
holmes/plugins/toolsets/mcp/toolset_mcp.py (1)

192-198: Logic looks good for gcloud command generation.

The implementation correctly handles the run_gcloud_command tool by joining args into a proper gcloud CLI command. The isinstance check provides good defensive typing for the args parameter.

docs/data-sources/builtin-toolsets/index.md (1)

29-29: MariaDB (MCP) entry looks good.

Consistent with other entries in the index.

helm/holmes/templates/mcp-servers/gcp/networkpolicy.yaml (2)

52-56: Egress rule is permissive by design.

The egress allows all external traffic except the GCP metadata service. While this is broad, it's pragmatic for GCP API access since Google's IP ranges change frequently. The metadata service block (169.254.169.254/32) is a good security measure to enforce Workload Identity.


22-41: Ingress rules correctly scoped to Holmes pods.

The conditional port exposure based on enabled flags is well-implemented.

helm/holmes/values.yaml (1)

184-262: GCP MCP configuration structure is comprehensive.

The configuration provides good flexibility with:

  • Multiple authentication methods (service account key, workload identity)
  • Per-service toggles and resource controls
  • Placement options (nodeSelector, tolerations, affinity)
  • LLM instruction overrides

The documentation comments explaining authentication setup and multi-project support are helpful.

helm/holmes/templates/mcp-servers/gcp/deployment.yaml (3)

91-96: Good security context configuration.

Running as non-root user (1000) with proper fsGroup is a security best practice.


206-273: Observability container configuration looks correct.

The container properly sets GOOGLE_APPLICATION_CREDENTIALS when service account key is provided, includes ConfigMap references for project/region, and has appropriate health probes.


343-379: Service definition correctly exposes all enabled MCP ports.

The conditional port exposure matches the container definitions.

helm/holmes/templates/toolset-config.yaml (1)

33-69: GCP MCP server configuration block is well-structured.

The implementation follows the established pattern from AWS and MariaDB MCP servers. The conditional merging of each subcomponent (gcloud, observability, storage) is clean and allows granular control. All referenced helper templates (holmes.gcpMcp.gcloud.llmInstructions, holmes.gcpMcp.observability.llmInstructions, holmes.gcpMcp.storage.llmInstructions) are properly defined in helm/holmes/templates/mcp-servers/gcp/_helpers.tpl.

helm/holmes/templates/mcp-servers/gcp/_helpers.tpl (3)

4-56: Well-structured template with comprehensive gcloud guidance.

The template definition follows proper Helm patterns and provides clear, detailed instructions for LLM consumption. The emphasis on args array format (excluding 'gcloud' prefix) and multi-project support with explicit --project flags is particularly valuable.


61-111: Excellent observability instructions with strong link guidance.

The template provides comprehensive guidance for GCP observability with particularly strong emphasis on:

  • Historical data retrieval for deleted resources (lines 67-68)
  • Descriptive link naming conventions (lines 79-85)
  • Practical filter patterns for common scenarios

The critical instruction at line 79 to "always answer log questions with actual log entries first" is especially valuable for ensuring LLMs provide actionable results before offering console links.


116-147: Clear and practical storage instructions.

The template provides focused guidance for Cloud Storage investigations with logical troubleshooting flows. The investigation patterns (lines 136-141) are particularly helpful, offering a clear sequence for common issues like access denied and missing data.

Comment thread docs/data-sources/builtin-toolsets/.nav.yml
Comment thread docs/data-sources/builtin-toolsets/index.md
Comment thread helm/holmes/templates/mcp-servers/gcp/_helpers.tpl
Comment thread helm/holmes/templates/mcp-servers/gcp/deployment.yaml
Comment thread helm/holmes/templates/mcp-servers/gcp/deployment.yaml
Comment thread helm/holmes/values.yaml Outdated
Comment thread helm/holmes/values.yaml
@arikalon1

arikalon1 commented Jan 2, 2026 •

Copy link
Copy Markdown
Collaborator Author

fixes #1195

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

♻️ Duplicate comments (2)
helm/holmes/values.yaml (2)

162-163: Placeholder URLs should be updated.

Lines 163 and 196 contain placeholder URLs (https://github.com/your-org/holmes-mcp-integrations) that should be replaced with the actual repository location or removed if images are already published.

Also applies to: 196-196


261-309: Azure MCP templates are still missing.

The Azure MCP values are defined here, but corresponding Helm templates are not present in helm/holmes/templates/mcp-servers/azure/. These values will have no effect without the associated templates.

🧹 Nitpick comments (2)
helm/holmes/values.yaml (2)

219-248: Consider adding CPU limits for consistency with other MCP addons.

The GCP MCP servers define CPU requests but omit CPU limits, while AWS and MariaDB addons in the same file specify both. Without CPU limits, containers can consume unbounded CPU, potentially causing resource contention.

🔎 Suggested addition of CPU limits
     gcloud:
       enabled: true
       image: "gcloud-cli-mcp:1.0.7"
       port: 8000
       resources:
         requests:
           memory: "256Mi"
           cpu: "100m"
         limits:
           memory: "1Gi"
+          cpu: "500m"

     # Observability MCP - Cloud Logging, Monitoring, Trace, Error Reporting
     observability:
       enabled: true
       image: "gcloud-observability-mcp:1.0.0"
       port: 8001
       resources:
         requests:
           memory: "256Mi"
           cpu: "100m"
         limits:
           memory: "512Mi"
+          cpu: "250m"

     # Storage MCP - Cloud Storage operations
     storage:
       enabled: true
       image: "gcloud-storage-mcp:1.0.0"
       port: 8002
       resources:
         requests:
           memory: "256Mi"
           cpu: "100m"
         limits:
           memory: "512Mi"
+          cpu: "250m"

288-293: Azure resources also lack CPU limits.

Same observation as the GCP section: CPU requests are specified but limits are omitted. For consistency with AWS and MariaDB addons, consider adding CPU limits.

📜 Review details

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 1d25e71229460a48f22e4434ec8ad30f1809d667 and 76140490295301514bfecd489e5a4d034d97d4ca.

📒 Files selected for processing (3)
  • docs/data-sources/builtin-toolsets/index.md
  • helm/holmes/values.yaml
  • holmes/plugins/toolsets/mcp/toolset_mcp.py
🚧 Files skipped from review as they are similar to previous changes (2)
  • holmes/plugins/toolsets/mcp/toolset_mcp.py
  • docs/data-sources/builtin-toolsets/index.md
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (1)
  • GitHub Check: llm_evals

Comment thread helm/holmes/values.yaml
@github-actions

github-actions Bot commented Jan 2, 2026

Copy link
Copy Markdown
Contributor

✅ Results of HolmesGPT evals

Automatically triggered by commit 7614049 on branch gcloud-mcp

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 31.8s ±0% 6 13 $0.1649
✅ 101_loki_historical_logs_pod_deleted 69.9s ↑33% 11 26 $0.2907
✅ 111_pod_names_contain_service 35.8s ↓15% 6 14 $0.1616
✅ 12_job_crashing 43.9s ±0% 8 17 $0.2044
✅ 162_get_runbooks 48.7s ±0% 7 20 $0.2373
✅ 176_network_policy_blocking_traffic_no_runbooks 42.8s ±0% 7 15 $0.1952
✅ 24_misconfigured_pvc 35.8s ±0% 7 17 $0.1744
✅ 43_current_datetime_from_prompt 3.4s ±0% 1 — $0.0621
✅ 61_exact_match_counting 10.5s ±0% 3 3 $0.0861
Total 35.8s avg 6.2 avg 15.6 avg $1.5766

Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%.

Historical Comparison Details

Filter: excluding branch 'gcloud-mcp'

Status: Success - 22 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref gcloud-mcp -f markers=regression

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /last to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

benchmark, chain-of-causation, compaction, context_window, coralogix, counting, database, datadog, datetime, easy, embeds, grafana-dashboard, hard, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /last · /list

@github-actions

github-actions Bot commented Jan 2, 2026

Copy link
Copy Markdown
Contributor

✅ Results of HolmesGPT evals

Automatically triggered by commit 2dd1dfd on branch gcloud-mcp

View workflow logs

⚠️ No eval report was generated.

📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref gcloud-mcp -f markers=regression

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /last to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

benchmark, chain-of-causation, compaction, context_window, coralogix, counting, database, datadog, datetime, easy, embeds, grafana-dashboard, hard, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /last · /list

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

♻️ Duplicate comments (1)
helm/holmes/values.yaml (1)

163-163: Placeholder URLs remain unfixed.

Both references to https://github.com/your-org/holmes-mcp-integrations are placeholders that should be updated to the actual repository location before merging.

Also applies to: 196-196

🧹 Nitpick comments (3)
helm/holmes/values.yaml (3)

219-225: Add CPU limit for consistency and resource isolation.

The gcloud service defines a CPU limit, but it's missing from the YAML. The observability (line 236) and storage (line 248) services both specify cpu limits under limits. Adding a CPU limit prevents unbounded CPU consumption and aligns with Kubernetes best practices.

🔎 Suggested fix
       resources:
         requests:
           memory: "256Mi"
           cpu: "100m"
         limits:
           memory: "1Gi"
+          cpu: "500m"

231-237: Add CPU limit for resource isolation.

The observability service is missing a cpu limit. Adding one prevents unbounded CPU consumption and follows Kubernetes best practices for resource management.

🔎 Suggested fix
       resources:
         requests:
           memory: "256Mi"
           cpu: "100m"
         limits:
           memory: "512Mi"
+          cpu: "500m"

243-249: Add CPU limit for resource isolation.

The storage service is missing a cpu limit. Adding one prevents unbounded CPU consumption and follows Kubernetes best practices for resource management.

🔎 Suggested fix
       resources:
         requests:
           memory: "256Mi"
           cpu: "100m"
         limits:
           memory: "512Mi"
+          cpu: "500m"
📜 Review details

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 76140490295301514bfecd489e5a4d034d97d4ca and 2dd1dfd80325d63459996e63089eb7f845c9a86d.

📒 Files selected for processing (1)
  • helm/holmes/values.yaml
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (4)
  • GitHub Check: llm_evals
  • GitHub Check: build (3.12)
  • GitHub Check: build (3.11)
  • GitHub Check: build (3.10)
🔇 Additional comments (2)
helm/holmes/values.yaml (2)

254-259: The nested llmInstructions structure is correctly implemented. The three GCP services (gcloud, observability, storage) each have dedicated template definitions in _helpers.tpl that properly check for nested overrides at .Values.mcpAddons.gcp.llmInstructions.{gcloud|observability|storage}, with sensible defaults provided when overrides are empty. All three templates are properly included in toolset-config.yaml with appropriate filters. This nested design is appropriate for GCP's multi-service architecture (unlike AWS and MariaDB's single-service flat structure).


189-189: Remove or update the Workload Identity gcloud CLI limitation comment—it's inaccurate.

Workload Identity DOES work with gcloud CLI in GKE. Pods with Workload Identity enabled can authenticate using gcloud auth print-access-token or Application Default Credentials. Workload Identity Federation also supports gcloud CLI (v363.0.0+, available as of Jan 2026). The claim that it "currently not working with gcloud CLI" is incorrect and should be removed or clarified.

Also applies to: 200-202

Likely an incorrect or invalid review comment.

Comment thread helm/holmes/values.yaml Outdated
Signed-off-by: Arik Alon <alon.arik@gmail.com>
Signed-off-by: Arik Alon <alon.arik@gmail.com>
Signed-off-by: Arik Alon <alon.arik@gmail.com>
Signed-off-by: Arik Alon <alon.arik@gmail.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🧹 Nitpick comments (1)
docs/data-sources/builtin-toolsets/gcp.md (1)

339-355: Consider indented code blocks in numbered lists for better Markdown style consistency.

Fenced code blocks (triple backticks) inside numbered lists can sometimes cause rendering issues. Markdownlint prefers indented code blocks (4+ space indentation) within lists for proper nesting. The current fenced blocks work, but converting them to indented style would align with Markdown best practices for list content.

Also applies to: 396-419

📜 Review details

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 8919b19 and 641743f.

📒 Files selected for processing (1)
  • docs/data-sources/builtin-toolsets/gcp.md
🧰 Additional context used
📓 Path-based instructions (1)
docs/**/*.md

📄 CodeRabbit inference engine (CLAUDE.md)

When writing documentation in the docs/ directory, always add a blank line between a header/bold text and a list, otherwise MkDocs won't render the list properly

Files:

  • docs/data-sources/builtin-toolsets/gcp.md
🪛 LanguageTool
docs/data-sources/builtin-toolsets/gcp.md

[grammar] ~7-~7: Ensure spelling is correct
Context: ...ed resources. ## Overview The GCP MCP addon consists of three specialized servers: ...

(QB_NEW_EN_ORTHOGRAPHY_ERROR_IDS_1)

🪛 markdownlint-cli2 (0.18.1)
docs/data-sources/builtin-toolsets/gcp.md

341-341: Code block style
Expected: indented; Actual: fenced

(MD046, code-block-style)


396-396: Code block style
Expected: indented; Actual: fenced

(MD046, code-block-style)


451-451: Fenced code blocks should have a language specified

(MD040, fenced-code-language)


451-451: Code block style
Expected: indented; Actual: fenced

(MD046, code-block-style)


457-457: Fenced code blocks should have a language specified

(MD040, fenced-code-language)


457-457: Code block style
Expected: indented; Actual: fenced

(MD046, code-block-style)


463-463: Fenced code blocks should have a language specified

(MD040, fenced-code-language)


463-463: Code block style
Expected: indented; Actual: fenced

(MD046, code-block-style)


469-469: Fenced code blocks should have a language specified

(MD040, fenced-code-language)


469-469: Code block style
Expected: indented; Actual: fenced

(MD046, code-block-style)


475-475: Fenced code blocks should have a language specified

(MD040, fenced-code-language)


475-475: Code block style
Expected: indented; Actual: fenced

(MD046, code-block-style)


483-483: Code block style
Expected: indented; Actual: fenced

(MD046, code-block-style)


495-495: Code block style
Expected: indented; Actual: fenced

(MD046, code-block-style)


502-502: Code block style
Expected: indented; Actual: fenced

(MD046, code-block-style)

⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (4)
  • GitHub Check: llm_evals
  • GitHub Check: build (3.12)
  • GitHub Check: build (3.10)
  • GitHub Check: build (3.11)

Comment thread docs/data-sources/builtin-toolsets/gcp.md
Comment thread docs/data-sources/builtin-toolsets/gcp.md
@github-actions

github-actions Bot commented Jan 2, 2026

Copy link
Copy Markdown
Contributor

✅ Results of HolmesGPT evals

Automatically triggered by commit 8919b19 on branch gcloud-mcp

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 28.7s ↓15% 5 12 $0.1527
✅ 101_loki_historical_logs_pod_deleted 53.9s ±0% 9 20 $0.2401
✅ 111_pod_names_contain_service 37.7s ↓10% 7 15 $0.1771
✅ 12_job_crashing 43.3s ±0% 8 17 $0.2009
✅ 162_get_runbooks 44.0s ↓11% 7 18 $0.2237
✅ 176_network_policy_blocking_traffic_no_runbooks 38.7s ±0% 6 15 $0.1833
✅ 24_misconfigured_pvc 32.0s ↓16% 6 15 $0.1610
✅ 43_current_datetime_from_prompt 3.4s ±0% 1 — $0.0621
✅ 61_exact_match_counting 10.4s ±0% 3 3 $0.0860
Total 32.5s avg 5.8 avg 14.4 avg $1.4870

Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%.

Historical Comparison Details

Filter: excluding branch 'gcloud-mcp'

Status: Success - 23 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref gcloud-mcp -f markers=regression

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /last to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

benchmark, chain-of-causation, compaction, context_window, coralogix, counting, database, datadog, datetime, easy, embeds, grafana-dashboard, hard, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /last · /list

Signed-off-by: Arik Alon <alon.arik@gmail.com>
@github-actions

github-actions Bot commented Jan 2, 2026

Copy link
Copy Markdown
Contributor

✅ Results of HolmesGPT evals

Automatically triggered by commit 641743f on branch gcloud-mcp

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 34.1s ±0% 6 12 $0.1112
✅ 101_loki_historical_logs_pod_deleted 47.9s ±0% 8 14 $0.1447
✅ 111_pod_names_contain_service 37.4s ↓11% 6 14 $0.1082
✅ 12_job_crashing 41.9s ↓11% 7 16 $0.1325
✅ 162_get_runbooks 52.0s ±0% 7 19 $0.1743
✅ 176_network_policy_blocking_traffic_no_runbooks 42.2s ±0% 6 14 $0.1284
✅ 24_misconfigured_pvc 34.2s ↓11% 6 16 $0.1060
✅ 43_current_datetime_from_prompt 3.6s ±0% 1 — $0.0086
✅ 61_exact_match_counting 11.7s ±0% 3 3 $0.0326
Total 33.9s avg 5.6 avg 13.5 avg $0.9465

Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%.

Historical Comparison Details

Filter: excluding branch 'gcloud-mcp'

Status: Success - 23 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref gcloud-mcp -f markers=regression

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /last to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

benchmark, chain-of-causation, compaction, context_window, coralogix, counting, database, datadog, datetime, easy, embeds, grafana-dashboard, hard, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /last · /list

@github-actions

github-actions Bot commented Jan 2, 2026

Copy link
Copy Markdown
Contributor

✅ Results of HolmesGPT evals

Automatically triggered by commit 382e479 on branch gcloud-mcp

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 31.7s ±0% 6 13 $0.1683
✅ 101_loki_historical_logs_pod_deleted 35.1s ↓32% 5 12 $0.1626
✅ 111_pod_names_contain_service 35.2s ↓16% 6 15 $0.1673
✅ 12_job_crashing 43.4s ±0% 9 16 $0.2033
✅ 162_get_runbooks 46.0s ±0% 7 19 $0.2305
✅ 176_network_policy_blocking_traffic_no_runbooks 31.8s ↓23% 5 13 $0.1623
✅ 24_misconfigured_pvc 36.1s ±0% 7 16 $0.1775
✅ 43_current_datetime_from_prompt 3.3s ±0% 1 — $0.0621
✅ 61_exact_match_counting 10.2s ±0% 3 3 $0.0860
Total 30.3s avg 5.4 avg 13.4 avg $1.4198

Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%.

Historical Comparison Details

Filter: excluding branch 'gcloud-mcp'

Status: Success - 23 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref gcloud-mcp -f markers=regression

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /last to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

benchmark, chain-of-causation, compaction, context_window, coralogix, counting, database, datadog, datetime, easy, embeds, grafana-dashboard, hard, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /last · /list

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

📜 Review details

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 641743f and 382e479.

📒 Files selected for processing (2)
  • helm/holmes/templates/mcp-servers/gcp/_helpers.tpl
  • helm/holmes/values.yaml
🚧 Files skipped from review as they are similar to previous changes (1)
  • helm/holmes/templates/mcp-servers/gcp/_helpers.tpl
🧰 Additional context used
🧠 Learnings (1)
📚 Learning: 2025-12-29T08:35:37.678Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-12-29T08:35:37.678Z
Learning: Applies to tests/llm/**/*.yaml : Never use `:latest` container tags in test infrastructure - use specific versions like `grafana/grafana:12.3.1`

Applied to files:

  • helm/holmes/values.yaml
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (4)
  • GitHub Check: build (3.12)
  • GitHub Check: llm_evals
  • GitHub Check: build (3.11)
  • GitHub Check: build (3.10)

Comment thread helm/holmes/values.yaml Outdated
Signed-off-by: Arik Alon <alon.arik@gmail.com>
@github-actions

github-actions Bot commented Jan 2, 2026

Copy link
Copy Markdown
Contributor

✅ Results of HolmesGPT evals

Automatically triggered by commit ec3e66a on branch gcloud-mcp

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 32.2s ±0% 5 14 $0.1619
✅ 101_loki_historical_logs_pod_deleted 53.9s ±0% 9 18 $0.2284
✅ 111_pod_names_contain_service 43.4s ±0% 7 15 $0.1764
✅ 12_job_crashing 43.2s ±0% 7 16 $0.1831
✅ 162_get_runbooks 53.8s ±0% 8 20 $0.2452
✅ 176_network_policy_blocking_traffic_no_runbooks 39.2s ±0% 6 13 $0.1683
✅ 24_misconfigured_pvc 37.2s ±0% 7 13 $0.1619
✅ 43_current_datetime_from_prompt 3.5s ±0% 1 — $0.0621
✅ 61_exact_match_counting 11.5s ±0% 3 3 $0.0861
Total 35.3s avg 5.9 avg 14.0 avg $1.4734

Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%.

Historical Comparison Details

Filter: excluding branch 'gcloud-mcp'

Status: Success - 23 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref gcloud-mcp -f markers=regression

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /last to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

benchmark, chain-of-causation, compaction, context_window, coralogix, counting, database, datadog, datetime, easy, embeds, grafana-dashboard, hard, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /last · /list

Comment thread docs/data-sources/builtin-toolsets/gcp.md
Avi-Robusta
Avi-Robusta previously approved these changes Jan 5, 2026

@Avi-Robusta Avi-Robusta left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good
The docs seem unclear, left a comment there.

Weird edge case, a user could apply this config which will make a deployment with no containers which i think would crash helm installations.
Not sure if we want to handle it or not, there is no reason for a user to do this

holmes:
  mcpAddons:
    gcp:
      enabled: true
...
      gcloud:
        enabled: false
      observability:
        enabled: false
      storage:
        enabled: false

Signed-off-by: Arik Alon <alon.arik@gmail.com>
@github-actions

github-actions Bot commented Jan 5, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

📜 Run @ 07db59e (#20717205995)

✅ Results of HolmesGPT evals

Automatically triggered by commit 07db59e on branch gcloud-mcp

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 44.8s ↑31% 7 14 $0.1835
✅ 101_loki_historical_logs_pod_deleted 68.3s ±0% 10 22 $0.2540
✅ 111_pod_names_contain_service 39.8s ±0% 6 14 $0.1596
✅ 12_job_crashing 53.3s ±0% 9 19 $0.2282
✅ 162_get_runbooks 51.4s ±0% 7 18 $0.2193
✅ 176_network_policy_blocking_traffic_no_runbooks 51.0s ↑24% 8 16 $0.2013
✅ 24_misconfigured_pvc 43.8s ↑16% 7 17 $0.1755
✅ 43_current_datetime_from_prompt 3.8s ↑13% 1 — $0.0618
✅ 61_exact_match_counting 13.2s ↑13% 3 3 $0.0859
Total 41.0s avg 6.4 avg 15.4 avg $1.5690

Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%.

Historical Comparison Details

Filter: excluding branch 'gcloud-mcp'

Status: Success - 15 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)

✅ Results of HolmesGPT evals

Automatically triggered by commit 2dc1b97 on branch gcloud-mcp

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 32.9s ±0% 6 12 $0.1624
✅ 101_loki_historical_logs_pod_deleted 41.7s ↓33% 6 11 $0.1739
✅ 111_pod_names_contain_service 43.5s ±0% 8 18 $0.1881
✅ 12_job_crashing 39.3s ↓27% 7 15 $0.1801
✅ 162_get_runbooks 52.9s ±0% 8 20 $0.2375
✅ 176_network_policy_blocking_traffic_no_runbooks 35.2s ↓14% 6 14 $0.1716
✅ 24_misconfigured_pvc 32.7s ↓16% 6 16 $0.1585
✅ 43_current_datetime_from_prompt 3.6s ±0% 1 — $0.0618
✅ 61_exact_match_counting 11.3s ±0% 3 3 $0.0859
Total 32.6s avg 5.7 avg 13.6 avg $1.4198

Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%.

Historical Comparison Details

Filter: excluding branch 'gcloud-mcp'

Status: Success - 26 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref gcloud-mcp -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

benchmark, chain-of-causation, compaction, context_window, coralogix, counting, database, datadog, datetime, easy, elasticsearch, embeds, grafana-dashboard, hard, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref gcloud-mcp -f markers=regression -f filter=

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 0

♻️ Duplicate comments (2)
docs/data-sources/builtin-toolsets/gcp.md (2)

370-382: Add blank lines between bold text and list items (MkDocs rendering requirement).

Per the coding guidelines, bold text must be followed by a blank line before any list items for proper MkDocs rendering. Currently missing blank lines at:

  • Line 370: **What's Included:** → needs blank line before line 371's bullet
  • Line 378: **Security Boundaries:** → needs blank line before line 379's bullet
🔎 Proposed fix
  **What's Included:**
+ 
  - ✅ Complete audit log visibility (who changed what)
  - ✅ Full networking troubleshooting (firewalls, load balancers, SSL)
  - ✅ Database and BigQuery metadata (schemas, configurations)
  - ✅ Security findings and IAM analysis
  - ✅ Container and Kubernetes visibility
  - ✅ Monitoring, logging, and tracing

  **Security Boundaries:**
+ 
  - ❌ NO actual data access (cannot read storage objects or BigQuery data)
  - ❌ NO secret values (only metadata)
  - ❌ NO write permissions

Based on coding guidelines: Add blank line between header/bold text and a list in MkDocs documentation files, otherwise lists won't render properly.


453-481: Add blank lines between subheaders and code blocks in Example Usage section.

The coding guideline requires blank lines between headers and lists/content blocks for proper MkDocs rendering. This is missing for all five example subsections (lines 453, 459, 465, 471, 477). Additionally, the example code blocks should specify a language identifier (use text for plain-text examples).

🔎 Proposed fix
  ### Investigating Deleted Pod Logs
+ 
- ```
+ ```text
  "Show me logs from the payment-service pod that was OOMKilled this morning"

Cross-Project Resource Discovery

  • "List all GKE clusters across our dev, staging, and prod projects"
    

    Audit Trail Investigation

  • "Who modified the firewall rules in the last 24 hours?"
    

    Storage Access Issues

  • "Why is my application getting 403 errors accessing the data-bucket?"
    

    SSL Certificate Problems

  • "Check the SSL certificates on our load balancers"
    
</details>

Based on coding guidelines: Add blank line between header/bold text and a list in MkDocs documentation files, otherwise lists won't render properly.

</blockquote></details>

</blockquote></details>

<details>
<summary>🧹 Nitpick comments (1)</summary><blockquote>

<details>
<summary>helm/holmes/values.yaml (1)</summary><blockquote>

`219-248`: **Consider adding CPU limits and documenting memory limit variation.**

The GCP MCP services have memory limits but no CPU limits, unlike the AWS MCP addon (lines 98-104) which sets both. Additionally, gcloud has a 1Gi memory limit while observability and storage have 512Mi, without explanation.



<details>
<summary>Considerations</summary>

1. **CPU limits**: Adding CPU limits prevents unbounded CPU usage:
   - Provides predictable resource consumption
   - Prevents noisy neighbor issues
   - Aligns with AWS MCP addon pattern (line 104: `cpu: "500m"`)

2. **Memory variation**: The 2x memory allocation for gcloud (1Gi vs 512Mi) might be intentional due to CLI overhead, but a comment would clarify the reasoning.

Example adjustment:

```yaml
gcloud:
  enabled: true
  image: "gcloud-cli-mcp:1.0.7"
  port: 8000
  resources:
    requests:
      memory: "256Mi"
      cpu: "100m"
    limits:
      memory: "1Gi"  # Higher limit due to gcloud CLI overhead
      cpu: "500m"
📜 Review details

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 382e479 and 07db59e.

📒 Files selected for processing (2)
  • docs/data-sources/builtin-toolsets/gcp.md
  • helm/holmes/values.yaml
🧰 Additional context used
📓 Path-based instructions (1)
docs/**/*.md

📄 CodeRabbit inference engine (CLAUDE.md)

Add blank line between header/bold text and a list in MkDocs documentation files, otherwise lists won't render properly

Files:

  • docs/data-sources/builtin-toolsets/gcp.md
🧠 Learnings (1)
📚 Learning: 2026-01-05T11:14:20.208Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: CLAUDE.md:0-0
Timestamp: 2026-01-05T11:14:20.208Z
Learning: Applies to docs/**/*.md : Add blank line between header/bold text and a list in MkDocs documentation files, otherwise lists won't render properly

Applied to files:

  • docs/data-sources/builtin-toolsets/gcp.md
🪛 LanguageTool
docs/data-sources/builtin-toolsets/gcp.md

[grammar] ~10-~10: Ensure spelling is correct
Context: ...on for setup instructions. The GCP MCP addon consists of three specialized servers: ...

(QB_NEW_EN_ORTHOGRAPHY_ERROR_IDS_1)

🪛 markdownlint-cli2 (0.18.1)
docs/data-sources/builtin-toolsets/gcp.md

344-344: Code block style
Expected: indented; Actual: fenced

(MD046, code-block-style)


399-399: Code block style
Expected: indented; Actual: fenced

(MD046, code-block-style)


454-454: Fenced code blocks should have a language specified

(MD040, fenced-code-language)


454-454: Code block style
Expected: indented; Actual: fenced

(MD046, code-block-style)


460-460: Fenced code blocks should have a language specified

(MD040, fenced-code-language)


460-460: Code block style
Expected: indented; Actual: fenced

(MD046, code-block-style)


466-466: Fenced code blocks should have a language specified

(MD040, fenced-code-language)


466-466: Code block style
Expected: indented; Actual: fenced

(MD046, code-block-style)


472-472: Fenced code blocks should have a language specified

(MD040, fenced-code-language)


472-472: Code block style
Expected: indented; Actual: fenced

(MD046, code-block-style)


478-478: Fenced code blocks should have a language specified

(MD040, fenced-code-language)


478-478: Code block style
Expected: indented; Actual: fenced

(MD046, code-block-style)


486-486: Code block style
Expected: indented; Actual: fenced

(MD046, code-block-style)


498-498: Code block style
Expected: indented; Actual: fenced

(MD046, code-block-style)


505-505: Code block style
Expected: indented; Actual: fenced

(MD046, code-block-style)

⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (5)
  • GitHub Check: build
  • GitHub Check: llm_evals
  • GitHub Check: build (3.11)
  • GitHub Check: build (3.10)
  • GitHub Check: build (3.12)
🔇 Additional comments (1)
helm/holmes/values.yaml (1)

254-259: GCP llmInstructions map structure is correctly handled by Helm templates.

The GCP addon's map structure for llmInstructions (with gcloud, observability, storage keys) is properly referenced in the templates. The helper functions in helm/holmes/templates/mcp-servers/gcp/_helpers.tpl correctly check and use .Values.mcpAddons.gcp.llmInstructions.gcloud, .Values.mcpAddons.gcp.llmInstructions.observability, and .Values.mcpAddons.gcp.llmInstructions.storage with fallback defaults, ensuring proper template rendering regardless of whether custom instructions are provided.

moshemorad
moshemorad previously approved these changes Jan 5, 2026
Avi-Robusta
Avi-Robusta previously approved these changes Jan 5, 2026
@arikalon1
arikalon1 dismissed stale reviews from Avi-Robusta and moshemorad via 2dc1b97 January 5, 2026 15:59

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 0

🧹 Nitpick comments (1)
helm/holmes/values.yaml (1)

213-249: Individual service configurations are well-structured.

Each GCP MCP service (gcloud, observability, storage) has dedicated resources, ports, and images. The higher memory limit for gcloud (1Gi vs 512Mi) is appropriate given it's the general-purpose CLI tool.

Note that all three services default to enabled: true. While this provides full functionality out of the box, users may want to disable services they don't need to conserve cluster resources. Consider if this is the desired behavior or if a more conservative default (e.g., only enabling gcloud) would be better.

📜 Review details

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 07db59e and 2dc1b97.

📒 Files selected for processing (3)
  • docs/data-sources/builtin-toolsets/.nav.yml
  • helm/holmes/templates/toolset-config.yaml
  • helm/holmes/values.yaml
🧰 Additional context used
🪛 YAMLlint (1.37.1)
helm/holmes/templates/toolset-config.yaml

[error] 1-1: syntax error: expected the node content, but found '-'

(syntax)

⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (4)
  • GitHub Check: llm_evals
  • GitHub Check: build (3.11)
  • GitHub Check: build (3.10)
  • GitHub Check: build (3.12)
🔇 Additional comments (7)
helm/holmes/templates/toolset-config.yaml (2)

1-1: LGTM! Condition properly extended for GCP addon.

The top-level condition correctly includes GCP MCP addon, maintaining consistency with AWS, MariaDB, and Azure patterns.


48-84: GCP MCP server configuration is well-structured and correct.

Port numbers are properly configured in values.yaml (gcloud: 8000, observability: 8001, storage: 8002) and correctly referenced in the template with explicit int casting for type safety. The SSE mode is intentional and consistent with the /sse endpoint paths used by the GCP MCP servers. Each service is properly guarded by individual enable flags, and the pattern matches the established implementation for other MCP addons.

docs/data-sources/builtin-toolsets/.nav.yml (1)

2-37: LGTM! GCP navigation entry added correctly.

The GCP (MCP) navigation entry is properly positioned in alphabetical order, and the indentation has been standardized to 2 spaces for consistency. The referenced gcp.md file issue from the previous review has been addressed.

helm/holmes/values.yaml (4)

162-181: Excellent documentation for GCP MCP setup.

The authentication and multi-project support sections provide clear, actionable instructions with concrete commands. This will significantly improve the user experience for configuring the GCP MCP addon.


182-206: Core configuration is well-structured and secure.

The configuration properly defaults to enabled: false for safe deployment, and networkPolicy.enabled: true for security. The note about Workload Identity not working with gcloud CLI (line 200-202) is valuable context for users.

Previous review comments regarding imagePullPolicy and placeholder URLs have been properly addressed.


207-212: LGTM! Config section provides appropriate defaults and flexibility.

The configuration allows for dynamic project/region selection via flags while providing sensible defaults. The comments clearly explain when each field is needed.


250-256: LGTM! Per-service LLM instruction customization provides good flexibility.

The ability to customize instructions for each service (gcloud, observability, storage) while defaulting to built-in instructions is well-designed. This aligns with the template's helper function usage.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants