Skip to content

ROB-3773 bump github app mcp and build for arm/amd - #1986

Merged
naomi-robusta merged 2 commits into
masterfrom
ROB-3773-github-mcp-works-only-on-arm2
May 3, 2026
Merged

naomi-robusta merged 2 commits into
masterfrom
ROB-3773-github-mcp-works-only-on-arm2

Conversation

@naomi-robusta

@naomi-robusta naomi-robusta commented May 3, 2026 •

Copy link
Copy Markdown
Collaborator

Summary by CodeRabbit

  • Chores
    • Updated GitHub integration addon image to version 1.0.2. This refreshes the deployed addon image reference so environments use the updated release.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This repository is configured for manual code reviews. Comment @claude review to trigger a review and subscribe this PR to future pushes, or @claude review once for a one-time review.

Tip: disable this comment in your organization's Code Review settings.

@naomi-robusta
naomi-robusta enabled auto-merge (squash) May 3, 2026 07:27
@github-actions

github-actions Bot commented May 3, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

📜 #1 · Run @ __2504295__ (#25273099199) — May 3, 07:34 UTC

✅ Results of HolmesGPT evals

Automatically triggered by commit 2504295 on branch ROB-3773-github-mcp-works-only-on-arm2

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 31.6s 5 10 $0.2443 104,825 102,721 23,336 2,104 893 78,020 24,701 — —
✅ 101_loki_historical_logs_pod_deleted 45.1s 6 11 $0.2826 133,158 130,458 25,255 2,700 884 104,167 26,291 — —
✅ 112_find_pvcs_by_uuid 19.2s 4 3 $0.1795 75,740 74,721 20,277 1,019 312 54,432 20,289 — —
✅ 12_job_crashing 38.4s 5 14 $0.2713 113,566 111,278 25,782 2,288 695 83,216 28,062 — —
✅ 176_network_policy_blocking_traffic_no_skills 42.1s 7 15 $0.2988 155,264 152,400 25,903 2,864 749 126,039 26,361 — —
✅ 227_count_configmaps_per_namespace[0] 20.3s 4 9 $0.1912 77,611 76,373 20,967 1,238 585 55,076 21,297 — —
✅ 243_pod_names_contain_service 31.5s 5 9 $0.2278 100,082 98,201 22,155 1,881 717 75,104 23,097 — —
✅ 24_misconfigured_pvc 30.3s 5 12 $0.2365 102,426 100,340 22,624 2,086 588 76,805 23,535 — —
✅ 43_current_datetime_from_prompt 4.1s 1 — $0.1092 17,105 16,982 16,982 123 123 0 16,982 — —
✅ 51_logs_summarize_errors 22.3s 4 5 $0.1893 77,654 76,455 21,050 1,199 446 55,393 21,062 — —
✅ 61_exact_match_counting 10.5s 3 3 $0.1394 53,044 52,657 17,980 387 240 34,666 17,991 — —
Total 26.9s avg 4.5 avg 9.1 avg $2.3699 1,010,475 992,586 25,903 17,889 893 742,918 249,668 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 73 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)

✅ Results of HolmesGPT evals

Automatically triggered by commit 0998fdf on branch ROB-3773-github-mcp-works-only-on-arm2

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 33.4s 5 10 $0.2463 105,267 103,081 23,637 2,186 605 78,520 24,561 — —
✅ 101_loki_historical_logs_pod_deleted 49.8s 7 11 $0.2801 148,968 146,294 24,151 2,674 518 121,861 24,433 — —
✅ 112_find_pvcs_by_uuid 16.8s 3 3 $0.1771 60,482 59,545 21,536 937 544 37,998 21,547 — —
✅ 12_job_crashing 33.1s 5 13 $0.2600 112,212 110,146 25,002 2,066 585 83,022 27,124 — —
✅ 176_network_policy_blocking_traffic_no_skills 38.9s 5 13 $0.2808 116,294 113,642 26,475 2,652 744 85,967 27,675 — —
✅ 227_count_configmaps_per_namespace[0] 24.4s 5 9 $0.2018 95,218 93,958 20,956 1,260 582 72,365 21,593 — —
✅ 243_pod_names_contain_service 30.3s 4 8 $0.2150 79,833 77,970 21,791 1,863 835 55,261 22,709 — —
✅ 24_misconfigured_pvc 37.0s 6 14 $0.2576 122,035 119,676 22,992 2,359 797 95,230 24,446 — —
✅ 43_current_datetime_from_prompt 4.6s 1 — $0.1093 17,108 16,982 16,982 126 126 0 16,982 — —
✅ 51_logs_summarize_errors 21.8s 4 5 $0.1895 77,885 76,715 21,187 1,170 415 55,516 21,199 — —
✅ 61_exact_match_counting 8.7s 2 1 $0.1246 34,698 34,408 17,417 290 222 16,981 17,427 — —
Total 27.2s avg 4.3 avg 8.7 avg $2.3420 970,000 952,417 26,475 17,583 835 702,721 249,696 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 73 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref ROB-3773-github-mcp-works-only-on-arm2 -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
tags: regression

Or with more options (one per line):

/eval
model: gpt-4o
tags: regression
id: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
tags: regression
Option Description
model Model(s) to test (default: same as automatic runs)
tags Pytest tags / markers (no default - runs all tests!)
id Eval ID / pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

Option 3: Add PR labels to include extra evals (applies to both automatic runs and /eval comments):

Label Effect
evals-tag-<name> Run tests with tag <name> alongside regression
evals-id-<name> Run a specific eval by test ID
evals-model-<name> Override the model (use model list name, e.g. sonnet-4.5)

Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5

🏷️ Valid tags

benchmark, chain-of-causation, compaction, confluence, context_window, conversation_worker, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, images, integration, kafka, kubernetes, leaked-information, logs, loki, manual, mcp, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, skills, slackbot, storage, token-limit, toolset-limitation, traces, transparency, victorialogs

🤖 Valid models

deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref ROB-3773-github-mcp-works-only-on-arm2 -f markers=regression -f filter=

@coderabbitai

coderabbitai Bot commented May 3, 2026 •

Copy link
Copy Markdown
Contributor

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 9d84ef1c-66a1-47cb-a958-60148ffef09b

📥 Commits

Reviewing files that changed from the base of the PR and between 2504295 and 0998fdf.

📒 Files selected for processing (1)
  • helm/holmes/values.yaml

Walkthrough

GitHub MCP addon image version updated in Helm values: mcpAddons.github.githubApp.image changed from github-app-mcp:1.0.0 to github-app-mcp:1.0.2 in helm/holmes/values.yaml.

Changes

Helm Values Configuration

Layer / File(s) Summary
Functional Change
helm/holmes/values.yaml
Bumped mcpAddons.github.githubApp.image from github-app-mcp:1.0.0 to github-app-mcp:1.0.2.

Estimated code review effort

🎯 1 (Trivial) | ⏱️ ~2 minutes

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title directly relates to the main change: bumping the GitHub App MCP image version from 1.0.0 to 1.0.2, with the mention of building for arm/amd indicating architectural support expansion.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share
Review rate limit: 6/8 reviews remaining, refill in 10 minutes and 10 seconds.

Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions

github-actions Bot commented May 3, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker images ready for f9cf8a5d (built in 4m 43s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use these tags to pull the images for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:f9cf8a5d
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:f9cf8a5d me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:f9cf8a5d
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:f9cf8a5d
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:f9cf8a5d
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:f9cf8a5d me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:f9cf8a5d
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:f9cf8a5d

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:f9cf8a5d \
  --set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set operator.image=holmes-operator-dev:f9cf8a5d

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:f9cf8a5d \
  --set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.operator.image=holmes-operator-dev:f9cf8a5d

@netlify

netlify Bot commented May 3, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit 0998fdf
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/69f6fa18678d980007984060
😎 Deploy Preview https://deploy-preview-1986--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
helm/holmes/values.yaml (1)

552-556: 💤 Low value

Consider block scalar (>-) instead of backslash-newline continuation for readability.

When you end a line with \, YAML treats the next line as a continuation of the same string — no space, no newline. The values parse correctly, but putting \ at the end of a line to escape the newline is valid in double-quoted strings, though the block scalar syntax is usually a nicer approach for wrapping long strings.

Block scalars with >- (folded, strip trailing newline) are the conventional Helm pattern for long single-line config values:

♻️ Suggested alternative using folded block scalars
-      allowedCommands: "edit,patch,delete,scale,rollout,cordon,uncordon,drain,taint,l\
-        abel,annotate"
+      allowedCommands: >-
+        edit,patch,delete,scale,rollout,cordon,uncordon,drain,taint,label,annotate
       # Comma-separated list of blocked flags
-      dangerousFlags: "--kubeconfig,--context,--cluster,--user,--token,--as,--as-grou\
-        p,--as-uid"
+      dangerousFlags: >-
+        --kubeconfig,--context,--cluster,--user,--token,--as,--as-group,--as-uid

The same applies to enabledTools at line 613:

-      enabledTools: "confluence_search,confluence_get_page,confluence_get_page_conten\
-        t,confluence_get_comments"
+      enabledTools: >-
+        confluence_search,confluence_get_page,confluence_get_page_content,confluence_get_comments
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@helm/holmes/values.yaml` around lines 552 - 556, The YAML uses
backslash-newline continuation inside double-quoted strings for long values (see
allowedCommands and dangerousFlags, also enabledTools) which is harder to read
and error-prone; replace those quoted, backslash-continued strings with folded
block scalars (e.g., use >- to fold into a single-line string) for each value so
the long comma-separated lists remain single-line when parsed but are readable
in the file — update the allowedCommands, dangerousFlags (and enabledTools)
entries to use folded block scalars instead of backslash escapes.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Nitpick comments:
In `@helm/holmes/values.yaml`:
- Around line 552-556: The YAML uses backslash-newline continuation inside
double-quoted strings for long values (see allowedCommands and dangerousFlags,
also enabledTools) which is harder to read and error-prone; replace those
quoted, backslash-continued strings with folded block scalars (e.g., use >- to
fold into a single-line string) for each value so the long comma-separated lists
remain single-line when parsed but are readable in the file — update the
allowedCommands, dangerousFlags (and enabledTools) entries to use folded block
scalars instead of backslash escapes.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 1c863e5e-6b3b-4ea5-90be-ba6912b25b20

📥 Commits

Reviewing files that changed from the base of the PR and between cf6ddb7 and 2504295.

📒 Files selected for processing (1)
  • helm/holmes/values.yaml

@naomi-robusta
naomi-robusta merged commit 89ca617 into master May 3, 2026
18 of 19 checks passed
@naomi-robusta
naomi-robusta deleted the ROB-3773-github-mcp-works-only-on-arm2 branch May 3, 2026 07:42
@coderabbitai coderabbitai Bot mentioned this pull request Jun 4, 2026
2 tasks
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants