Skip to content

Remove deprecated workload health - #1380

Merged
arikalon1 merged 9 commits into
masterfrom
claude/merge-master-gbQOa
Jan 31, 2026
Merged

arikalon1 merged 9 commits into
masterfrom
claude/merge-master-gbQOa

Conversation

@aantn

@aantn aantn commented Jan 19, 2026 •

Copy link
Copy Markdown
Collaborator

Summary by CodeRabbit

  • Documentation

    • Removed workload health endpoints from the API reference and updated the overview wording.
  • Tests

    • Removed the workload health test suite, fixtures, and associated test tooling, mocks, and test-case data.
  • Chores

    • Removed workload health endpoints, request/response models, prompt templates, data-retrieval helpers, and reporting/counting for workload health.

✏️ Tip: You can customize this high-level summary in your review settings.

aantn and others added 4 commits December 30, 2025 12:17
Signed-off-by: Codex <codex@openai.com>
Signed-off-by: Codex <codex@openai.com>
Signed-off-by: Codex <codex@openai.com>
Resolve merge conflicts by keeping HEAD's removal of workload health endpoints:
- Remove WorkloadHealthRequest and WorkloadHealthChatRequest imports
- Remove /api/workload_health_check and /api/workload_health_chat endpoints
- Delete tests/llm/test_workload_health.py test file
- Remove test_workload_health from LLM_TEST_TYPES

Keep ToolCallConversationResult import in conversations.py as it's used elsewhere.

Signed-off-by: Claude <noreply@anthropic.com>
@linux-foundation-easycla

linux-foundation-easycla Bot commented Jan 19, 2026 •

Copy link
Copy Markdown

CLA Signed

The committers listed above are authorized under a signed CLA.

  • ✅ login: aantn / name: Natan Yellin (f2a1770)

@netlify

netlify Bot commented Jan 19, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit f2a1770
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/697db2be974a410008cd06c2
😎 Deploy Preview https://deploy-preview-1380--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@coderabbitai

coderabbitai Bot commented Jan 19, 2026 •

Copy link
Copy Markdown
Contributor

Walkthrough

This PR removes the workload health feature across the codebase: API endpoints, core models, conversation builders, Supabase DAL method, prompt templates, LLM fixtures/tests, and related reporting/test-infrastructure hooks.

Changes

Cohort / File(s) Summary
API & Documentation
server.py, docs/reference/http-api.md
Removed /api/workload_health_check and /api/workload_health_chat endpoints; updated API overview and removed workload-health imports/types.
Core Models & Conversations
holmes/core/models.py, holmes/core/conversations.py
Deleted WorkloadHealthRequest, WorkloadHealthInvestigationResult, WorkloadHealthChatRequest, workload_health_structured_output, and build_workload_health_chat_messages.
Data Access Layer
holmes/core/supabase_dal.py
Removed SupabaseDal.get_workload_issues method and its queries/processing.
Prompt Templates
holmes/plugins/prompts/...workload*.jinja2
Deleted kubernetes_workload_ask.jinja2 and kubernetes_workload_chat.jinja2 templates.
LLM Test Framework & Fixtures
tests/llm/conftest.py, tests/llm/test_workload_health.py, tests/llm/fixtures/test_workload_health/*, tests/llm/utils/mock_dal.py, tests/llm/utils/test_case_utils.py
Removed workload_health from LLM_TEST_TYPES; deleted workload-health test module, fixtures, mock DAL stub, test-case helpers and related loaders.
Unit & Integration Tests / Reporting
tests/test_server_endpoints.py, tests/core/test_prompt.py, tests/test_ai_safety_prompt.py, tests/llm/utils/reporting/github_reporter.py
Deleted workload-health-specific tests and removed workload_health counters/summary from GitHub report generation.

Sequence Diagram(s)

(omitted)

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~22 minutes

Possibly related PRs

Suggested reviewers

  • Sheeproid
🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The PR title 'Remove deprecated workload health' directly and clearly describes the main change—it removes workload health functionality across the codebase including API endpoints, models, prompts, tests, and documentation.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions

github-actions Bot commented Jan 19, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

📜 Run @ 9d90443 (#21492814514)

✅ Results of HolmesGPT evals

Automatically triggered by commit 9d90443 on branch claude/merge-master-gbQOa

View workflow logs

⚠️ No eval report was generated.

📜 Run @ e19dcb0 (#21397191995)

✅ Results of HolmesGPT evals

Automatically triggered by commit e19dcb0 on branch claude/merge-master-gbQOa

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 29.4s ±0% 5 11 $0.2298
✅ 101_loki_historical_logs_pod_deleted 51.3s ↓11% 7 14 $0.3050
✅ 111_pod_names_contain_service 22.6s ±0% 4 8 $0.1821
✅ 12_job_crashing 34.0s ±0% 5 15 $0.2586
✅ 162_get_runbooks 38.7s ±0% 6 11 $0.2710
✅ 176_network_policy_blocking_traffic_no_runbooks 38.7s ↓16% 6 16 $0.2725
✅ 24_misconfigured_pvc 34.5s ±0% 6 14 $0.2473
✅ 43_current_datetime_from_prompt 3.3s ↓18% 1 — $0.1019
✅ 61_exact_match_counting 15.7s ±0% 4 4 $0.1589
Total 29.8s avg 4.9 avg 11.6 avg $2.0272

Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%.

Historical Comparison Details

Filter: excluding branch 'claude/merge-master-gbQOa'

Status: Success - 12 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 Run @ 69c1c64 (#21227284890)

✅ Results of HolmesGPT evals

Automatically triggered by commit 69c1c64 on branch claude/merge-master-gbQOa

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 34.3s ±0% 6 14 $0.1755
✅ 101_loki_historical_logs_pod_deleted 87.5s ↑56% 11 26 $0.3634
✅ 111_pod_names_contain_service 41.8s ±0% 8 18 $0.1936
✅ 12_job_crashing 38.0s ↓22% 7 15 $0.1804
✅ 162_get_runbooks 44.0s ±0% 7 17 $0.2105
✅ 176_network_policy_blocking_traffic_no_runbooks 41.1s ±0% 7 14 $0.1757
✅ 24_misconfigured_pvc 41.1s ↑11% 8 18 $0.1930
✅ 43_current_datetime_from_prompt 3.3s ±0% 1 — $0.0617
✅ 61_exact_match_counting 10.6s ±0% 3 3 $0.0856
Total 38.0s avg 6.4 avg 15.6 avg $1.6392

Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%.

Historical Comparison Details

Filter: excluding branch 'claude/merge-master-gbQOa'

Status: Success - 29 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 Run @ 5cbf1a2 (#21227188757)

✅ Results of HolmesGPT evals

Automatically triggered by commit 5cbf1a2 on branch claude/merge-master-gbQOa

View workflow logs

⚠️ No eval report was generated.

📜 Run @ 2ad5afc (#21131599992)

✅ Results of HolmesGPT evals

Automatically triggered by commit 2ad5afc on branch claude/merge-master-gbQOa

View workflow logs

⚠️ No eval report was generated.


✅ Results of HolmesGPT evals

Automatically triggered by commit f2a1770 on branch claude/merge-master-gbQOa

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 106.7s ↑236% 5 11 $0.2309
✅ 101_loki_historical_logs_pod_deleted 49.9s ↑29% 7 12 $0.2785
✅ 111_pod_names_contain_service 32.3s ±0% 5 12 $0.2327
✅ 12_job_crashing 31.4s ±0% 5 11 $0.2354
✅ 162_get_runbooks 36.9s ±0% 6 11 $0.2592
✅ 176_network_policy_blocking_traffic_no_runbooks 41.4s ±0% 7 17 $0.2913
✅ 24_misconfigured_pvc 35.9s ±0% 6 16 $0.2536
✅ 43_current_datetime_from_prompt 5.7s ↑11% 1 — $0.1063
✅ 61_exact_match_counting 16.3s ±0% 4 4 $0.1589
Total 39.6s avg 5.1 avg 11.8 avg $2.0468

Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%.

Historical Comparison Details

Filter: excluding branch 'claude/merge-master-gbQOa'

Status: Success - 9 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/merge-master-gbQOa -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, easy, elasticsearch, embeds, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/merge-master-gbQOa -f markers=regression -f filter=

@github-actions

github-actions Bot commented Jan 19, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker image ready for e795a40 (built in 7m 28s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use this tag to pull the image for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:e795a40
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:e795a40 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:e795a40
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:e795a40

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:e795a40

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:e795a40

aantn and others added 5 commits January 21, 2026 23:58
These functions were incorrectly removed during merge conflict resolution.
They are new features from master unrelated to workload health removal.

Signed-off-by: Claude <noreply@anthropic.com>
Resolve merge conflicts:
- server.py: Remove workload health endpoints (intentionally removed in HEAD)
- tests/test_server_endpoints.py: Remove workload health test, keep new
  TestExtractPassthroughHeaders test class

Maintain workload health removal while incorporating new features:
- Request context passthrough for all endpoints
- New bash validation tests and kubectl-run toolset
- Scheduled prompts feature
- Various toolset improvements

Signed-off-by: Claude <noreply@anthropic.com>

@arikalon1 arikalon1 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@arikalon1
arikalon1 merged commit 40d1839 into master Jan 31, 2026
15 of 17 checks passed
@arikalon1
arikalon1 deleted the claude/merge-master-gbQOa branch January 31, 2026 10:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants