Skip to content

Add generation-based re-execution and rerun annotation support - #1862

Open
aantn wants to merge 7 commits into
masterfrom
claude/kubectl-apply-operator-cr-vkzvj
Open

aantn wants to merge 7 commits into
masterfrom
claude/kubectl-apply-operator-cr-vkzvj

Conversation

@aantn

@aantn aantn commented Mar 31, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

This PR implements generation-based re-execution for HealthCheck resources, allowing checks to automatically re-run when the spec is modified. It also adds support for manual re-execution via a holmesgpt.dev/rerun annotation, and tracks the last processed generation in the status via observedGeneration.

Key Changes

  • Generation-based re-execution: When metadata.generation differs from status.observedGeneration, the check automatically re-executes. This leverages Kubernetes' built-in generation counter that increments whenever the spec changes.

  • Rerun annotation support: Users can set the holmesgpt.dev/rerun=true annotation to manually trigger re-execution even without spec changes. The annotation is automatically cleared after processing.

  • Refactored execution logic: Extracted common execution logic into _execute_healthcheck() shared by both create and update handlers, reducing code duplication.

  • Status tracking: Added observedGeneration field to HealthCheckStatus to track which generation was last processed, preventing infinite retry loops on failures.

  • Update handler enhancement: The on_healthcheck_update handler now implements two re-execution triggers (in order):

    1. Generation mismatch (spec changed)
    2. Rerun annotation (manual trigger)
  • CRD schema update: Added observedGeneration field definition to the HealthCheck CRD schema with documentation.

  • Comprehensive test coverage: Added test cases for:

    • Generation-based re-triggering
    • Skipping execution when generation matches
    • Rerun annotation triggering and clearing

Implementation Details

  • The _execute_healthcheck() function now accepts an optional generation parameter that gets stored as observedGeneration in the status, even on failure.
  • The rerun annotation is cleared via a separate _clear_rerun_annotation() helper that patches the resource metadata.
  • All status updates (Completed, Failed) now include the observedGeneration to prevent re-execution loops.
  • Tests use a _make_body() helper to construct proper HealthCheck resource bodies with metadata.

https://claude.ai/code/session_01QY6zsEHPW9CfSKJBtE6WCd

Summary by CodeRabbit

  • New Features

    • HealthChecks now record observedGeneration in status (set on success or failure); operator re-runs when spec generation changes or when rerun annotation is used. Rerun annotation is cleared after execution.
  • Documentation

    • Guides updated with generation-based re-execution, annotation-based re-run flows, Helm/ArgoCD examples, and verification tips.
  • Tests

    • Expanded tests for create, update (generation-triggered), annotation-triggered reruns, and observedGeneration in status.

claude added 2 commits March 31, 2026 16:32
When a user runs `kubectl apply -f` with a modified spec, Kubernetes
increments metadata.generation. The operator now tracks
status.observedGeneration and re-executes the health check whenever
generation != observedGeneration, giving one-shot-per-apply semantics.

Changes:
- Add observedGeneration field to CRD status schema and Pydantic model
- Extract shared _execute_healthcheck() from create/update handlers
- on.create sets observedGeneration after execution
- on.update re-runs when generation != observedGeneration (spec change)
- Preserve existing annotation-based rerun as fallback for same-spec reruns
- Add tests for generation-based trigger, skip, and annotation fallback

https://claude.ai/code/session_01QY6zsEHPW9CfSKJBtE6WCd
Signed-off-by: Claude <noreply@anthropic.com>
After the operator processes a rerun annotation, it now patches the
resource to remove the annotation (set to null). This lets users
simply re-add the annotation to trigger another run, without needing
to manually remove and re-add it.

https://claude.ai/code/session_01QY6zsEHPW9CfSKJBtE6WCd
Signed-off-by: Claude <noreply@anthropic.com>

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This repository is configured for manual code reviews. Comment @claude review to trigger a review and subscribe this PR to future pushes, or @claude review once for a one-time review.

Tip: disable this comment in your organization's Code Review settings.

@aantn

aantn commented Mar 31, 2026

Copy link
Copy Markdown
Collaborator Author

@claude review

@coderabbitai

coderabbitai Bot commented Mar 31, 2026 •

Copy link
Copy Markdown
Contributor

Walkthrough

Added observedGeneration to HealthCheck status and wired generation- and annotation-based re-execution: CRD, data model, status utilities, handlers updated; create/update handlers now delegate execution to a shared _execute_healthcheck and clear the rerun annotation after use.

Changes

Cohort / File(s) Summary
CRD Schema
helm/holmes/crds/healthcheck.yaml
Added status.observedGeneration: integer (int64) to HealthCheck CRD status schema.
Data Model
holmes_operator/models.py
Added observedGeneration: Optional[int] to HealthCheckStatus.
Handler Implementation
holmes_operator/handlers/healthcheck.py
Added _execute_healthcheck(...) and _clear_rerun_annotation(...); create handler extracts metadata.generation and body and delegates to _execute_healthcheck; update handler re-executes when metadata.generation != status.observedGeneration or when holmesgpt.dev/rerun="true", and clears the rerun annotation after execution.
Status Utilities
holmes_operator/utils.py
Extended update_healthcheck_status, set_healthcheck_completed, and set_healthcheck_failed with optional observed_generation to include observedGeneration in status patches.
Tests
tests/holmes_operator/test_healthcheck_component.py
Added _make_body helper; updated create tests to include metadata.generation and assert status.observedGeneration; added TestHealthCheckUpdate covering generation-based re-triggering, annotation-based rerun, and annotation clearing.
Documentation
docs/operator/health-checks.md, docs/operator/deployment-verification.md
Documented status.observedGeneration, automatic re-execution on spec changes, and holmesgpt.dev/rerun: "true" annotation behavior with Helm/ArgoCD/CI examples and updated guidance.

Sequence Diagram

sequenceDiagram
    participant K8s as Kubernetes API
    participant Handler as HealthCheck Handler
    participant Holmes as Holmes API
    participant StatusAPI as Status Patch API

    K8s->>Handler: create/update event (body, metadata.generation)
    Handler->>Handler: read status.observedGeneration\ncheck holmesgpt.dev/rerun annotation
    alt generation changed OR rerun annotation present
        Handler->>Handler: call _execute_healthcheck(spec, name, ns, uid, generation, body)
        Handler->>Holmes: invoke Holmes API (check execution)
        Holmes-->>Handler: result / error
        Handler->>StatusAPI: patch status (include observedGeneration = generation)
        StatusAPI->>K8s: apply status patch
        alt rerun annotation was present
            Handler->>K8s: patch to clear holmesgpt.dev/rerun annotation
            K8s->>Handler: patch result
        end
    else no action
        Handler-->>K8s: skip execution
    end
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Suggested reviewers

  • Sheeproid
  • arikalon1
🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately and concisely describes the main changes: adding generation-based re-execution and rerun annotation support to the HealthCheck operator.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@netlify

netlify Bot commented Mar 31, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit ee321a6
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/69cc1e6f648fea0008a2e78f
😎 Deploy Preview https://deploy-preview-1862--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@github-actions

github-actions Bot commented Mar 31, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

📜 #4 · Run @ __fb882c1__ (#23814824032) — Mar 31, 19:27 UTC

✅ Results of HolmesGPT evals

Automatically triggered by commit fb882c1 on branch claude/kubectl-apply-operator-cr-vkzvj

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 34.5s 5 10 $0.2370 100,393 98,498 23,068 1,895 934 73,721 24,777 — —
✅ 101_loki_historical_logs_pod_deleted 53.1s 6 10 $0.2720 131,302 128,729 24,128 2,573 968 103,577 25,152 — —
✅ 112_find_pvcs_by_uuid 17.6s 3 3 $0.1772 60,685 59,761 21,596 924 547 38,154 21,607 — —
✅ 12_job_crashing 38.5s 5 12 $0.2569 110,188 107,897 24,893 2,291 673 82,441 25,456 — —
✅ 176_network_policy_blocking_traffic_no_runbooks 55.6s 7 16 $0.3207 162,351 159,362 27,402 2,989 774 130,028 29,334 — —
✅ 227_count_configmaps_per_namespace[0] 25.0s 5 9 $0.2051 95,578 94,316 21,017 1,262 585 72,064 22,252 — —
✅ 243_pod_names_contain_service 32.8s 5 8 $0.2170 99,305 97,633 21,744 1,672 534 75,578 22,055 — —
✅ 24_misconfigured_pvc 44.2s 6 15 $0.2772 127,070 124,448 24,405 2,622 712 98,030 26,418 — —
✅ 43_current_datetime_from_prompt 5.4s 1 — $0.1097 17,183 17,057 17,057 126 126 0 17,057 — —
✅ 51_logs_summarize_errors 25.0s 4 5 $0.1843 77,222 76,165 20,820 1,057 322 55,333 20,832 — —
✅ 61_exact_match_counting 10.5s 2 1 $0.1252 34,853 34,561 17,495 292 224 17,056 17,505 — —
Total 31.1s avg 4.5 avg 8.9 avg $2.3823 1,016,130 998,427 27,402 17,703 968 745,982 252,445 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 27 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 #3 · Run @ __5803c3f__ (#23814426380) — Mar 31, 19:17 UTC

✅ Results of HolmesGPT evals

Automatically triggered by commit 5803c3f on branch claude/kubectl-apply-operator-cr-vkzvj

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 30.4s 4 9 $0.2276 82,589 80,692 23,090 1,897 970 56,041 24,651 — —
✅ 101_loki_historical_logs_pod_deleted 62.8s 9 15 $0.3413 203,574 200,158 26,822 3,416 545 172,991 27,167 — —
✅ 112_find_pvcs_by_uuid 17.0s 3 3 $0.1770 60,608 59,686 21,570 922 497 38,105 21,581 — —
✅ 12_job_crashing 33.3s 5 11 $0.2388 107,602 105,682 23,630 1,920 535 81,598 24,084 — —
✅ 176_network_policy_blocking_traffic_no_runbooks 44.7s 5 17 $0.2929 116,542 113,526 27,286 3,016 964 85,395 28,131 — —
✅ 227_count_configmaps_per_namespace[0] 24.3s 5 9 $0.1997 95,632 94,374 21,046 1,258 587 73,315 21,059 — —
✅ 243_pod_names_contain_service 36.2s 5 11 $0.2386 103,523 101,444 22,892 2,079 594 77,613 23,831 — —
✅ 24_misconfigured_pvc 39.8s 5 14 $0.2687 108,804 106,096 24,466 2,708 787 80,019 26,077 — —
✅ 43_current_datetime_from_prompt 5.0s 1 — $0.1096 17,177 17,057 17,057 120 120 0 17,057 — —
✅ 51_logs_summarize_errors 22.3s 4 5 $0.1891 78,299 77,191 21,342 1,108 339 55,837 21,354 — —
✅ 61_exact_match_counting 8.5s 2 1 $0.1233 34,731 34,499 17,433 232 163 17,056 17,443 — —
Total 29.5s avg 4.4 avg 9.5 avg $2.4063 1,009,081 990,405 27,286 18,676 970 737,970 252,435 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 114 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 #2 · Run @ __ddf309c__ (#23814230173) — Mar 31, 19:02 UTC

✅ Results of HolmesGPT evals

Automatically triggered by commit ddf309c on branch claude/kubectl-apply-operator-cr-vkzvj

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 44.7s 6 12 $0.2696 126,824 124,355 24,625 2,469 769 98,821 25,534 — —
✅ 101_loki_historical_logs_pod_deleted 56.0s 7 13 $0.3066 160,654 157,698 25,997 2,956 856 130,728 26,970 — —
✅ 112_find_pvcs_by_uuid 17.8s 3 3 $0.1753 60,326 59,439 21,460 887 465 37,968 21,471 — —
✅ 12_job_crashing 39.5s 6 15 $0.2775 138,158 135,820 25,636 2,338 678 109,349 26,471 — —
✅ 176_network_policy_blocking_traffic_no_runbooks 52.2s 6 16 $0.3047 142,619 139,729 27,287 2,890 756 111,173 28,556 — —
✅ 227_count_configmaps_per_namespace[0] 25.7s 5 9 $0.1968 95,128 93,935 20,853 1,193 567 73,069 20,866 — —
✅ 243_pod_names_contain_service 35.5s 5 9 $0.2256 101,376 99,500 22,409 1,876 586 77,078 22,422 — —
✅ 24_misconfigured_pvc 37.3s 5 13 $0.2482 104,803 102,609 23,411 2,194 674 77,557 25,052 — —
✅ 43_current_datetime_from_prompt 5.2s 1 — $0.1098 17,185 17,057 17,057 128 128 0 17,057 — —
✅ 51_logs_summarize_errors 24.6s 4 5 $0.1871 77,491 76,346 20,909 1,145 404 55,425 20,921 — —
✅ 61_exact_match_counting 10.0s 2 1 $0.1240 34,775 34,521 17,455 254 186 17,056 17,465 — —
Total 31.7s avg 4.5 avg 9.6 avg $2.4251 1,059,339 1,041,009 27,287 18,330 856 788,224 252,785 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 114 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 #1 · Run @ __29546b1__ (#23812193213) — Mar 31, 18:12 UTC

✅ Results of HolmesGPT evals

Automatically triggered by commit 29546b1 on branch claude/kubectl-apply-operator-cr-vkzvj

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 40.1s 5 11 $0.2538 108,696 106,340 24,189 2,356 995 81,568 24,772 — —
✅ 101_loki_historical_logs_pod_deleted 46.8s 5 9 $0.2551 108,491 106,098 24,510 2,393 868 81,307 24,791 — —
✅ 112_find_pvcs_by_uuid 29.9s 5 4 $0.2105 98,841 97,513 22,340 1,328 323 75,160 22,353 — —
✅ 12_job_crashing 43.5s 6 13 $0.2773 136,641 134,354 25,160 2,287 594 107,352 27,002 — —
✅ 176_network_policy_blocking_traffic_no_runbooks 49.0s 8 14 $0.3111 179,737 177,015 26,356 2,722 554 150,003 27,012 — —
✅ 227_count_configmaps_per_namespace[0] 25.6s 5 9 $0.2032 95,551 94,311 21,014 1,240 578 72,353 21,958 — —
✅ 243_pod_names_contain_service 34.7s 4 8 $0.2178 80,605 78,730 21,880 1,875 609 55,571 23,159 — —
✅ 24_misconfigured_pvc 35.6s 5 12 $0.2278 101,552 99,659 22,238 1,893 599 76,813 22,846 — —
✅ 43_current_datetime_from_prompt 5.5s 1 — $0.1098 17,185 17,057 17,057 128 128 0 17,057 — —
✅ 51_logs_summarize_errors 24.4s 4 5 $0.1870 77,998 76,935 21,202 1,063 324 55,721 21,214 — —
✅ 61_exact_match_counting 8.7s 2 1 $0.1239 34,767 34,516 17,450 251 183 17,056 17,460 — —
Total 31.3s avg 4.5 avg 8.6 avg $2.3772 1,040,064 1,022,528 26,356 17,536 995 772,904 249,624 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 114 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)

✅ Results of HolmesGPT evals

Automatically triggered by commit ee321a6 on branch claude/kubectl-apply-operator-cr-vkzvj

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 38.5s 7 11 $0.2680 146,553 144,456 23,798 2,097 565 119,209 25,247 — —
✅ 101_loki_historical_logs_pod_deleted 40.5s 5 10 $0.2405 104,978 102,674 22,714 2,304 880 79,754 22,920 — —
✅ 112_find_pvcs_by_uuid 30.4s 6 6 $0.2341 122,183 120,641 23,203 1,542 380 97,038 23,603 — —
✅ 12_job_crashing 42.7s 6 13 $0.2769 137,365 135,004 25,415 2,361 622 108,651 26,353 — —
✅ 176_network_policy_blocking_traffic_no_runbooks 49.4s 7 15 $0.3117 159,993 157,262 27,006 2,731 629 128,146 29,116 — —
✅ 227_count_configmaps_per_namespace[0] 23.2s 5 9 $0.2021 95,542 94,290 21,006 1,252 585 72,653 21,637 — —
✅ 243_pod_names_contain_service 25.9s 4 6 $0.1907 77,193 75,897 20,705 1,296 453 54,895 21,002 — —
✅ 24_misconfigured_pvc 45.3s 7 17 $0.3016 156,258 153,440 25,802 2,818 896 126,290 27,150 — —
✅ 43_current_datetime_from_prompt 4.0s 1 — $0.1096 17,179 17,057 17,057 122 122 0 17,057 — —
✅ 51_logs_summarize_errors 21.8s 4 5 $0.1883 78,067 76,954 21,209 1,113 351 55,733 21,221 — —
✅ 61_exact_match_counting 7.5s 2 1 $0.1231 34,713 34,486 17,420 227 158 17,056 17,430 — —
Total 29.9s avg 4.9 avg 9.3 avg $2.4466 1,130,024 1,112,161 27,006 17,863 896 859,425 252,736 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 44 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/kubectl-apply-operator-cr-vkzvj -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
tags: regression

Or with more options (one per line):

/eval
model: gpt-4o
tags: regression
id: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
tags: regression
Option Description
model Model(s) to test (default: same as automatic runs)
tags Pytest tags / markers (no default - runs all tests!)
id Eval ID / pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

Option 3: Add PR labels to include extra evals (applies to both automatic runs and /eval comments):

Label Effect
evals-tag-<name> Run tests with tag <name> alongside regression
evals-id-<name> Run a specific eval by test ID
evals-model-<name> Override the model (use model list name, e.g. sonnet-4.5)

Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5

🏷️ Valid tags

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, images, integration, kafka, kubernetes, leaked-information, logs, loki, mcp, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency

🤖 Valid models

deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/kubectl-apply-operator-cr-vkzvj -f markers=regression -f filter=

@github-actions

github-actions Bot commented Mar 31, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker images ready for b6594ed2 (built in 1m 0s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use these tags to pull the images for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:b6594ed2
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:b6594ed2 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:b6594ed2
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:b6594ed2
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:b6594ed2
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:b6594ed2 me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:b6594ed2
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:b6594ed2

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:b6594ed2 \
  --set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set operator.image=holmes-operator-dev:b6594ed2

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:b6594ed2 \
  --set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.operator.image=holmes-operator-dev:b6594ed2

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (3)
tests/holmes_operator/test_healthcheck_component.py (1)

81-95: Consider adding type hints to the helper function.

The helper function lacks type hints which are required per coding guidelines.

📝 Suggested fix
-def _make_body(name, namespace, uid, spec, generation=1):
+def _make_body(
+    name: str, namespace: str, uid: str, spec: dict, generation: int = 1
+) -> dict:
     """Helper to build a HealthCheck resource body with metadata.generation."""

As per coding guidelines: **/*.py: Type hints are required throughout the codebase.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@tests/holmes_operator/test_healthcheck_component.py` around lines 81 - 95,
The helper function _make_body lacks type hints; update its signature to add
parameter and return type annotations (e.g., name: str, namespace: str, uid:
str, spec: Dict[str, Any] or Mapping[str, Any], generation: int = 1) and
annotate the return as Dict[str, Any]; also add any required typing imports
(from typing import Any, Dict, Mapping) at the top of the module so the function
and its returned dictionary conform to the project's type-hinting guidelines.
holmes_operator/utils.py (1)

34-53: Consider updating the docstring to document the new parameter.

The observed_generation parameter is added but not documented in the Args section of the docstring.

📝 Suggested docstring update
     Args:
         api: Kubernetes CustomObjectsApi instance
         name: Name of the HealthCheck resource
         namespace: Namespace of the HealthCheck resource
         phase: Execution phase (Pending, Running, Completed, Failed)
         result: Check result (pass, fail, error)
         message: Human-readable summary
         rationale: LLM explanation
         duration: Execution duration in seconds
         error: Error details
         model_used: Model that was used
         notifications: List of notification statuses
         start_time: ISO format start time
         completion_time: ISO format completion time
+        observed_generation: Generation to record as processed
     """
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@holmes_operator/utils.py` around lines 34 - 53, The docstring for the
function that updates HealthCheck status (update_healthcheck_status) is missing
documentation for the new observed_generation parameter; update the Args section
to add an entry for observed_generation (type Optional[int]) describing it as
the resource's observedGeneration value used to indicate the controller's
observed revision of the HealthCheck, and note that it should be set when
reporting status to help reconcile loops detect changes.
holmes_operator/handlers/healthcheck.py (1)

176-183: Consider using f-string conversion flag.

Static analysis suggests using explicit conversion flag instead of str(e).

📝 Suggested fix
-            message=f"Operator error: {str(e)}",
-            error=str(e),
+            message=f"Operator error: {e!s}",
+            error=f"{e!s}",
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@holmes_operator/handlers/healthcheck.py` around lines 176 - 183, Replace
explicit str(e) calls with f-string conversion flags to make formatting clearer:
update the f"Operator error: {str(e)}" to use f"Operator error: {e!s}" and
change error=str(e) to error=f"{e!s}" in the set_healthcheck_failed call so the
exception is converted via the f-string conversion flag; this touches the
set_healthcheck_failed invocation in the healthcheck handler where variables
name, namespace, generation and context.k8s_api are passed.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Nitpick comments:
In `@holmes_operator/handlers/healthcheck.py`:
- Around line 176-183: Replace explicit str(e) calls with f-string conversion
flags to make formatting clearer: update the f"Operator error: {str(e)}" to use
f"Operator error: {e!s}" and change error=str(e) to error=f"{e!s}" in the
set_healthcheck_failed call so the exception is converted via the f-string
conversion flag; this touches the set_healthcheck_failed invocation in the
healthcheck handler where variables name, namespace, generation and
context.k8s_api are passed.

In `@holmes_operator/utils.py`:
- Around line 34-53: The docstring for the function that updates HealthCheck
status (update_healthcheck_status) is missing documentation for the new
observed_generation parameter; update the Args section to add an entry for
observed_generation (type Optional[int]) describing it as the resource's
observedGeneration value used to indicate the controller's observed revision of
the HealthCheck, and note that it should be set when reporting status to help
reconcile loops detect changes.

In `@tests/holmes_operator/test_healthcheck_component.py`:
- Around line 81-95: The helper function _make_body lacks type hints; update its
signature to add parameter and return type annotations (e.g., name: str,
namespace: str, uid: str, spec: Dict[str, Any] or Mapping[str, Any], generation:
int = 1) and annotate the return as Dict[str, Any]; also add any required typing
imports (from typing import Any, Dict, Mapping) at the top of the module so the
function and its returned dictionary conform to the project's type-hinting
guidelines.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: ace34798-353e-4e34-a44c-01e033d8269a

📥 Commits

Reviewing files that changed from the base of the PR and between 1ef4643 and 29546b1.

📒 Files selected for processing (5)
  • helm/holmes/crds/healthcheck.yaml
  • holmes_operator/handlers/healthcheck.py
  • holmes_operator/models.py
  • holmes_operator/utils.py
  • tests/holmes_operator/test_healthcheck_component.py

aantn and others added 4 commits March 31, 2026 21:54
- deployment-verification.md: Change example to use fixed check name
  with version in query (not name), explain auto re-run on spec change,
  update tips section
- health-checks.md: Expand "Re-running Checks" section to document both
  spec-change trigger and annotation trigger, add observedGeneration to
  status fields reference

https://claude.ai/code/session_01QY6zsEHPW9CfSKJBtE6WCd
Signed-off-by: Claude <noreply@anthropic.com>
Kept remote's open-ended query style but added version reference
(v2.4.1) so the query naturally changes between deploys, triggering
the generation-based re-execution. Updated CI/CD gating script to
use fixed check name.

https://claude.ai/code/session_01QY6zsEHPW9CfSKJBtE6WCd
Signed-off-by: Claude <noreply@anthropic.com>
Instead of requiring users to template version strings into queries
or use run-id annotations, the operator now clears holmesgpt.dev/rerun
on create too (not just update). This enables a simple toggle pattern:

1. Manifest includes holmesgpt.dev/rerun: "true"
2. Operator runs check, clears annotation
3. Next kubectl apply restores annotation from manifest → re-run

Works with Helm, ArgoCD, or plain manifests with zero templating.
Reverted the unused lastRunId/run-id approach.

Updated docs with Helm and ArgoCD examples.

https://claude.ai/code/session_01QY6zsEHPW9CfSKJBtE6WCd
Signed-off-by: Claude <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
docs/operator/deployment-verification.md (1)

81-87: Helpful tips section with clear guidance.

The updated tips correctly guide users toward the fixed-name pattern with query updates for most cases, while noting that versioned names remain an option for audit trails. The annotation-based manual re-run is accurately documented as an alternative.

Optional style refinement: The phrase "exact same" on line 87 could be simplified to "same" for conciseness, but the current wording is clear.

Minor wording refinement
-- **Force re-run without spec changes:** If you need to re-run the exact same check, use `kubectl annotate hc/checkout-api-deploy-check holmesgpt.dev/rerun=true`. The annotation is cleared automatically after execution.
+- **Force re-run without spec changes:** If you need to re-run the same check, use `kubectl annotate hc/checkout-api-deploy-check holmesgpt.dev/rerun=true`. The annotation is cleared automatically after execution.
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@docs/operator/deployment-verification.md` around lines 81 - 87, In the "Force
re-run without spec changes:" tip, simplify the phrase "re-run the exact same
check" to "re-run the same check" for conciseness; update the sentence that
follows the header (the one that references the annotation example `kubectl
annotate hc/checkout-api-deploy-check holmesgpt.dev/rerun=true`) to use "same"
instead of "exact same" and keep the rest of the line and example unchanged.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Nitpick comments:
In `@docs/operator/deployment-verification.md`:
- Around line 81-87: In the "Force re-run without spec changes:" tip, simplify
the phrase "re-run the exact same check" to "re-run the same check" for
conciseness; update the sentence that follows the header (the one that
references the annotation example `kubectl annotate hc/checkout-api-deploy-check
holmesgpt.dev/rerun=true`) to use "same" instead of "exact same" and keep the
rest of the line and example unchanged.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: c41c2dcf-14f6-4fb5-b2fd-c0705f576a51

📥 Commits

Reviewing files that changed from the base of the PR and between 29546b1 and 5803c3f.

📒 Files selected for processing (2)
  • docs/operator/deployment-verification.md
  • docs/operator/health-checks.md

- Use mychart.fullname helper instead of .Release.Name for the
  Deployment name reference (release name != app Deployment name)
- Add namespace to Helm and ArgoCD examples
- Add labels block to Helm template
- Minor wording improvements

https://claude.ai/code/session_01QY6zsEHPW9CfSKJBtE6WCd
Signed-off-by: Claude <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
docs/operator/deployment-verification.md (1)

107-123: ⚠️ Potential issue | 🟠 Major

Wait for the new generation before trusting .status.result.

This example now reuses a fixed HealthCheck name, so the first poll after kubectl apply can still read the previous run's pass/fail before the operator has started the new execution. Gate on status.observedGeneration == metadata.generation (or a changed completion timestamp) before acting on status.result.

🧪 Suggested update
+# Capture the generation created by this apply
+TARGET_GEN=$(kubectl get hc checkout-api-deploy-check -n production -o jsonpath='{.metadata.generation}')
+
 # Wait for the check to complete, then read the result
 for i in $(seq 1 30); do
+  OBSERVED_GEN=$(kubectl get hc checkout-api-deploy-check -n production -o jsonpath='{.status.observedGeneration}' 2>/dev/null)
   RESULT=$(kubectl get hc checkout-api-deploy-check -n production -o jsonpath='{.status.result}' 2>/dev/null)
-  if [ "$RESULT" = "pass" ]; then
+  if [ "$OBSERVED_GEN" = "$TARGET_GEN" ] && [ "$RESULT" = "pass" ]; then
     echo "Deploy verified healthy"
     exit 0
-  elif [ "$RESULT" = "fail" ] || [ "$RESULT" = "error" ]; then
+  elif [ "$OBSERVED_GEN" = "$TARGET_GEN" ] && { [ "$RESULT" = "fail" ] || [ "$RESULT" = "error" ]; }; then
     echo "Deploy check failed:"
     kubectl get hc checkout-api-deploy-check -n production -o jsonpath='{.status.message}'
     exit 1
   fi
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@docs/operator/deployment-verification.md` around lines 107 - 123, The poll
loop can read a previous HealthCheck run's result; modify the logic that queries
the HealthCheck (named checkout-api-deploy-check) to first fetch and compare
status.observedGeneration with metadata.generation (or alternatively wait for a
changed status.completedAt) and only then consider status.result; in practice,
update the loop that reads kubectl get hc ... -o jsonpath='{.status.result}' to
first retrieve both metadata.generation and status.observedGeneration and
continue sleeping until they match (or until completedAt advances) before
evaluating status.result for "pass"/"fail"/"error".
🧹 Nitpick comments (1)
docs/operator/deployment-verification.md (1)

3-14: Document the ArgoCD drift caveat for this pattern.

Because the operator removes holmesgpt.dev/rerun from the live object, GitOps controllers will see this resource as drifted after every run. In ArgoCD with selfHeal enabled, that can turn into a sync/rerun loop instead of “once per deploy”, so this section should mention either ignoring diffs for this annotation or limiting the pattern to manual/commit-driven syncs.

Also applies to: 84-103

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@docs/operator/deployment-verification.md` around lines 3 - 14, Update the
paragraph about the holmesgpt.dev/rerun annotation to add an ArgoCD caveat: note
that because the operator clears the holmesgpt.dev/rerun annotation from the
live object, GitOps controllers (ArgoCD) will detect drift and—if selfHeal is
enabled—may enter a sync loop; instruct users to either add an ArgoCD
ignoreDifferences rule for the holmesgpt.dev/rerun annotation
(resource.customizations or argocd-cm diff/ignore settings) or restrict this
pattern to manual/commit-driven syncs (disable selfHeal for the resource) to
avoid repeated re-runs.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@holmes_operator/handlers/healthcheck.py`:
- Around line 236-250: The post-execution cleanup that clears the
holmesgpt.dev/rerun annotation is skipped if _execute_healthcheck raises or if
the generation-mismatch early return occurs; wrap the call to
_execute_healthcheck (and any surrounding logic) in a try/finally so
_clear_rerun_annotation(name=..., namespace=..., logger=...) is always awaited
in the finally block, and modify the generation-mismatch branch inside the same
handler (the update handler around lines handling generation and returning
early) to call _clear_rerun_annotation before returning so the annotation is
removed in all paths.

---

Outside diff comments:
In `@docs/operator/deployment-verification.md`:
- Around line 107-123: The poll loop can read a previous HealthCheck run's
result; modify the logic that queries the HealthCheck (named
checkout-api-deploy-check) to first fetch and compare status.observedGeneration
with metadata.generation (or alternatively wait for a changed
status.completedAt) and only then consider status.result; in practice, update
the loop that reads kubectl get hc ... -o jsonpath='{.status.result}' to first
retrieve both metadata.generation and status.observedGeneration and continue
sleeping until they match (or until completedAt advances) before evaluating
status.result for "pass"/"fail"/"error".

---

Nitpick comments:
In `@docs/operator/deployment-verification.md`:
- Around line 3-14: Update the paragraph about the holmesgpt.dev/rerun
annotation to add an ArgoCD caveat: note that because the operator clears the
holmesgpt.dev/rerun annotation from the live object, GitOps controllers (ArgoCD)
will detect drift and—if selfHeal is enabled—may enter a sync loop; instruct
users to either add an ArgoCD ignoreDifferences rule for the holmesgpt.dev/rerun
annotation (resource.customizations or argocd-cm diff/ignore settings) or
restrict this pattern to manual/commit-driven syncs (disable selfHeal for the
resource) to avoid repeated re-runs.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: f4610c6b-6ec8-4302-91fd-4ef39e771c0f

📥 Commits

Reviewing files that changed from the base of the PR and between 5803c3f and fb882c1.

📒 Files selected for processing (4)
  • docs/operator/deployment-verification.md
  • docs/operator/health-checks.md
  • holmes_operator/handlers/healthcheck.py
  • tests/holmes_operator/test_healthcheck_component.py
🚧 Files skipped from review as they are similar to previous changes (1)
  • tests/holmes_operator/test_healthcheck_component.py

Comment on lines +236 to +250
await _execute_healthcheck(
spec=spec,
name=name,
namespace=namespace,
uid=uid,
generation=generation,
logger=logger,
body=body,
)

# Clear rerun annotation if present, so that the next kubectl apply
# (which restores it from the manifest) triggers a re-run
annotations = body.get("metadata", {}).get("annotations", {})
if annotations.get("holmesgpt.dev/rerun") == "true":
await _clear_rerun_annotation(name=name, namespace=namespace, logger=logger)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

Always clear holmesgpt.dev/rerun when an execution consumed it.

Two paths still leave the flag stuck at "true": the generation-mismatch branch returns before cleanup, and any _execute_healthcheck() exception skips the post-call cleanup entirely. After that, later kubectl apply / kubectl annotate ... rerun=true calls stop creating a fresh transition, so the rerun mechanism quietly stops working for that object.

♻️ Suggested direction
-    await _execute_healthcheck(
-        spec=spec,
-        name=name,
-        namespace=namespace,
-        uid=uid,
-        generation=generation,
-        logger=logger,
-        body=body,
-    )
-
-    # Clear rerun annotation if present, so that the next kubectl apply
-    # (which restores it from the manifest) triggers a re-run
-    annotations = body.get("metadata", {}).get("annotations", {})
-    if annotations.get("holmesgpt.dev/rerun") == "true":
-        await _clear_rerun_annotation(name=name, namespace=namespace, logger=logger)
+    annotations = body.get("metadata", {}).get("annotations", {})
+    should_clear_rerun = annotations.get("holmesgpt.dev/rerun") == "true"
+    try:
+        await _execute_healthcheck(
+            spec=spec,
+            name=name,
+            namespace=namespace,
+            uid=uid,
+            generation=generation,
+            logger=logger,
+            body=body,
+        )
+    finally:
+        if should_clear_rerun:
+            await _clear_rerun_annotation(name=name, namespace=namespace, logger=logger)

Apply the same try/finally pattern to the update handler, and make sure the generation-mismatch branch clears the annotation before returning.

Also applies to: 277-313

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@holmes_operator/handlers/healthcheck.py` around lines 236 - 250, The
post-execution cleanup that clears the holmesgpt.dev/rerun annotation is skipped
if _execute_healthcheck raises or if the generation-mismatch early return
occurs; wrap the call to _execute_healthcheck (and any surrounding logic) in a
try/finally so _clear_rerun_annotation(name=..., namespace=..., logger=...) is
always awaited in the finally block, and modify the generation-mismatch branch
inside the same handler (the update handler around lines handling generation and
returning early) to call _clear_rerun_annotation before returning so the
annotation is removed in all paths.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
docs/operator/deployment-verification.md (1)

3-14: Document generation-based re-execution alongside annotation-based mechanism.

The documentation focuses exclusively on the holmesgpt.dev/rerun annotation but doesn't mention that the operator also automatically re-executes when the HealthCheck spec changes (metadata.generation != status.observedGeneration). This could confuse users who see re-execution happen without the annotation, or lead them to believe the annotation is always required.

Consider adding a note explaining both mechanisms:

  • Generation-based (automatic): Triggers when spec fields change (query, timeout, mode, etc.)
  • Annotation-based (manual): Triggers re-execution even when spec is unchanged — useful for deploying the same HealthCheck definition repeatedly

For this "Deployment Verification" use case where the HealthCheck spec is typically static, the annotation approach is correct, but documenting both mechanisms would improve clarity.

Based on context snippets from holmes_operator/handlers/healthcheck.py:274-292 showing the generation check as the first trigger, and from helm/holmes/crds/healthcheck.yaml:69-72 documenting the observedGeneration field.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@docs/operator/deployment-verification.md` around lines 3 - 14, Docs only
describe the holmesgpt.dev/rerun annotation but omit the automatic
generation-based re-execution; update deployment-verification.md to mention both
triggers: explain that re-execution also happens automatically when
metadata.generation != status.observedGeneration (i.e., spec changes such as
query, timeout, mode) and that the holmesgpt.dev/rerun: "true" annotation forces
re-run even when spec is unchanged; reference the operator check in
holmes_operator/handlers/healthcheck.py (the generation vs observedGeneration
logic) and the observedGeneration field documented in
helm/holmes/crds/healthcheck.yaml so readers understand when to use the
annotation vs relying on spec changes.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Nitpick comments:
In `@docs/operator/deployment-verification.md`:
- Around line 3-14: Docs only describe the holmesgpt.dev/rerun annotation but
omit the automatic generation-based re-execution; update
deployment-verification.md to mention both triggers: explain that re-execution
also happens automatically when metadata.generation != status.observedGeneration
(i.e., spec changes such as query, timeout, mode) and that the
holmesgpt.dev/rerun: "true" annotation forces re-run even when spec is
unchanged; reference the operator check in
holmes_operator/handlers/healthcheck.py (the generation vs observedGeneration
logic) and the observedGeneration field documented in
helm/holmes/crds/healthcheck.yaml so readers understand when to use the
annotation vs relying on spec changes.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 52f7f010-ce8e-4b39-8e09-f0153dd88c55

📥 Commits

Reviewing files that changed from the base of the PR and between fb882c1 and ee321a6.

📒 Files selected for processing (1)
  • docs/operator/deployment-verification.md

Comment on lines +271 to +285
body = new
metadata = body.get("metadata", {})
status = body.get("status", {})
generation = metadata.get("generation")
observed_generation = status.get("observedGeneration")

# Trigger 1: Generation changed (spec was modified via kubectl apply)
if generation is not None and generation != observed_generation:
logger.info(
f"Re-running HealthCheck {namespace}/{name}: "
f"generation={generation} != observedGeneration={observed_generation}"
)
await _execute_healthcheck(
spec=new.get("spec", {}),
name=name,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 When Trigger 1 (generation mismatch) fires in on_healthcheck_update, it executes the check and does an early return without checking or clearing the holmesgpt.dev/rerun annotation, leaving it permanently as true in K8s. This breaks the documented Helm/ArgoCD CI/CD toggle pattern: on the next deploy with the same spec, kubectl apply sees no annotation diff and fires no update event, so no re-run ever happens again. Fix by calling _clear_rerun_annotation before the return in Trigger 1 if the annotation is present.

Extended reasoning...

Bug: Trigger 1 early return skips annotation clearing

What the bug is

In on_healthcheck_update (lines 271-285 of holmes_operator/handlers/healthcheck.py), the handler checks two triggers in order. Trigger 1 (generation mismatch) executes _execute_healthcheck and then does return, completely bypassing the Trigger 2 block that checks for and clears the holmesgpt.dev/rerun annotation.

The specific code path

Trigger 1 code:

Trigger 2 (never reached when Trigger 1 fires):

Why existing code does not prevent it

The annotation-clearing logic lives exclusively in the Trigger 2 block. Neither the _execute_healthcheck helper nor the Trigger 1 code path has any awareness of the annotation. Any code path that returns before reaching Trigger 2 silently skips the clearing.

Impact

This breaks the key Helm/ArgoCD rerun toggle pattern advertised in the PR documentation. The docs explicitly tell users to add holmesgpt.dev/rerun: "true" permanently to their manifest so that every helm upgrade or ArgoCD sync triggers a fresh health check. Once a user does a spec-changing deploy (bumping generation), this pattern permanently stops working.

Step-by-step proof

  1. User has holmesgpt.dev/rerun: "true" permanently in Helm chart template.
  2. Deploy 1 (initial): generation=1, observedGeneration=null. Trigger 1 fires (1 != null), check runs, observedGeneration set to 1, return executed — annotation stays true in K8s, never cleared.
  3. Deploy 2: user changes the spec query. Generation bumps to 2. Trigger 1 fires (2 != 1), check runs, observedGeneration set to 2, return — annotation remains true in K8s, never cleared.
  4. Deploy 3: same spec as Deploy 2, annotation still true in manifest. kubectl apply computes the diff: K8s annotation = true, manifest annotation = true — no change detected, no PATCH sent, no update event fires at all. Health check is silently skipped.
  5. Even if an update event fires for another reason: old_annotations["holmesgpt.dev/rerun"] = "true" (never cleared), new_annotations["holmesgpt.dev/rerun"] = "true" so Trigger 2 condition old != "true" is False and still no re-run happens.

Fix

Before the return in Trigger 1, check for and clear the rerun annotation if present:

Comment on lines +280 to +306
f"Re-running HealthCheck {namespace}/{name}: "
f"generation={generation} != observedGeneration={observed_generation}"
)
await _execute_healthcheck(
spec=new.get("spec", {}),
name=name,
namespace=namespace,
uid=metadata.get("uid", ""),
generation=generation,
logger=logger,
body=body,
)
return

# Trigger 2: Rerun annotation (for re-running without spec changes)
annotations = metadata.get("annotations", {})
old_annotations = old.get("metadata", {}).get("annotations", {})
if (
annotations.get("holmesgpt.dev/rerun") == "true"
and old_annotations.get("holmesgpt.dev/rerun") != "true"
):
logger.info(f"Re-running HealthCheck: {namespace}/{name}")

# Trigger re-execution by calling create handler
await on_healthcheck_create(
logger.info(f"Re-running HealthCheck via annotation: {namespace}/{name}")
await _execute_healthcheck(
spec=new.get("spec", {}),
name=name,
namespace=namespace,
uid=new.get("metadata", {}).get("uid", ""),
uid=metadata.get("uid", ""),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 The on_healthcheck_update handler has two annotation-clearing gaps that can permanently break the Helm/ArgoCD rerun toggle pattern. In Trigger 1 (generation mismatch), an early return skips the annotation-clearing code entirely; and in Trigger 2 (annotation-based rerun), _clear_rerun_annotation is called after _execute_healthcheck which re-raises on failure, leaving the annotation stuck as true after any operator error. Both paths should clear the annotation (or use try/finally) before returning/raising.

Extended reasoning...

Bug 1 - Trigger 1 early return skips annotation clearing (lines 280-285)

When metadata.generation != status.observedGeneration, the handler calls _execute_healthcheck and immediately does return, bypassing the Trigger 2 block that is responsible for clearing the holmesgpt.dev/rerun annotation. This means any resource that has holmesgpt.dev/rerun: "true" in its manifest AND receives a spec change at the same time will exit Trigger 1 with the annotation still set to true in the live Kubernetes object.

Concrete proof (Bug 1):

  1. User has holmesgpt.dev/rerun: "true" in their Helm/ArgoCD manifest (the toggle pattern documented in this PR).
  2. User also changes spec.query - generation bumps from 1 to 2.
  3. Update event fires: generation=2, observedGeneration=1 -> Trigger 1 fires, check runs, set_healthcheck_completed sets observedGeneration=2, handler does return. Annotation remains "true" in K8s.
  4. Next helm upgrade or ArgoCD sync with no spec change: kubectl apply computes the diff - manifest says rerun=true, K8s already has rerun=true -> no diff, no PATCH, no update event. The check never re-runs. The toggle is permanently broken until someone manually removes the annotation.

Bug 2 - Trigger 2 annotation not cleared on _execute_healthcheck exception (lines 295-306)

In Trigger 2, _clear_rerun_annotation is called sequentially after _execute_healthcheck. But _execute_healthcheck re-raises any exception after setting observedGeneration via set_healthcheck_failed. The exception propagates before _clear_rerun_annotation is ever called.

Concrete proof (Bug 2):

  1. User sets kubectl annotate hc my-check holmesgpt.dev/rerun=true. Old annotation is absent, new annotation is "true" -> Trigger 2 fires.
  2. Holmes API returns 500. _execute_healthcheck calls set_healthcheck_failed(..., observed_generation=generation), then re-raises the exception.
  3. _clear_rerun_annotation is never reached. Now status.observedGeneration == generation AND annotation == "true" in K8s.
  4. On the next kubectl apply or retry: Trigger 1 does not fire (generation == observedGeneration). Trigger 2 does not fire (old_annotations.get("holmesgpt.dev/rerun") != "true" evaluates False since the old state in kopf's cache also has "true"). The user is stuck and must manually delete then re-add the annotation to retry.

Why existing code does not prevent this:
The annotation-clearing logic in Trigger 2 assumes _execute_healthcheck always succeeds (or that exceptions propagate without side effects). The early return in Trigger 1 was probably intentional to avoid double-execution, but it accidentally skips annotation cleanup. There is no try/finally guard in either path.

Impact:
Both bugs break the primary CI/CD workflow advertised in this PR's documentation (the Helm/ArgoCD rerun toggle pattern). Bug 1 triggers reliably any time a user combines a spec change with the persistent annotation pattern. Bug 2 makes failure recovery impossible via annotation - users must perform a two-step manual kubectl annotate delete+restore.

Suggested fix:

  • Trigger 1: Before return, check if the annotation is present and call _clear_rerun_annotation if so.
  • Trigger 2: Wrap in try/finally to ensure _clear_rerun_annotation is always called regardless of whether _execute_healthcheck raises.

Comment thread holmes_operator/handlers/healthcheck.py
Comment thread tests/holmes_operator/test_healthcheck_component.py
Comment thread holmes_operator/handlers/healthcheck.py
Comment on lines 120 to 127
echo "Deploy verified healthy"
exit 0
elif [ "$RESULT" = "fail" ] || [ "$RESULT" = "error" ]; then
echo "Deploy check failed:"
kubectl get hc checkout-api-deploy-v2-4-1 -n production -o jsonpath='{.status.message}'
kubectl get hc checkout-api-deploy-check -n production -o jsonpath='{.status.message}'
exit 1
fi
sleep 10

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 The CI/CD polling script in deployment-verification.md checks only status.result without verifying status.phase == Completed, causing it to report false success during a re-run. Because set_healthcheck_pending only patches phase=Pending without clearing the old result field, the stale result=pass from a prior run is visible immediately when a new check starts — the script exits 0 before the new check completes. Fix by updating the polling script to also assert phase=Completed before trusting the result, or by having set_healthcheck_pending explicitly null out the result/message/etc. fields.

Extended reasoning...

What the bug is and how it manifests

set_healthcheck_pending in utils.py only patches phase=Pending and startTime. Kubernetes strategic merge PATCH preserves every field not explicitly set, so the old result, message, rationale, duration, completionTime, and modelUsed remain in the live resource status from the previous run. The CI/CD polling script in deployment-verification.md (lines 120-127) polls status.result directly, with no check that status.phase == Completed before trusting the result.

The specific code path that triggers it

When a re-run fires (either via generation mismatch in on_healthcheck_update Trigger 1, or via the holmesgpt.dev/rerun annotation in Trigger 2), _execute_healthcheck is called which immediately invokes set_healthcheck_pending. That function calls update_healthcheck_status with only phase=Pending and start_time. The status subresource PATCH does not include result, message, or completionTime, so Kubernetes preserves those fields at their previous values. The polling script immediately queries status.result which returns the stale value and exits 0.

Why existing code does not prevent it

Neither set_healthcheck_pending nor _execute_healthcheck explicitly nulls out the stale result fields. The polling script has no guard checking phase — it will immediately see the stale result=pass on the very first poll iteration, which may happen within milliseconds of kubectl apply before the operator has even processed the update event, let alone completed the check.

Impact

The documented CI/CD gating use case is silently broken for repeat deploys. A pipeline will declare "Deploy verified healthy" before the new health check completes, bypassing the entire safety gate. The previous docs used versioned resource names (e.g., checkout-api-deploy-v2-4-1) meaning each deploy started with a fresh resource and no prior result. This PR explicitly changes the guidance to a persistent name with the rerun toggle, making the stale-result window the default for every deploy after the first.

Step-by-step proof

  1. Deploy 1: HealthCheck checkout-api-deploy-check runs and passes. Status: {phase: Completed, result: pass, message: "All healthy"}
  2. Deploy 2: CI runs kubectl apply with holmesgpt.dev/rerun: "true" restored by Helm/ArgoCD. Operator fires on_healthcheck_update -> _execute_healthcheck -> set_healthcheck_pending patches only phase=Pending, startTime=now. Status is now: {phase: Pending, result: pass (STALE), ...}
  3. CI polling script starts its loop immediately after kubectl apply. On the first iteration it queries status.result -> sees "pass" -> prints "Deploy verified healthy" -> exit 0.
  4. The new check has not yet completed or even started running. The CI gate reports success on a stale result.

How to fix

Either: (1) Update the polling script to check phase=Completed before trusting result — add a PHASE check alongside the RESULT check so the script only exits when both phase=Completed AND result=pass. Or: (2) have set_healthcheck_pending explicitly null out result, message, rationale, completionTime, duration, modelUsed so there is no stale data window at all.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants