Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
82 changes: 68 additions & 14 deletions docs/operator/deployment-verification.md
Original file line number Diff line number Diff line change
@@ -1,10 +1,19 @@
# Deployment Verification

A common pattern is deploying a HealthCheck alongside your application to verify the new version is working correctly. Since HealthChecks run immediately when created, you can include one in the same manifest (or CI/CD step) as your deployment and use the result to gate rollout progression.
A common pattern is deploying a HealthCheck alongside your application to verify the new version is working correctly. By adding the `holmesgpt.dev/rerun: "true"` annotation to your HealthCheck manifest, the operator will **re-run the check on every deploy** — no manual intervention needed.

## One-Time Verification with HealthCheck
## How It Works

Include a [HealthCheck](health-checks.md) in the same manifest as your deployment. It runs immediately after `kubectl apply` and reports whether the new version started correctly.
The `holmesgpt.dev/rerun` annotation creates an automatic toggle:

1. Your manifest includes `holmesgpt.dev/rerun: "true"`
2. On `kubectl apply`, the operator runs the check and **clears the annotation**
3. On the next deploy, `kubectl apply` sees the annotation is missing (operator cleared it) but the manifest says it should be `"true"` → restores it → triggers re-run
4. Repeat forever

This means every `helm upgrade` or ArgoCD sync automatically triggers a fresh health check.

## Example: HealthCheck Alongside a Deployment

```yaml
# app-deployment.yaml
Expand All @@ -30,11 +39,12 @@ spec:
apiVersion: holmesgpt.dev/v1alpha1
kind: HealthCheck
metadata:
name: checkout-api-deploy-v2-4-1
name: checkout-api-deploy-check
namespace: production
annotations:
holmesgpt.dev/rerun: "true"
labels:
app: checkout-api
deploy-version: v2.4.1
spec:
query: "We just rolled out a new version of checkout-api to production. Compare logs, error rates, latency, and resource usage before and after the deploy. Flag anything that changed or looks off."
timeout: 120
Expand All @@ -45,13 +55,58 @@ spec:
channel: "#deploy-alerts"
```

Apply both together:

```bash
kubectl apply -f app-deployment.yaml
```

If pods crash or fail readiness, the check fails and alerts your team.
## Helm

Add a HealthCheck template to your application's Helm chart:

```yaml
# templates/healthcheck.yaml
apiVersion: holmesgpt.dev/v1alpha1
kind: HealthCheck
metadata:
name: {{ include "mychart.fullname" . }}-deploy-check
namespace: {{ .Release.Namespace }}
annotations:
holmesgpt.dev/rerun: "true"
labels:
{{- include "mychart.labels" . | nindent 4 }}
spec:
query: "Is the {{ include "mychart.fullname" . }} deployment in '{{ .Release.Namespace }}' fully rolled out and healthy? Check pod status, logs, and error rates."
timeout: 120
mode: alert
destinations:
- type: slack
config:
channel: "#deploy-alerts"
```

Every `helm upgrade` triggers a fresh check.

## ArgoCD

Add a HealthCheck to the same repo/path as your Application source. Every ArgoCD sync triggers a re-run:

```yaml
apiVersion: holmesgpt.dev/v1alpha1
kind: HealthCheck
metadata:
name: checkout-api-deploy-check
namespace: production
annotations:
holmesgpt.dev/rerun: "true"
spec:
query: "Is the checkout-api deployment in 'production' fully rolled out and healthy? Check pod status, logs, and error rates."
timeout: 120
mode: alert
destinations:
- type: slack
config:
channel: "#deploy-alerts"
```

## Gating CI/CD on the Result

Expand All @@ -60,13 +115,13 @@ After applying, poll for the result to gate your pipeline:
```bash
# Wait for the check to complete, then read the result
for i in $(seq 1 30); do
RESULT=$(kubectl get hc checkout-api-deploy-v2-4-1 -n production -o jsonpath='{.status.result}' 2>/dev/null)
RESULT=$(kubectl get hc checkout-api-deploy-check -n production -o jsonpath='{.status.result}' 2>/dev/null)
if [ "$RESULT" = "pass" ]; then
echo "Deploy verified healthy"
exit 0
elif [ "$RESULT" = "fail" ] || [ "$RESULT" = "error" ]; then
echo "Deploy check failed:"
kubectl get hc checkout-api-deploy-v2-4-1 -n production -o jsonpath='{.status.message}'
kubectl get hc checkout-api-deploy-check -n production -o jsonpath='{.status.message}'
exit 1
fi
sleep 10
Comment on lines 120 to 127

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 The CI/CD polling script in deployment-verification.md checks only status.result without verifying status.phase == Completed, causing it to report false success during a re-run. Because set_healthcheck_pending only patches phase=Pending without clearing the old result field, the stale result=pass from a prior run is visible immediately when a new check starts — the script exits 0 before the new check completes. Fix by updating the polling script to also assert phase=Completed before trusting the result, or by having set_healthcheck_pending explicitly null out the result/message/etc. fields.

Extended reasoning...

What the bug is and how it manifests

set_healthcheck_pending in utils.py only patches phase=Pending and startTime. Kubernetes strategic merge PATCH preserves every field not explicitly set, so the old result, message, rationale, duration, completionTime, and modelUsed remain in the live resource status from the previous run. The CI/CD polling script in deployment-verification.md (lines 120-127) polls status.result directly, with no check that status.phase == Completed before trusting the result.

The specific code path that triggers it

When a re-run fires (either via generation mismatch in on_healthcheck_update Trigger 1, or via the holmesgpt.dev/rerun annotation in Trigger 2), _execute_healthcheck is called which immediately invokes set_healthcheck_pending. That function calls update_healthcheck_status with only phase=Pending and start_time. The status subresource PATCH does not include result, message, or completionTime, so Kubernetes preserves those fields at their previous values. The polling script immediately queries status.result which returns the stale value and exits 0.

Why existing code does not prevent it

Neither set_healthcheck_pending nor _execute_healthcheck explicitly nulls out the stale result fields. The polling script has no guard checking phase — it will immediately see the stale result=pass on the very first poll iteration, which may happen within milliseconds of kubectl apply before the operator has even processed the update event, let alone completed the check.

Impact

The documented CI/CD gating use case is silently broken for repeat deploys. A pipeline will declare "Deploy verified healthy" before the new health check completes, bypassing the entire safety gate. The previous docs used versioned resource names (e.g., checkout-api-deploy-v2-4-1) meaning each deploy started with a fresh resource and no prior result. This PR explicitly changes the guidance to a persistent name with the rerun toggle, making the stale-result window the default for every deploy after the first.

Step-by-step proof

  1. Deploy 1: HealthCheck checkout-api-deploy-check runs and passes. Status: {phase: Completed, result: pass, message: "All healthy"}
  2. Deploy 2: CI runs kubectl apply with holmesgpt.dev/rerun: "true" restored by Helm/ArgoCD. Operator fires on_healthcheck_update -> _execute_healthcheck -> set_healthcheck_pending patches only phase=Pending, startTime=now. Status is now: {phase: Pending, result: pass (STALE), ...}
  3. CI polling script starts its loop immediately after kubectl apply. On the first iteration it queries status.result -> sees "pass" -> prints "Deploy verified healthy" -> exit 0.
  4. The new check has not yet completed or even started running. The CI gate reports success on a stale result.

How to fix

Either: (1) Update the polling script to check phase=Completed before trusting result — add a PHASE check alongside the RESULT check so the script only exits when both phase=Completed AND result=pass. Or: (2) have set_healthcheck_pending explicitly null out result, message, rationale, completionTime, duration, modelUsed so there is no stale data window at all.

Expand All @@ -79,9 +134,8 @@ exit 1

One-time deploy checks catch immediate failures, but some problems only appear later — memory leaks, connection pool exhaustion, gradual performance degradation. [Scheduled Health Checks](scheduled-health-checks.md) run on a cron schedule to catch these regressions automatically.

## Tips for One-Time HealthChecks
## Tips

- **Version the check name** (e.g., `checkout-api-deploy-v2-4-1`) so each deploy creates a distinct resource and you keep an audit trail. This applies to one-time `HealthCheck` resources only — `ScheduledHealthCheck` resources use a fixed name and create child HealthChecks automatically.
- **Set a longer timeout** (60–120s) to give the rollout time to complete before Holmes evaluates.
- **Use labels** like `deploy-version` to query checks for a specific release: `kubectl get hc -l deploy-version=v2.4.1`.
- **Combine with ArgoCD**: If you use ArgoCD, the query can reference sync status — e.g., *"Is the ArgoCD application 'checkout-api' synced and healthy with no degraded resources?"* — since Holmes has access to the [ArgoCD toolset](../data-sources/builtin-toolsets/argocd.md).
- **Use labels** to query checks for a specific app: `kubectl get hc -l app=checkout-api`.
- **Combine with ArgoCD**: The query can reference sync status — e.g., *"Is the ArgoCD application 'checkout-api' synced and healthy with no degraded resources?"* — since Holmes has access to the [ArgoCD toolset](../data-sources/builtin-toolsets/argocd.md).
20 changes: 17 additions & 3 deletions docs/operator/health-checks.md
Original file line number Diff line number Diff line change
Expand Up @@ -124,6 +124,10 @@ After execution, the HealthCheck status contains:

### Execution Tracking

**observedGeneration** (integer)

The `metadata.generation` value that was last processed. When `metadata.generation != status.observedGeneration`, the operator will re-execute the check. This field is set on both successful and failed executions to prevent infinite retry loops.

**phase** (string)

Current execution state:
Expand Down Expand Up @@ -239,13 +243,23 @@ kubectl get hc check-pod-health -o jsonpath='{.status.rationale}'

## Re-running Checks

To re-execute a check, add the rerun annotation:
**Re-run on every apply (recommended for CI/CD):** Add `holmesgpt.dev/rerun: "true"` to your manifest. The operator clears the annotation after each run, so the next `kubectl apply` restores it and triggers a re-run. This works with Helm, ArgoCD, or plain manifests — see [Deployment Verification](deployment-verification.md) for examples.

```yaml
metadata:
annotations:
holmesgpt.dev/rerun: "true"
```

**Re-run on spec change:** When you modify any spec field (query, timeout, mode, etc.) and run `kubectl apply`, the operator detects the change via `metadata.generation` and re-executes automatically.

**Manual re-run via kubectl:** To re-execute a check ad-hoc:

```bash
kubectl annotate hc check-pod-health holmesgpt.dev/rerun=true --overwrite
kubectl annotate hc check-pod-health holmesgpt.dev/rerun=true
```

This triggers a new execution while preserving the original resource. The status will be updated with new results.
The annotation is cleared automatically after execution, so you can repeat this as many times as needed.

## Practical Examples

Expand Down
4 changes: 4 additions & 0 deletions helm/holmes/crds/healthcheck.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -66,6 +66,10 @@ spec:
status:
type: object
properties:
observedGeneration:
type: integer
format: int64
description: "Last spec generation that was processed. When metadata.generation != observedGeneration, the check will be re-executed"
phase:
type: string
description: "Execution phase: Pending, Running, Completed, Failed"
Expand Down
Loading
Loading