Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions docs/operator/.nav.yml
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,7 @@ nav:
- Deployment Verification: deployment-verification.md
- Health Checks: health-checks.md
- Scheduled Health Checks: scheduled-health-checks.md
- Triggered Health Checks: triggered-health-checks.md
- Alert Destinations: destinations.md
- Configuration: configuration.md
- Development Guide: development.md
107 changes: 90 additions & 17 deletions docs/operator/deployment-verification.md
Original file line number Diff line number Diff line change
@@ -1,10 +1,75 @@
# Deployment Verification

A common pattern is deploying a HealthCheck alongside your application to verify the new version is working correctly. Since HealthChecks run immediately when created, you can include one in the same manifest (or CI/CD step) as your deployment and use the result to gate rollout progression.
Verifying that a new version is healthy right after it ships is one of the most common
uses of the operator. There are two ways to do it, and they compose:

## One-Time Verification with HealthCheck
- **Automatically, on every rollout** — declare one [TriggeredHealthCheck](triggered-health-checks.md)
and Holmes investigates *every* future rollout of the service, no matter how it was
triggered (CI, GitOps/Argo sync, `kubectl set image`, or a rollback). This is the
recommended default — declare once, no per-deploy wiring.
- **Inline, to gate a pipeline** — include a one-time [HealthCheck](health-checks.md) in
the deploy manifest (or CI/CD step) and block the pipeline on its result. Use this when
CI must wait synchronously for the verdict before proceeding.

Include a [HealthCheck](health-checks.md) in the same manifest as your deployment. It runs immediately after `kubectl apply` and reports whether the new version started correctly.
A typical setup uses both: a `TriggeredHealthCheck` for hands-off coverage of all
rollouts, plus an inline `HealthCheck` in the specific pipeline stage where you want a
hard gate.

## Automatic verification with TriggeredHealthCheck

Apply this once. From then on, any rollout of a Deployment matching the selector
automatically spawns a check — including deploys you didn't make through CI.

```yaml
# verify-checkout-deploys.yaml — apply once, verifies every future rollout
apiVersion: holmesgpt.dev/v1alpha1
kind: TriggeredHealthCheck
metadata:
name: verify-checkout-deploys
namespace: production
spec:
deploymentRollout:
selector:
matchLabels:
app: checkout-api
delaySeconds: 300 # wait 5m after the rollout, then check (default)
query: |
checkout-api was just rolled out to {{ .new.image }} (previously {{ .old.image }}).
Is the new version healthy? Compare error rates, latency, restarts, and logs
before vs after the rollout and flag any regressions.
timeout: 120
mode: alert
destinations:
- type: slack
config:
channel: "#deploy-alerts"
```

```bash
kubectl apply -f verify-checkout-deploys.yaml

# After a deploy, see the check it produced and the verdict
kubectl get hc -n production -l holmesgpt.dev/triggered-by=verify-checkout-deploys
kubectl describe thc verify-checkout-deploys -n production
```

The `{{ .new.image }}` / `{{ .old.image }}` tokens are substituted with the rollout's
before/after images, so the investigation knows exactly what changed. See
[Triggered Health Checks](triggered-health-checks.md) for the full field reference and
[how long the check waits](triggered-health-checks.md#how-long-to-wait) after a rollout.

!!! tip "Catch slow-burn regressions too"

Some problems (memory leaks, connection-pool exhaustion) only appear after the new
version has run for a while. Add a second trigger with a delay — e.g.
`delaySeconds: 86400` — to re-investigate the same rollout a day later, or use a
[ScheduledHealthCheck](scheduled-health-checks.md) for continuous coverage.

## Gating CI/CD with an inline HealthCheck

When a pipeline must **wait for the verdict** before promoting a release, include a
one-time `HealthCheck` in the same manifest as your deployment. It runs immediately after
`kubectl apply` and reports whether the new version started correctly.

```yaml
# app-deployment.yaml
Expand Down Expand Up @@ -45,17 +110,11 @@ spec:
channel: "#deploy-alerts"
```

Apply both together:

```bash
kubectl apply -f app-deployment.yaml
```

If pods crash or fail readiness, the check fails and alerts your team.

## Gating CI/CD on the Result

After applying, poll for the result to gate your pipeline:
Then poll for the result to gate the pipeline:

```bash
# Wait for the check to complete, then read the result
Expand All @@ -75,13 +134,27 @@ echo "Timed out waiting for health check"
exit 1
```

## Ongoing Monitoring with ScheduledHealthCheck
If pods crash or fail readiness, the check fails and the pipeline stops.

## When to use which

One-time deploy checks catch immediate failures, but some problems only appear later — memory leaks, connection pool exhaustion, gradual performance degradation. [Scheduled Health Checks](scheduled-health-checks.md) run on a cron schedule to catch these regressions automatically.
| | TriggeredHealthCheck | Inline HealthCheck |
|---|---|---|
| Runs on | *Every* rollout, automatically | Only when you apply it |
| Setup | Declare once per service | Added to each deploy manifest/step |
| Covers out-of-band deploys (`kubectl set image`, GitOps, rollback) | Yes | No |
| Blocks a CI/CD pipeline | No (fire-and-forget) | Yes (poll the result to gate) |

## Tips for One-Time HealthChecks
## Tips

- **Version the check name** (e.g., `checkout-api-deploy-v2-4-1`) so each deploy creates a distinct resource and you keep an audit trail. This applies to one-time `HealthCheck` resources only — `ScheduledHealthCheck` resources use a fixed name and create child HealthChecks automatically.
- **Set a longer timeout** (60–120s) to give the rollout time to complete before Holmes evaluates.
- **Use labels** like `deploy-version` to query checks for a specific release: `kubectl get hc -l deploy-version=v2.4.1`.
- **Combine with ArgoCD**: If you use ArgoCD, the query can reference sync status — e.g., *"Is the ArgoCD application 'checkout-api' synced and healthy with no degraded resources?"* — since Holmes has access to the [ArgoCD toolset](../data-sources/builtin-toolsets/argocd.md).
- **Version the inline check name** (e.g., `checkout-api-deploy-v2-4-1`) so each deploy
creates a distinct resource and you keep an audit trail. This applies to one-time
`HealthCheck` resources only — `TriggeredHealthCheck` and `ScheduledHealthCheck` use a
fixed name and create child HealthChecks automatically.
- **Set a longer timeout** (60–120s) to give the investigation time to gather data.
- **Use labels** like `deploy-version` to query checks for a specific release:
`kubectl get hc -l deploy-version=v2.4.1`.
- **Combine with ArgoCD**: the query can reference sync status — e.g., *"Is the ArgoCD
application 'checkout-api' synced and healthy with no degraded resources?"* — since
Holmes has access to the [ArgoCD toolset](../data-sources/builtin-toolsets/argocd.md).
</content>
2 changes: 2 additions & 0 deletions docs/operator/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,7 @@ Under the hood, it uses Kubernetes CRDs to declaratively define one-time and sch
- **[Deployment Verification](deployment-verification.md)**: Deploy a HealthCheck alongside your app to verify the new version is healthy — and gate CI/CD on the result
- **[One-time Health Checks](health-checks.md)**: Create `HealthCheck` resources that run immediately and report results
- **[Scheduled Health Checks](scheduled-health-checks.md)**: Create `ScheduledHealthCheck` resources that run on cron schedules for continuous monitoring
- **[Triggered Health Checks](triggered-health-checks.md)**: Create `TriggeredHealthCheck` resources that run automatically whenever a matching Deployment is rolled out
- **Not just Kubernetes**: Health checks can query any connected data source — Prometheus, Datadog, AWS, databases, and [more](../data-sources/builtin-toolsets/index.md)
- **Kubernetes-native**: Uses standard CRDs with kubectl support
- **Status Tracking**: Full execution history and results stored in resource status
Expand Down Expand Up @@ -154,6 +155,7 @@ kubectl describe hc example-check
- **[Deployment Verification](deployment-verification.md)** - Verify new deploys are healthy and gate CI/CD pipelines on the result
- **[Health Checks](health-checks.md)** - Learn how to create and manage one-time HealthCheck resources
- **[Scheduled Health Checks](scheduled-health-checks.md)** - Set up recurring health checks with cron schedules
- **[Triggered Health Checks](triggered-health-checks.md)** - Run checks automatically on every Deployment rollout
- **[Alert Destinations](destinations.md)** - Configure Slack and PagerDuty notifications
- **[Configuration](configuration.md)** - Explore advanced configuration options
- **[Development Guide](development.md)** - Build and test operator changes locally
Expand Down
152 changes: 152 additions & 0 deletions docs/operator/triggered-health-checks.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,152 @@
# Triggered Health Checks

A `TriggeredHealthCheck` runs an investigation **automatically when a Deployment rolls
out a new version** — no per-deploy wiring, no CI polling. Declare it once, and every
rollout of a matching Deployment (from CI, Argo, `kubectl set image`, or a rollback)
fires a check.

It is the event-driven sibling of the [ScheduledHealthCheck](scheduled-health-checks.md):
both are self-contained (they embed the check definition inline) and both spawn a
[HealthCheck](health-checks.md) per run, which becomes the execution record.

!!! info "Alpha"

`TriggeredHealthCheck` currently supports a single trigger type — `deploymentRollout`.
More event sources (pod crashloops, failed Jobs, alerts) are planned.

## How it works

1. The operator watches Deployments in namespaces where `TriggeredHealthCheck`
resources exist.
2. When a matching Deployment's **pod template changes** (a rollout), the operator waits
`delaySeconds` (default 5 minutes) and then runs the check. The wait gives the rollout
time to finish and gives any crashes or errors time to show up.
3. It creates a `HealthCheck` (owned by the trigger) with your query, having
substituted the rollout context into it.
4. Holmes investigates using every connected data source; in `alert` mode it notifies
your [destinations](destinations.md) on failure.

## Example

```yaml
Comment thread
aantn marked this conversation as resolved.
apiVersion: holmesgpt.dev/v1alpha1
kind: TriggeredHealthCheck
metadata:
name: verify-checkout-rollouts
namespace: production
spec:
deploymentRollout:
selector:
matchLabels:
app: checkout-api
delaySeconds: 300 # wait 5m after the rollout, then check (default)
cooldownSeconds: 600 # don't re-fire for the same Deployment within 10m
query: |
checkout-api was rolled out to {{ .new.image }} (was {{ .old.image }}).
Compare error rates, latency, restarts, and logs before vs after the rollout
and flag any regressions.
timeout: 120
mode: alert
destinations:
- type: slack
config:
channel: "#deploy-alerts"
```

Apply it once:

```bash
kubectl apply -f triggeredhealthcheck.yaml

# List triggers (short name: thc)
kubectl get thc

# See fire history and the HealthChecks each rollout produced
kubectl describe thc verify-checkout-rollouts
kubectl get hc -l holmesgpt.dev/triggered-by=verify-checkout-rollouts
```

## Query and rollout context

The rollout facts are **always** prepended to the query the check runs, so even a terse
query like `"Is the new version healthy?"` gives the model what it needs:

```
This health check was triggered automatically by a Kubernetes Deployment rollout. Use this context when investigating:
- Deployment: checkout-api
- Namespace: production
- Previous image(s): myregistry/checkout-api:v2.4.0
- New image(s): myregistry/checkout-api:v2.4.1

<your query>
```

You can also reference the same facts inline in your `query` with these tokens:

| Token | Replaced with |
|-------|---------------|
| `{{ .deployment }}` | Name of the Deployment that rolled out |
| `{{ .namespace }}` | Its namespace |
| `{{ .old.image }}` | Container image(s) before the rollout |
| `{{ .new.image }}` | Container image(s) after the rollout |

## Spec reference

| Field | Default | Description |
|-------|---------|-------------|
| `enabled` | `true` | Whether the trigger is active |
| `deploymentRollout.selector.matchLabels` | `{}` | Deployment labels that must all match. **Empty matches every Deployment in the namespace.** |
| `delaySeconds` | `300` | How long to wait after a rollout before running the check. Default 5 minutes; `0` checks immediately; `86400` checks a day later (max 7 days). See [How long to wait](#how-long-to-wait). |
| `cooldownSeconds` | `0` | Suppress re-firing for the same Deployment within this window. `0` disables. |
| `query` | — | Natural-language investigation (supports the tokens above). Required. |
| `timeout` | `120` | Check execution timeout in seconds. |
| `mode` | `monitor` | `alert` notifies destinations on failure; `monitor` only records the result. |
| `model` | — | Override the default LLM model for this check. |
| `destinations` | `[]` | Alert destinations (used in `alert` mode). See [Destinations](destinations.md). |

## How long to wait

There is one knob: **`delaySeconds`** — how long to wait after a rollout before running
the check. That's it.

The check does not run the instant a new version is deployed, because a brand-new
rollout hasn't finished and problems haven't surfaced yet. So Holmes waits `delaySeconds`
first, then looks. Pick a value that fits what you want to catch:

- **`300` (default, 5 minutes)** — enough time for the rollout to finish and for crashes
or startup errors to appear.
- **`0`** — check right away (useful if you only care that the deploy was accepted).
- **`86400` (a day)** — catch slower problems like memory leaks or resource creep that
only show up after the version has been running a while.

If you set the wait too short for your app's rollout, the check may run while pods are
still starting — Holmes will simply report that the rollout hasn't finished yet.

The wait is saved on the resource, so it still completes if the operator restarts — even
a day-long wait. If the same Deployment is rolled out again while a check is still
waiting, the waiting check is replaced so only the newest version is checked.

A common setup is one trigger with the default 5-minute wait, plus a second with
`delaySeconds: 86400` to re-check the same rollout a day later.

## Notes & limitations

- **Rollout = pod-template change.** Scaling and HPA changes (which only touch
`spec.replicas`) do **not** fire the trigger; only changes to the pod template do.
- **Restart behavior.** A check that was *already scheduled* before a restart still runs
(the pending fire is persisted in `status.pending`). *Detecting* new rollouts, however,
relies on an in-memory baseline of each Deployment's last-seen pod template (the operator
does not annotate your Deployments). After a restart the first observation of each
Deployment just re-establishes that baseline, so a rollout that happens *during* the
restart window is not detected. Use a
[ScheduledHealthCheck](scheduled-health-checks.md) for continuous coverage.
- **Cost.** Every fire is at least one LLM call. Use `cooldownSeconds` and a specific
`selector` to bound spend on busy namespaces.

## Next Steps

- **[Deployment Verification](deployment-verification.md)** — patterns for gating and
verifying deploys
- **[Health Checks](health-checks.md)** — the one-time checks this spawns
- **[Alert Destinations](destinations.md)** — Slack and PagerDuty configuration
</content>
Loading
Loading