Skip to content

feat(alerts): complete alerting system overhaul — retry limits, HMAC signing, severity, resolution notifications - #2

Merged
AbdulmalikAlayande merged 9 commits into
devfrom
feat/alerting-system-overhaul
Jun 13, 2026
Merged

AbdulmalikAlayande merged 9 commits into
devfrom
feat/alerting-system-overhaul

Conversation

@AbdulmalikAlayande

@AbdulmalikAlayande AbdulmalikAlayande commented Jun 13, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Complete overhaul of the alerting system to address 10 identified reliability, security, and usability gaps. The previous alerting system had critical issues: infinite retry loops for failed deliveries, no webhook authentication, silent email channel acceptance that never delivered, missing resolution notifications, and no way to verify channel connectivity or view alert history.

This PR transforms the alerting system from a basic notification mechanism into a production-grade, fault-tolerant alerting pipeline.

What Changed

Critical Fixes (P0)

  • Retry limits with tracking — Failed alert deliveries now increment a retry_count column on alerts_fired. After 5 failed attempts (MAX_RETRY_COUNT), the alert is excluded from future delivery cycles. Previously, failed alerts retried every daemon cycle forever with no limit or backoff. DeliveryResult now reports abandoned count for observability.

  • Slack token wired from config — SentinelConfig had a slackToken field that was never used. The Slack handler always read SENTINEL_SLACK_TOKEN from env directly, creating a confusing disconnect. Now resolveSlackToken() checks env first, then falls back to config.slackToken from ~/.soroban-sentinel/config.yaml. Error message now mentions both configuration methods.

  • Email channel blocked — sentinel alerts add --type email now exits immediately with "Email alerting is not yet implemented. Use 'webhook' or 'slack'." Previously, email configs were silently accepted and stored in the database, counted as "attempted" during delivery, but never actually sent. The --email CLI option has been removed. Schema CHECK constraint updated to only allow slack and webhook.

Important Improvements (P1)

  • Resolution notifications — When a monitored entry's TTL recovers above the alert threshold (e.g., after a successful extension), the monitor now sends alert_resolved events to all channels that received the original threshold_crossed alert. Previously, resolveAlerts() only flipped resolved=1 in the database — users were told something was wrong but never told when it recovered. Resolution delivery is fire-and-forget (best-effort, failures logged but never block the cycle).

  • Webhook HMAC-SHA256 signing — Webhook payloads now include an X-Sentinel-Signature: sha256=<hex> header when a secret is configured. Receivers can recompute the HMAC to verify payload authenticity. Follows the same pattern used by GitHub, Stripe, and Slack webhooks. Secrets are auto-generated (32 random bytes, hex-encoded) when adding webhook configs and displayed once to the user. Can be overridden via --secret flag.

  • Alert severity levels — AlertEvent now includes a severity field (critical | warning | info). Critical = remaining TTL is at or below 25% of threshold or expired. Warning = below threshold but above 25%. Info = resolution events. Slack Block Kit formatting adapts: red circle emoji for critical, warning emoji for warning, green check for resolved. Previously all alerts looked identical regardless of urgency.

New CLI Commands (P2)

  • sentinel alerts test --id <configId> — Sends a synthetic test alert to verify channel connectivity. Validates that webhook URLs are reachable and Slack tokens have correct permissions. Reports success/failure immediately. Essential for production verification before relying on alerts.

  • sentinel alerts history --contract <id> [--limit N] — Shows chronological alert history with status icons (resolved/active, delivered/undelivered), TTL at fire time, channel info, retry count, and resolution timestamps. Previously alert data existed in the DB but was inaccessible to users.

Schema Changes

Table Column Type Purpose
alerts_fired retry_count INTEGER NOT NULL DEFAULT 0 Track delivery attempts per alert
alert_configs webhook_secret TEXT (nullable) HMAC signing secret per webhook config
alert_configs channel_type CHECK Removed 'email' Block unsupported channel type at DB level

All schema changes include live migrations in database.ts using ALTER TABLE with try/catch for idempotent upgrades of existing sentinel.db files.

Files Changed (15 files, +685/-155 lines)

Source (9 files)

  • src/db/schema.sql — Schema changes (retry_count, webhook_secret, email removal)
  • src/db/database.ts — Live migration statements for existing databases
  • src/db/repositories.ts — New functions (incrementRetryCount, getAlertConfigById, getAlertHistory), updated queries (retry cap, webhook_secret join), resolveAlerts now returns config IDs
  • src/alerts/types.ts — AlertSeverity type, computeSeverity(), severity in AlertEvent
  • src/alerts/webhook.ts — HMAC-SHA256 signing with X-Sentinel-Signature header
  • src/alerts/slack.ts — Token resolution from config, severity-aware Block Kit formatting
  • src/alerts/dispatcher.ts — Retry limits, incrementRetryCount on failure, deliverSingleAlert(), email channel removal
  • src/core/monitor.ts — Resolution notification delivery via deliverSingleAlert()
  • src/commands/alerts.ts — Email blocking, alerts test command, alerts history command, webhook secret generation

Tests (6 files)

  • tests/alerts/dispatcher.test.ts — Retry limits suite, severity assertions, webhook secret verification
  • tests/alerts/webhook.test.ts — HMAC signing suite (4 tests), severity in body structure
  • tests/alerts/slack.test.ts — Severity field, updated token error message pattern
  • tests/commands/alerts.test.ts — Email rejection test replaces email success test
  • tests/db/alert_delivery.test.ts — retryCount and webhookSecret field assertions

Commit Strategy

9 granular commits, each self-contained and individually reviewable:

  1. feat(db): add retry_count and webhook_secret columns to alert schema
  2. feat(db): add retry tracking, alert history, and webhook secret to repositories
  3. feat(alerts): add severity levels to AlertEvent
  4. feat(alerts): add HMAC-SHA256 signing to webhook deliveries
  5. feat(alerts): wire Slack token from config and add severity formatting
  6. feat(alerts): add retry limits and resolution delivery to dispatcher
  7. feat(core): send resolution notifications when alerts recover
  8. feat(cli): overhaul alerts command with test, history, and email blocking
  9. test(alerts): update test suites for alerting system overhaul

Test Plan

  • All 227 tests pass (0 failures, 5 skipped)
  • TypeScript compiles cleanly with tsc --noEmit
  • Verify sentinel alerts add --type webhook --contract <id> --url <url> --threshold 10000 auto-generates and displays HMAC secret
  • Verify sentinel alerts add --type email is rejected with clear error
  • Verify sentinel alerts test --id <configId> delivers test alert to webhook/Slack
  • Verify sentinel alerts history --contract <id> displays alert records
  • Verify failed webhook deliveries increment retry_count and stop after 5 attempts
  • Verify resolution notifications are sent when TTL recovers above threshold
  • Verify Slack token resolves from config.yaml when env var is not set

Breaking Changes

  • Email alert configs — Existing email alert configs in the database will no longer match the updated CHECK constraint on new databases. On existing databases, the constraint is preserved (SQLite limitation), but email configs will never be delivered since the email branch was removed from the dispatcher. Users should remove email configs and reconfigure as webhook or Slack.
  • AlertEvent payload — Now includes a severity field. Webhook consumers parsing the JSON body will see the new field. This is additive and should not break existing consumers.
  • DeliveryResult — Now includes an abandoned field (number). Code that destructures DeliveryResult will need to account for this.

AbdulmalikAlayande and others added 9 commits June 13, 2026 11:38
Add retry_count column (INTEGER NOT NULL DEFAULT 0) to alerts_fired table
to track how many delivery attempts have been made per alert. This enables
the dispatcher to enforce a maximum retry limit and stop retrying permanently
broken channels.

Add webhook_secret column (TEXT, nullable) to alert_configs table to store
per-config HMAC signing secrets for webhook authentication. Receivers can
verify payload authenticity using the X-Sentinel-Signature header.

Remove 'email' from alert_configs.channel_type CHECK constraint since email
alerting is not implemented — silently accepting email configs was misleading
users into thinking alerts would be delivered.

Add live migration statements to database.ts so existing sentinel.db files
created before these columns are seamlessly upgraded on next startup (ALTER
TABLE with try/catch for idempotency).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
…positories

Add incrementRetryCount() to bump retry_count on failed delivery attempts.

Add getAlertConfigById() for direct config lookup by ID, needed by the
monitor's resolution notification logic and the new 'alerts test' command.

Add getAlertHistory() with AlertHistoryRecord interface — returns fired alerts
joined with config and entry data for the 'alerts history' CLI command. Supports
optional limit parameter.

Update getUndeliveredAlerts() query:
- Filter out alerts where retry_count >= MAX_RETRY_COUNT (5) so permanently
  failed alerts stop being retried every cycle
- Include webhook_secret and retry_count in the joined result
- Update UndeliveredAlert interface to include new fields

Update insertAlertConfig() to accept optional webhook_secret parameter.

Update resolveAlerts() to return the list of alert_config_ids that were
resolved, enabling the monitor to send resolution notifications to the
correct channels.

Update AlertConfig interface: remove 'email' from channel_type union,
add webhook_secret field.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Introduce AlertSeverity type ("critical" | "warning" | "info") and
computeSeverity() function that classifies alerts based on how close
the entry is to expiry relative to the configured threshold:

- critical: remaining TTL is <= 0 or below 25% of threshold
- warning:  remaining TTL is below threshold but above 25%
- info:     used exclusively for resolution events

The severity field is now included in every AlertEvent payload, enabling
downstream consumers (Slack, webhooks) to prioritize and format alerts
differently based on urgency. Previously all alerts had identical urgency
regardless of how dangerously close to expiry the entry was.

buildAlertEvent() now calls computeSeverity() automatically based on the
event type and TTL values.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Webhook payloads previously had zero authentication — receivers could not
verify that a request actually came from Sentinel. Any entity that discovered
the webhook URL could send fake alert payloads.

sendWebhookAlert() now accepts an optional secret parameter. When provided:
1. Computes HMAC-SHA256 of the JSON body using the secret as key
2. Sends the signature as X-Sentinel-Signature: sha256=<hex digest>
3. Receivers can recompute the HMAC and compare to verify authenticity

This follows the same pattern used by GitHub webhooks, Stripe, and Slack
for payload verification. When no secret is configured, the header is
omitted for backwards compatibility.

Import node:crypto's createHmac for the signing operation.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Fix Slack token disconnect: SentinelConfig had a slackToken field that was
never used — the Slack handler always read SENTINEL_SLACK_TOKEN from env
directly. Now resolveSlackToken() checks both sources with clear priority:

1. SENTINEL_SLACK_TOKEN environment variable (takes precedence)
2. config.slackToken from ~/.soroban-sentinel/config.yaml (fallback)

This means users can configure their token once in config.yaml instead of
needing to export an env var before every daemon start. The error message
now mentions both configuration methods.

Add severity-aware formatting to Slack Block Kit messages:
- Critical alerts: red circle emoji (🔴) + "TTL CRITICAL" header
- Warning alerts: warning emoji (⚠️) + "TTL Warning" header
- Resolution events: green check (✅) + "Alert Resolved" header
- Footer now includes severity label for quick scanning

Previously all alerts looked identical regardless of urgency — a contract
1 ledger from expiry got the same formatting as one just barely below
threshold.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Fix infinite retry loop: previously failed alerts retried every daemon cycle
forever with no backoff or limit. If a webhook endpoint went permanently
offline, Sentinel would waste resources attempting delivery indefinitely.

Now the dispatcher:
1. Increments retry_count on each failed delivery via incrementRetryCount()
2. Alerts exceeding MAX_RETRY_COUNT (5) are excluded from getUndeliveredAlerts()
   queries automatically at the DB level
3. Reports abandoned count in DeliveryResult for observability
4. Logs escalating messages: warn for retryable failures, error for abandoned

Add deliverSingleAlert() function for sending one-off alerts outside the
batch delivery cycle. Used by the monitor to send resolution notifications
and by the new 'alerts test' CLI command. Returns boolean for simple
success/failure signaling.

Remove email channel from route() — email was silently skipped but counted
as "attempted", making delivery metrics misleading. Email configs are now
blocked at the CLI level.

Pass webhook_secret through to sendWebhookAlert() for HMAC signing.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Fix silent resolution: previously resolveAlerts() only flipped resolved=1
in the database. Users were told something was wrong (threshold_crossed)
but never told when it recovered. This left operators uncertain whether
manual intervention was still needed.

When an entry's TTL recovers above the threshold:
1. resolveAlerts() now returns the list of alert_config_ids that were resolved
2. For each resolved config, build an alert_resolved AlertEvent
3. Deliver via deliverSingleAlert() — fire-and-forget, best-effort
4. Resolution delivery failures are logged but never block the monitor cycle

This means every channel (webhook, Slack) that received a threshold_crossed
alert will also receive an alert_resolved event when the TTL is extended
back above the threshold, completing the alert lifecycle.

Pass network parameter through to processContract() for correct event
building (previously tried to access client.network which was private).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
…king

Block email channel: 'sentinel alerts add --type email' now exits with a
clear error message instead of silently creating a config that would never
deliver. Users are directed to use 'webhook' or 'slack' instead.

Add 'sentinel alerts test --id <configId>' command:
- Sends a synthetic test alert through the configured channel
- Validates webhook URLs are reachable and Slack tokens are valid
- Reports success/failure immediately
- Essential for verifying channel connectivity before relying on it
  in production

Add 'sentinel alerts history --contract <id> [--limit N]' command:
- Shows chronological alert history with delivery and resolution status
- Status icons: ✓ resolved / ● active, ✓ delivered / ✗ undelivered
- Includes TTL at fire time, channel info, and retry count
- Defaults to 20 most recent records

Add webhook HMAC secret support to 'alerts add':
- Webhooks auto-generate a 32-byte hex secret on creation
- Secret is displayed once and must be saved by the user
- Can be overridden via --secret flag
- 'alerts list' shows [signed] indicator for configs with secrets

Remove --email option from 'alerts add' since the channel is blocked.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
dispatcher.test.ts:
- Add retry limits test suite: verifies alerts stop retrying after
  MAX_RETRY_COUNT failures and reports abandoned count
- Update DeliveryResult assertions to include abandoned field
- Remove email channel routing test (email no longer supported)
- Update webhook call assertions to verify secret parameter is passed
- Add severity field assertion to payload correctness tests
- Update seedContractWithAlert helper to accept webhookSecret

webhook.test.ts:
- Add HMAC signing test suite: verifies X-Sentinel-Signature header
  presence/absence, correct sha256 digest computation, and null secret
  handling
- Add severity field to makeAlertEvent helper
- Update body structure test to include severity in expected keys

slack.test.ts:
- Add severity field to makeAlertEvent helper
- Update token validation tests to match new error message pattern
  (now mentions both env var and config.yaml as token sources)

alerts.test.ts (commands):
- Replace email success test with email rejection test — verifies
  'alerts add --type email' exits with "not yet implemented" error

alert_delivery.test.ts:
- Update channelType union to remove 'email'
- Add retryCount and webhookSecret assertions to shape tests

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Jun 13, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 7953a13c-d4e5-44e3-93ca-471876c2ac1b

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/alerting-system-overhaul

Comment @coderabbitai help to get the list of available commands and usage tips.

@AbdulmalikAlayande
AbdulmalikAlayande merged commit 09bbf7b into dev Jun 13, 2026
3 checks passed
AbdulmalikAlayande added a commit that referenced this pull request Aug 2, 2026
…erhaul

feat(alerts): complete alerting system overhaul — retry limits, HMAC signing, severity, resolution notifications
AbdulmalikAlayande added a commit that referenced this pull request Aug 2, 2026
…erhaul

feat(alerts): complete alerting system overhaul — retry limits, HMAC signing, severity, resolution notifications
AbdulmalikAlayande pushed a commit that referenced this pull request Aug 7, 2026
Adds a new periodic fleet-wide health digest that teams can configure
instead of (or alongside) per-entry threshold alerts.

Schema
- New digest_configs table (separate from alert_configs — no threshold_ledgers,
  no alert_config_id FK; has interval_ms instead). Justified separately because
  a digest is semantically different from a per-entry TTL alert.

src/core/digest.ts (new)
- DigestPayload type — deliberately NOT part of the AlertEvent discriminated
  union in alerts/types.ts, per the issue's explicit instruction.
- buildFleetDigest(db, network, currentLedger, options?) — reads live DB state
  synchronously, classifying every tracked entry by severity (critical/warning/ok),
  computing topExpiring contracts sorted by min remaining TTL, and summing
  extension costs for the period.

src/db/repositories.ts
- insertDigestConfig / getDigestConfigs repository helpers.
- DigestConfig interface.

src/daemon/loop.ts
- digestIntervalMs option added to DaemonOptions.
- Module-level lastDigestAt timestamp gate, following runScheduledVacuum pattern.
- runScheduledDigest() fires buildFleetDigest + deliverSingleAlert for each
  enabled digest_configs row when the interval has elapsed.
- Called from scheduledTick() before executeCycle(), isolated from cycle errors.

Tests (TDD — tests written before implementation)
- tests/core/digest.test.ts — 23 tests covering payload shape, severity
  classification, network isolation, topExpiring ordering/cap, cost aggregation,
  live fleet-state accuracy (AC #2), inactive contract exclusion, and the
  digest_configs repository functions.
- tests/daemon/digest-loop.test.ts — 8 tests covering AC #1 (fires once per
  interval, not once per monitor cycle), multi-config delivery, delivery failure
  isolation, and network filtering.

All 1342 existing tests continue to pass. Build is clean.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant