Repository navigation
feat(alerts): complete alerting system overhaul — retry limits, HMAC signing, severity, resolution notifications - #2
Merged
Conversation
Add retry_count column (INTEGER NOT NULL DEFAULT 0) to alerts_fired table to track how many delivery attempts have been made per alert. This enables the dispatcher to enforce a maximum retry limit and stop retrying permanently broken channels. Add webhook_secret column (TEXT, nullable) to alert_configs table to store per-config HMAC signing secrets for webhook authentication. Receivers can verify payload authenticity using the X-Sentinel-Signature header. Remove 'email' from alert_configs.channel_type CHECK constraint since email alerting is not implemented — silently accepting email configs was misleading users into thinking alerts would be delivered. Add live migration statements to database.ts so existing sentinel.db files created before these columns are seamlessly upgraded on next startup (ALTER TABLE with try/catch for idempotency). Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
…positories Add incrementRetryCount() to bump retry_count on failed delivery attempts. Add getAlertConfigById() for direct config lookup by ID, needed by the monitor's resolution notification logic and the new 'alerts test' command. Add getAlertHistory() with AlertHistoryRecord interface — returns fired alerts joined with config and entry data for the 'alerts history' CLI command. Supports optional limit parameter. Update getUndeliveredAlerts() query: - Filter out alerts where retry_count >= MAX_RETRY_COUNT (5) so permanently failed alerts stop being retried every cycle - Include webhook_secret and retry_count in the joined result - Update UndeliveredAlert interface to include new fields Update insertAlertConfig() to accept optional webhook_secret parameter. Update resolveAlerts() to return the list of alert_config_ids that were resolved, enabling the monitor to send resolution notifications to the correct channels. Update AlertConfig interface: remove 'email' from channel_type union, add webhook_secret field. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Introduce AlertSeverity type ("critical" | "warning" | "info") and
computeSeverity() function that classifies alerts based on how close
the entry is to expiry relative to the configured threshold:
- critical: remaining TTL is <= 0 or below 25% of threshold
- warning: remaining TTL is below threshold but above 25%
- info: used exclusively for resolution events
The severity field is now included in every AlertEvent payload, enabling
downstream consumers (Slack, webhooks) to prioritize and format alerts
differently based on urgency. Previously all alerts had identical urgency
regardless of how dangerously close to expiry the entry was.
buildAlertEvent() now calls computeSeverity() automatically based on the
event type and TTL values.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Webhook payloads previously had zero authentication — receivers could not verify that a request actually came from Sentinel. Any entity that discovered the webhook URL could send fake alert payloads. sendWebhookAlert() now accepts an optional secret parameter. When provided: 1. Computes HMAC-SHA256 of the JSON body using the secret as key 2. Sends the signature as X-Sentinel-Signature: sha256=<hex digest> 3. Receivers can recompute the HMAC and compare to verify authenticity This follows the same pattern used by GitHub webhooks, Stripe, and Slack for payload verification. When no secret is configured, the header is omitted for backwards compatibility. Import node:crypto's createHmac for the signing operation. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Fix Slack token disconnect: SentinelConfig had a slackToken field that was never used — the Slack handler always read SENTINEL_SLACK_TOKEN from env directly. Now resolveSlackToken() checks both sources with clear priority: 1. SENTINEL_SLACK_TOKEN environment variable (takes precedence) 2. config.slackToken from ~/.soroban-sentinel/config.yaml (fallback) This means users can configure their token once in config.yaml instead of needing to export an env var before every daemon start. The error message now mentions both configuration methods. Add severity-aware formatting to Slack Block Kit messages: - Critical alerts: red circle emoji (🔴) + "TTL CRITICAL" header - Warning alerts: warning emoji (⚠️ ) + "TTL Warning" header - Resolution events: green check (✅) + "Alert Resolved" header - Footer now includes severity label for quick scanning Previously all alerts looked identical regardless of urgency — a contract 1 ledger from expiry got the same formatting as one just barely below threshold. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Fix infinite retry loop: previously failed alerts retried every daemon cycle forever with no backoff or limit. If a webhook endpoint went permanently offline, Sentinel would waste resources attempting delivery indefinitely. Now the dispatcher: 1. Increments retry_count on each failed delivery via incrementRetryCount() 2. Alerts exceeding MAX_RETRY_COUNT (5) are excluded from getUndeliveredAlerts() queries automatically at the DB level 3. Reports abandoned count in DeliveryResult for observability 4. Logs escalating messages: warn for retryable failures, error for abandoned Add deliverSingleAlert() function for sending one-off alerts outside the batch delivery cycle. Used by the monitor to send resolution notifications and by the new 'alerts test' CLI command. Returns boolean for simple success/failure signaling. Remove email channel from route() — email was silently skipped but counted as "attempted", making delivery metrics misleading. Email configs are now blocked at the CLI level. Pass webhook_secret through to sendWebhookAlert() for HMAC signing. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Fix silent resolution: previously resolveAlerts() only flipped resolved=1 in the database. Users were told something was wrong (threshold_crossed) but never told when it recovered. This left operators uncertain whether manual intervention was still needed. When an entry's TTL recovers above the threshold: 1. resolveAlerts() now returns the list of alert_config_ids that were resolved 2. For each resolved config, build an alert_resolved AlertEvent 3. Deliver via deliverSingleAlert() — fire-and-forget, best-effort 4. Resolution delivery failures are logged but never block the monitor cycle This means every channel (webhook, Slack) that received a threshold_crossed alert will also receive an alert_resolved event when the TTL is extended back above the threshold, completing the alert lifecycle. Pass network parameter through to processContract() for correct event building (previously tried to access client.network which was private). Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
…king Block email channel: 'sentinel alerts add --type email' now exits with a clear error message instead of silently creating a config that would never deliver. Users are directed to use 'webhook' or 'slack' instead. Add 'sentinel alerts test --id <configId>' command: - Sends a synthetic test alert through the configured channel - Validates webhook URLs are reachable and Slack tokens are valid - Reports success/failure immediately - Essential for verifying channel connectivity before relying on it in production Add 'sentinel alerts history --contract <id> [--limit N]' command: - Shows chronological alert history with delivery and resolution status - Status icons: ✓ resolved / ● active, ✓ delivered / ✗ undelivered - Includes TTL at fire time, channel info, and retry count - Defaults to 20 most recent records Add webhook HMAC secret support to 'alerts add': - Webhooks auto-generate a 32-byte hex secret on creation - Secret is displayed once and must be saved by the user - Can be overridden via --secret flag - 'alerts list' shows [signed] indicator for configs with secrets Remove --email option from 'alerts add' since the channel is blocked. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
dispatcher.test.ts: - Add retry limits test suite: verifies alerts stop retrying after MAX_RETRY_COUNT failures and reports abandoned count - Update DeliveryResult assertions to include abandoned field - Remove email channel routing test (email no longer supported) - Update webhook call assertions to verify secret parameter is passed - Add severity field assertion to payload correctness tests - Update seedContractWithAlert helper to accept webhookSecret webhook.test.ts: - Add HMAC signing test suite: verifies X-Sentinel-Signature header presence/absence, correct sha256 digest computation, and null secret handling - Add severity field to makeAlertEvent helper - Update body structure test to include severity in expected keys slack.test.ts: - Add severity field to makeAlertEvent helper - Update token validation tests to match new error message pattern (now mentions both env var and config.yaml as token sources) alerts.test.ts (commands): - Replace email success test with email rejection test — verifies 'alerts add --type email' exits with "not yet implemented" error alert_delivery.test.ts: - Update channelType union to remove 'email' - Add retryCount and webhookSecret assertions to shape tests Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Organization UI Review profile: ASSERTIVE Plan: Pro Plus Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
AbdulmalikAlayande
added a commit
that referenced
this pull request
Aug 2, 2026
…erhaul feat(alerts): complete alerting system overhaul — retry limits, HMAC signing, severity, resolution notifications
AbdulmalikAlayande
added a commit
that referenced
this pull request
Aug 2, 2026
…erhaul feat(alerts): complete alerting system overhaul — retry limits, HMAC signing, severity, resolution notifications
AbdulmalikAlayande
pushed a commit
that referenced
this pull request
Aug 7, 2026
Adds a new periodic fleet-wide health digest that teams can configure instead of (or alongside) per-entry threshold alerts. Schema - New digest_configs table (separate from alert_configs — no threshold_ledgers, no alert_config_id FK; has interval_ms instead). Justified separately because a digest is semantically different from a per-entry TTL alert. src/core/digest.ts (new) - DigestPayload type — deliberately NOT part of the AlertEvent discriminated union in alerts/types.ts, per the issue's explicit instruction. - buildFleetDigest(db, network, currentLedger, options?) — reads live DB state synchronously, classifying every tracked entry by severity (critical/warning/ok), computing topExpiring contracts sorted by min remaining TTL, and summing extension costs for the period. src/db/repositories.ts - insertDigestConfig / getDigestConfigs repository helpers. - DigestConfig interface. src/daemon/loop.ts - digestIntervalMs option added to DaemonOptions. - Module-level lastDigestAt timestamp gate, following runScheduledVacuum pattern. - runScheduledDigest() fires buildFleetDigest + deliverSingleAlert for each enabled digest_configs row when the interval has elapsed. - Called from scheduledTick() before executeCycle(), isolated from cycle errors. Tests (TDD — tests written before implementation) - tests/core/digest.test.ts — 23 tests covering payload shape, severity classification, network isolation, topExpiring ordering/cap, cost aggregation, live fleet-state accuracy (AC #2), inactive contract exclusion, and the digest_configs repository functions. - tests/daemon/digest-loop.test.ts — 8 tests covering AC #1 (fires once per interval, not once per monitor cycle), multi-config delivery, delivery failure isolation, and network filtering. All 1342 existing tests continue to pass. Build is clean.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Complete overhaul of the alerting system to address 10 identified reliability, security, and usability gaps. The previous alerting system had critical issues: infinite retry loops for failed deliveries, no webhook authentication, silent email channel acceptance that never delivered, missing resolution notifications, and no way to verify channel connectivity or view alert history.
This PR transforms the alerting system from a basic notification mechanism into a production-grade, fault-tolerant alerting pipeline.
What Changed
Critical Fixes (P0)
Retry limits with tracking — Failed alert deliveries now increment a
retry_countcolumn onalerts_fired. After 5 failed attempts (MAX_RETRY_COUNT), the alert is excluded from future delivery cycles. Previously, failed alerts retried every daemon cycle forever with no limit or backoff.DeliveryResultnow reportsabandonedcount for observability.Slack token wired from config —
SentinelConfighad aslackTokenfield that was never used. The Slack handler always readSENTINEL_SLACK_TOKENfrom env directly, creating a confusing disconnect. NowresolveSlackToken()checks env first, then falls back toconfig.slackTokenfrom~/.soroban-sentinel/config.yaml. Error message now mentions both configuration methods.Email channel blocked —
sentinel alerts add --type emailnow exits immediately with "Email alerting is not yet implemented. Use 'webhook' or 'slack'." Previously, email configs were silently accepted and stored in the database, counted as "attempted" during delivery, but never actually sent. The--emailCLI option has been removed. Schema CHECK constraint updated to only allowslackandwebhook.Important Improvements (P1)
Resolution notifications — When a monitored entry's TTL recovers above the alert threshold (e.g., after a successful extension), the monitor now sends
alert_resolvedevents to all channels that received the originalthreshold_crossedalert. Previously,resolveAlerts()only flippedresolved=1in the database — users were told something was wrong but never told when it recovered. Resolution delivery is fire-and-forget (best-effort, failures logged but never block the cycle).Webhook HMAC-SHA256 signing — Webhook payloads now include an
X-Sentinel-Signature: sha256=<hex>header when a secret is configured. Receivers can recompute the HMAC to verify payload authenticity. Follows the same pattern used by GitHub, Stripe, and Slack webhooks. Secrets are auto-generated (32 random bytes, hex-encoded) when adding webhook configs and displayed once to the user. Can be overridden via--secretflag.Alert severity levels —
AlertEventnow includes aseverityfield (critical|warning|info). Critical = remaining TTL is at or below 25% of threshold or expired. Warning = below threshold but above 25%. Info = resolution events. Slack Block Kit formatting adapts: red circle emoji for critical, warning emoji for warning, green check for resolved. Previously all alerts looked identical regardless of urgency.New CLI Commands (P2)
sentinel alerts test --id <configId>— Sends a synthetic test alert to verify channel connectivity. Validates that webhook URLs are reachable and Slack tokens have correct permissions. Reports success/failure immediately. Essential for production verification before relying on alerts.sentinel alerts history --contract <id> [--limit N]— Shows chronological alert history with status icons (resolved/active, delivered/undelivered), TTL at fire time, channel info, retry count, and resolution timestamps. Previously alert data existed in the DB but was inaccessible to users.Schema Changes
alerts_firedretry_countINTEGER NOT NULL DEFAULT 0alert_configswebhook_secretTEXT(nullable)alert_configschannel_typeCHECK'email'All schema changes include live migrations in
database.tsusingALTER TABLEwith try/catch for idempotent upgrades of existingsentinel.dbfiles.Files Changed (15 files, +685/-155 lines)
Source (9 files)
src/db/schema.sql— Schema changes (retry_count, webhook_secret, email removal)src/db/database.ts— Live migration statements for existing databasessrc/db/repositories.ts— New functions (incrementRetryCount, getAlertConfigById, getAlertHistory), updated queries (retry cap, webhook_secret join), resolveAlerts now returns config IDssrc/alerts/types.ts— AlertSeverity type, computeSeverity(), severity in AlertEventsrc/alerts/webhook.ts— HMAC-SHA256 signing with X-Sentinel-Signature headersrc/alerts/slack.ts— Token resolution from config, severity-aware Block Kit formattingsrc/alerts/dispatcher.ts— Retry limits, incrementRetryCount on failure, deliverSingleAlert(), email channel removalsrc/core/monitor.ts— Resolution notification delivery via deliverSingleAlert()src/commands/alerts.ts— Email blocking, alerts test command, alerts history command, webhook secret generationTests (6 files)
tests/alerts/dispatcher.test.ts— Retry limits suite, severity assertions, webhook secret verificationtests/alerts/webhook.test.ts— HMAC signing suite (4 tests), severity in body structuretests/alerts/slack.test.ts— Severity field, updated token error message patterntests/commands/alerts.test.ts— Email rejection test replaces email success testtests/db/alert_delivery.test.ts— retryCount and webhookSecret field assertionsCommit Strategy
9 granular commits, each self-contained and individually reviewable:
feat(db): add retry_count and webhook_secret columns to alert schemafeat(db): add retry tracking, alert history, and webhook secret to repositoriesfeat(alerts): add severity levels to AlertEventfeat(alerts): add HMAC-SHA256 signing to webhook deliveriesfeat(alerts): wire Slack token from config and add severity formattingfeat(alerts): add retry limits and resolution delivery to dispatcherfeat(core): send resolution notifications when alerts recoverfeat(cli): overhaul alerts command with test, history, and email blockingtest(alerts): update test suites for alerting system overhaulTest Plan
tsc --noEmitsentinel alerts add --type webhook --contract <id> --url <url> --threshold 10000auto-generates and displays HMAC secretsentinel alerts add --type emailis rejected with clear errorsentinel alerts test --id <configId>delivers test alert to webhook/Slacksentinel alerts history --contract <id>displays alert recordsBreaking Changes
emailalert configs in the database will no longer match the updated CHECK constraint on new databases. On existing databases, the constraint is preserved (SQLite limitation), but email configs will never be delivered since the email branch was removed from the dispatcher. Users should remove email configs and reconfigure as webhook or Slack.severityfield. Webhook consumers parsing the JSON body will see the new field. This is additive and should not break existing consumers.abandonedfield (number). Code that destructures DeliveryResult will need to account for this.