Skip to content

Add a public status page: status.heykody.dev worker + /health/components + operator email alerts - #1230

Merged
kody-bot merged 7 commits into
mainfrom
cursor/status-page-da8d
Aug 5, 2026
Merged

kody-bot merged 7 commits into
mainfrom
cursor/status-page-da8d

Conversation

@kentcdodds

@kentcdodds kentcdodds commented Aug 4, 2026 •

Copy link
Copy Markdown
Owner

Public status page for kody

Implements the status page per decision record docs/contributing/decisions/0004-status-page-separate-worker.md: an independently deployed worker at status.heykody.dev that observes the product from the outside, so it stays up when the main worker's deploys, code, or database are broken.

What's in here

1. Component-level health endpoint (main worker). New public GET /health/components runs cheap per-binding checks (APP_DB, AUDIT_DB, OAUTH_KV, COMMUNITY_ASSETS) with 3s timeouts, memoizes results for 10s (with coalesced concurrent refresh) so public traffic cannot amplify load, and returns 503 when any component fails.

2. Status worker (packages/status/). Mirrors the backup-control-plane package shape. A cron trigger runs one probe pass per minute against public endpoints only: /health, /health/components, the unauthenticated OAuth bearer challenge on /mcp, and kodyapps.dev. A single StatusStore Durable Object (SQLite) keeps minute samples (24h), daily uptime rollups (90 days), incidents, and notification history — deliberately never APP_DB. Components open an incident after 2 consecutive failures and resolve after 2 consecutive successes. The public page (server-rendered, no client JS, 60s meta refresh) shows a banner, per-component status with 90-day uptime bars, and open/past incidents; /status.json exposes the same snapshot, with a controlled 503 fallback if the Durable Object itself is unavailable.

3. Operator email alerts. Sent to me@kentcdodds.com from kody@heykody.dev via the Cloudflare Email REST API (same mechanism as existing ops alerts). Policy: one email when an outage episode starts, then silence except one reminder per 24h while unresolved, one "all is well" email once everything recovers, all under a daily cap (STATUS_ALERT_DAILY_LIMIT, default 6). Every send attempt counts toward the cap (a permanently failing token cannot retry forever), and a cap-suppressed all-clear still sends on the first uncapped tick.

Deploy

New path-filtered deploy-status-worker job in .github/workflows/deploy.yml (mirrors the backup control plane job) that runs after the main worker deploy so probes never hit a production worker missing /health/components. It deploys with the production CLOUDFLARE_API_TOKEN, provisions the status.heykody.dev custom domain via wrangler routes, sets BUILD_COMMIT, syncs the token as a Worker secret for email sending, and healthchecks https://status.heykody.dev/health for the deployed SHA. npm run validate gains a status:build dry-run; the status package's *.node.test.ts run in the existing node-unit project.

Review notes

All CodeRabbit majors and Bugbot's deploy-ordering race are addressed (stricter /mcp and package-apps probe validation, degraded DO fallback, capped send attempts, coalesced health refresh, needs: deploy). Skipped: widening the status deploy path filter to package.json/lockfile (mirrors the backup-control-plane precedent; manual workflow_dispatch forces a status deploy) and the banner color-contrast nit.

Evidence

Local end-to-end test: main dev worker + status worker (probes pointed at the local main worker), forced an outage by stopping the main worker, ran probe ticks, then recovered:

  • 6 components opened incidents (external package_apps stayed operational), exactly one combined incident email decision was made, and on recovery all incidents resolved with exactly one all-clear decision (sends skipped locally without the token, but recorded — see status-alert-email-* log lines).
  • 30 unit tests across incidents state machine, email policy, probes, page rendering, and /health/components. Full npm run validate green locally (including the new deploy-guardrails check from main).

status_page_incident_and_recovery_demo.mp4

Status page during incident
Status page recovered

System recap — adds a new primitive (high risk)

Mode: recap · Base: main @ bde1cc1e · Head: 19c92607

Classification: adds — new status-page primitive (independently deployed worker + Durable Object); app-ui is extended with a new public route.

Primitives touched

Primitive Group Impact
status-page surfaces adds — packages/status/ worker, cron prober, StatusStore DO, alerts
app-ui surfaces extends — new public GET /health/components route + handler

System map

The status worker probes the app's public endpoints every minute and stores results in its own Durable Object; nothing in the main worker depends on it.

Legend: green = composes (wiring only) · amber = extended by this PR · red = new primitive · gray = context (unchanged, included only when an edge crosses it).

flowchart LR
	statusPage["status-page<br/>Public status page"]:::added
	appUi["app-ui<br/>Browser app (Remix 3)"]:::extended
	mcpServer["mcp-server<br/>MCP endpoint (/mcp)"]:::untouched
	statusPage -->|"GET /health + /health/components every minute"| appUi
	statusPage -->|"unauthenticated Bearer 401 challenge probe"| mcpServer
	statusPage -->|"operator alert email via Cloudflare Email REST API"| statusPage
	classDef touched fill:#1a7f37,color:#fff
	classDef extended fill:#9a6700,color:#fff
	classDef added fill:#cf222e,color:#fff
	classDef untouched fill:#57606a,color:#fff
Loading

Change flow

sequenceDiagram
	participant Cron as status worker cron (1/min)
	participant DO as StatusStore DO (SQLite)
	participant App as heykody.dev (public endpoints)
	participant Email as Cloudflare Email API
	Cron->>DO: runProbes()
	DO->>App: /health, /health/components, /mcp, kodyapps.dev
	DO->>DO: samples + daily rollups + incident transitions (2-fail open / 2-success resolve)
	DO->>Email: incident / daily reminder / all-clear (attempts capped at 6/day)
Loading

Invariants

Status data is global operational telemetry, not user data: the status worker holds no user state, no APP_DB access, and no per-user Durable Objects, so the per-user isolation invariant is not implicated. /health/components exposes only per-binding ok/latency, never data.

Open in Web Open in Cursor 

Summary by CodeRabbit

  • New Features

    • Added a public status page with overall health, component availability, uptime, response times, incidents, and automatic refresh.
    • Added machine-readable status and component health endpoints for monitoring.
    • Added automated probing to detect outages, track recovery, and maintain incident history.
    • Added email notifications for incidents, reminders, and service recovery with delivery limits.
    • Added independent scheduled monitoring with resilient status storage.
    • Health checks now report the status and latency of key application services.
  • Documentation

    • Added setup and architecture documentation for status monitoring and alerting.

@coderabbitai

coderabbitai Bot commented Aug 4, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

This change adds an independently deployed status worker. It probes application components, stores incidents and uptime data, sends rate-limited alerts, renders status endpoints, and integrates health checks, builds, deployment, and architecture documentation.

Changes

Status page system

Layer / File(s) Summary
Application health endpoint
packages/worker/src/app/handlers/health-components.ts, packages/worker/src/app/routes.ts, packages/worker/src/app/router.ts, packages/worker/src/app/handlers/*test.ts
Adds /health/components with parallel binding checks, timeout handling, cached reports, commit metadata, and 200/503 responses.
Probe and incident model
packages/status/status-types.ts, packages/status/probes.ts, packages/status/incidents.ts, packages/status/*test.ts, packages/status/package.json, packages/status/tsconfig.json, packages/status/vitest.config.ts
Defines status data types, endpoint probes, concurrent probe execution, and threshold-based incident transitions with coverage.
Durable Object storage and alerts
packages/status/status-store.ts, packages/status/email-policy.ts, packages/status/alert-email.ts, packages/status/*test.ts
Persists samples, uptime, incidents, and notifications in SQLite. Applies email policy limits and sends Cloudflare Email API alerts with retries.
Status worker routes and page
packages/status/worker.ts, packages/status/status-page.ts, packages/status/wrangler.jsonc, packages/status/readme.md, packages/status/*test.ts
Adds scheduled probing, /health, /, and /status.json routes, Durable Object configuration, and escaped server-rendered status pages.
Build, validation, and deployment integration
package.json, .github/workflows/validate.yml, .github/workflows/deploy.yml, docs/contributing/architecture/primitives.yaml, docs/contributing/decisions/*, docs/contributing/setup-manifest.md
Adds workspace scripts and validation. Deploys the status worker conditionally, syncs its secret, and verifies the deployed commit through its health endpoint. Adds architecture and setup records.

Estimated code review effort: 4 (Complex) | ~60 minutes

Sequence Diagram(s)

sequenceDiagram
  participant StatusWorker
  participant StatusStore
  participant MainWorker
  participant SQLite
  participant CloudflareEmailAPI
  StatusWorker->>StatusStore: Run scheduled probes
  StatusStore->>MainWorker: Request health endpoints
  MainWorker-->>StatusStore: Return component health
  StatusStore->>SQLite: Store samples and incidents
  StatusStore->>CloudflareEmailAPI: Send rate-limited alert
  StatusWorker-->>StatusStore: Serve status snapshot
Loading

Possibly related PRs

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 2.78% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly summarizes the public status worker, health endpoint, and operator email alerts.
Description check ✅ Passed The description covers intent, implementation summary, testing evidence, deployment, and system impact in sufficient detail.
✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch cursor/status-page-da8d

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@kody-bot
kody-bot marked this pull request as ready for review August 4, 2026 23:35

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 10

🧹 Nitpick comments (3)
packages/status/status-page.node.test.ts (1)

43-45: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value

Assert a literal escaped name instead of reapplying the escape rule.

Line 44 reapplies part of the escapeHtml logic to build the expectation. The assertion then tracks whatever the implementation does. If escapeHtml stopped escaping &, both sides would change together and this test would still pass.

Keep the loop for presence of each component, and add one literal assertion for the name that contains &.

The static analysis warning on this line is a false positive. This is a test expectation, not an output-encoding control.

💚 Proposed refactor
-	for (const component of statusComponents) {
-		expect(html).toContain(component.name.replaceAll('&', '&amp;'))
-	}
+	expect(html).toContain('App &amp; API')
+	expect(html).not.toContain('App & API')
+	for (const component of statusComponents) {
+		expect(html).toContain(`class="dot operational"`)
+		expect(html).toContain(component.id)
+	}

Adjust the per-component assertion to whatever identifier renderComponent emits for each component.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@packages/status/status-page.node.test.ts` around lines 43 - 45, Update the
assertions in the status component rendering test to compare each component
against the literal identifier emitted by renderComponent, without reapplying
escapeHtml or replaceAll. Preserve the loop verifying every component, and add a
separate literal assertion for the component name containing “&” to ensure the
expected escaped HTML is fixed independently of the implementation.

Source: Linters/SAST tools

packages/status/status-page.ts (1)

10-16: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Duplicated escapeHtml in two files of the same package. Both files define a byte-identical escaping helper. The shared root cause is a missing shared module, so any future correction to the escaping rule must be applied twice.

  • packages/status/status-page.ts#L10-L16: delete the local escapeHtml and import it from a new shared module, for example packages/status/html.ts.
  • packages/status/email-policy.ts#L64-L70: delete the local escapeHtml and import the same shared helper.

The escaping logic itself is correct in both files. Every dynamic value lands in element content or a double-quoted attribute, and " is escaped. The static analysis suggestion to adopt DOMPurify or sanitize-html does not apply, because those libraries sanitize markup and need a DOM, while this code encodes untrusted text in a Workers runtime.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@packages/status/status-page.ts` around lines 10 - 16, Extract the existing
byte-identical escapeHtml helper into a shared module, then delete the local
definitions and import the shared helper in
packages/status/status-page.ts#L10-L16 and
packages/status/email-policy.ts#L64-L70. Preserve the current escaping behavior
unchanged; both sites should use the same shared symbol.

Source: Linters/SAST tools

packages/status/email-policy.ts (1)

119-129: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value

Make the dropped footer paragraph explicit.

toHtml drops the last paragraph with .slice(0, -1). That works only because every caller appends footerText as the final paragraph. The coupling is not visible at this function boundary. If a future caller adds text after the footer, the HTML body silently loses a paragraph.

Pass the body and the footer separately.

♻️ Proposed refactor
-function toHtml(text: string, footerHtml: string): string {
-	const paragraphs = text
-		.split('\n\n')
-		.slice(0, -1)
+function toHtml(bodyText: string, footerHtml: string): string {
+	const paragraphs = bodyText
+		.split('\n\n')
 		.map(
 			(paragraph) =>
 				`<p>${escapeHtml(paragraph).replaceAll('\n', '<br />')}</p>`,
 		)
 		.join('')
 	return `<div>${paragraphs}${footerHtml}</div>`
 }

Then build each case from a body string and append footerText only for text:

 		case 'all_clear': {
 			const subject = '[kody status] All systems operational again'
-			const text = `Every kody component has recovered. All systems are operational as of ${new Date(input.now).toISOString()}.\n\n${footerText}`
-			return { subject, text, html: toHtml(text, footerHtml) }
+			const body = `Every kody component has recovered. All systems are operational as of ${new Date(input.now).toISOString()}.`
+			return {
+				subject,
+				text: `${body}\n\n${footerText}`,
+				html: toHtml(body, footerHtml),
+			}
 		}
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@packages/status/email-policy.ts` around lines 119 - 129, Refactor toHtml so
it receives the body text and footer separately instead of relying on
text.split('\n\n').slice(0, -1) to discard the final paragraph. Remove the
implicit last-paragraph drop, render all supplied body paragraphs, and append
footerHtml explicitly only where the text case requires it; update each caller
to pass the separated values.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In @.github/workflows/deploy.yml:
- Around line 100-107: Update the path filter in the status worker deployment
condition to include the root package.json and the npm lockfile consumed by npm
ci, alongside the existing deploy workflow and packages/status/ paths. Keep the
existing deploy_status_worker outputs and messages unchanged.

In `@packages/status/email-policy.node.test.ts`:
- Around line 75-77: Rename the test describing the default input in the “no
email” case so it reflects that the healthy state has not already been
announced. Keep the decideStatusEmail(input()) assertion unchanged; only update
the test name near that case.

In `@packages/status/probes.ts`:
- Around line 110-114: Update the OAuth validation in the probe’s ok calculation
to require a 401 response with a WWW-Authenticate header using the expected
OAuth scheme and required challenge parameters, rather than accepting any header
value. Add a test covering a 401 response with a Basic challenge and ensure it
is reported as unhealthy.
- Line 136: Update the health evaluation around result.response.status so only
2xx and 3xx responses are considered healthy, while 4xx and 5xx responses remain
unhealthy; preserve redirect validity and add a probe test covering a 404
response.

In `@packages/status/status-page.ts`:
- Around line 135-142: Update the latency rendering in the component card to
include the latency value only when the component is operational, while still
hiding it when latencyMs is null. Use component.status in the existing latency
expression so down components display only their status label and preserve the
current formatting for operational components.
- Around line 51-60: Update the .banner.unknown background to use a darker slate
color that provides WCAG AA contrast with the existing white text, while leaving
the --unknown color used by the status dot unchanged.

In `@packages/status/status-store.ts`:
- Around line 306-318: Update the send result contract used by sendAlertEmail
and the status-store notification flow to distinguish retryable failures from
permanent failures: mark network and 5xx errors retryable, while treating
permanent API rejections such as 403/422 as non-retryable. In the notification
recording logic around decideStatusEmail, continue skipping retryable failures
but insert permanent failures with delivered set to false so the daily cap
advances.

In `@packages/status/worker.ts`:
- Around line 22-36: Wrap the getSnapshot calls in both the root-page and
/status.json branches with error handling so Durable Object or storage failures
return a controlled degraded response instead of propagating a runtime error.
Preserve the existing successful renderStatusPage and Response.json behavior,
and apply the same fallback semantics to both routes.

In `@packages/status/wrangler.jsonc`:
- Line 41: Remove the personal value from the ALERT_EMAIL_TO entry in the
committed wrangler configuration and configure the recipient through the deploy
workflow as a Worker secret, following the existing CLOUDFLARE_API_TOKEN secret
pattern; alternatively, replace it with the approved role address
alerts@heykody.dev.

In `@packages/worker/src/app/handlers/health-components.ts`:
- Around line 128-130: Update the cache refresh logic around
collectHealthComponents to retain an in-flight Promise<HealthComponentsReport>
and have concurrent expired-cache requests await the same promise instead of
starting duplicate probes. Set expiresAt only after the shared report resolves,
clear the in-flight state when completion allows a later refresh, and add a test
covering concurrent requests during cache expiration.

---

Nitpick comments:
In `@packages/status/email-policy.ts`:
- Around line 119-129: Refactor toHtml so it receives the body text and footer
separately instead of relying on text.split('\n\n').slice(0, -1) to discard the
final paragraph. Remove the implicit last-paragraph drop, render all supplied
body paragraphs, and append footerHtml explicitly only where the text case
requires it; update each caller to pass the separated values.

In `@packages/status/status-page.node.test.ts`:
- Around line 43-45: Update the assertions in the status component rendering
test to compare each component against the literal identifier emitted by
renderComponent, without reapplying escapeHtml or replaceAll. Preserve the loop
verifying every component, and add a separate literal assertion for the
component name containing “&” to ensure the expected escaped HTML is fixed
independently of the implementation.

In `@packages/status/status-page.ts`:
- Around line 10-16: Extract the existing byte-identical escapeHtml helper into
a shared module, then delete the local definitions and import the shared helper
in packages/status/status-page.ts#L10-L16 and
packages/status/email-policy.ts#L64-L70. Preserve the current escaping behavior
unchanged; both sites should use the same shared symbol.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 66a94c64-2d4d-44d5-866f-4984aee61872

📥 Commits

Reviewing files that changed from the base of the PR and between bde1cc1 and e917ed1.

⛔ Files ignored due to path filters (1)
  • package-lock.json is excluded by !**/package-lock.json
📒 Files selected for processing (28)
  • .github/workflows/deploy.yml
  • .github/workflows/validate.yml
  • docs/contributing/architecture/primitives.yaml
  • docs/contributing/decisions/0004-status-page-separate-worker.md
  • docs/contributing/decisions/index.md
  • docs/contributing/setup-manifest.md
  • package.json
  • packages/status/alert-email.ts
  • packages/status/email-policy.node.test.ts
  • packages/status/email-policy.ts
  • packages/status/incidents.node.test.ts
  • packages/status/incidents.ts
  • packages/status/package.json
  • packages/status/probes.node.test.ts
  • packages/status/probes.ts
  • packages/status/readme.md
  • packages/status/status-page.node.test.ts
  • packages/status/status-page.ts
  • packages/status/status-store.ts
  • packages/status/status-types.ts
  • packages/status/tsconfig.json
  • packages/status/vitest.config.ts
  • packages/status/worker.ts
  • packages/status/wrangler.jsonc
  • packages/worker/src/app/handlers/health-components-handler.node.test.ts
  • packages/worker/src/app/handlers/health-components.ts
  • packages/worker/src/app/router.ts
  • packages/worker/src/app/routes.ts

Comment on lines +100 to +107
if git diff --name-only "${DEPLOY_SHA}^" "${DEPLOY_SHA}" | grep -E \
'^(\.github/workflows/deploy\.yml|packages/status/)'; then
echo "deploy_status_worker=true" >> "$GITHUB_OUTPUT"
echo "Status worker deploy: path filter matched."
else
echo "deploy_status_worker=false" >> "$GITHUB_OUTPUT"
echo "Status worker deploy: skipped (no relevant path changes)."
fi

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

Deploy when the root status deployment input changes.

Line 101 does not match root package.json. A change to package.json Line 39 changes the command executed at Line 419, but the workflow skips the status Worker deployment. Match package.json and the npm lockfile used by npm ci.

Proposed fix
-            '^(\.github/workflows/deploy\.yml|packages/status/)'; then
+            '^(\.github/workflows/deploy\.yml|package(-lock)?\.json|npm-shrinkwrap\.json|packages/status/)'; then
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
if git diff --name-only "${DEPLOY_SHA}^" "${DEPLOY_SHA}" | grep -E \
'^(\.github/workflows/deploy\.yml|packages/status/)'; then
echo "deploy_status_worker=true" >> "$GITHUB_OUTPUT"
echo "Status worker deploy: path filter matched."
else
echo "deploy_status_worker=false" >> "$GITHUB_OUTPUT"
echo "Status worker deploy: skipped (no relevant path changes)."
fi
if git diff --name-only "${DEPLOY_SHA}^" "${DEPLOY_SHA}" | grep -E \
'^(\.github/workflows/deploy\.yml|package(-lock)?\.json|npm-shrinkwrap\.json|packages/status/)'; then
echo "deploy_status_worker=true" >> "$GITHUB_OUTPUT"
echo "Status worker deploy: path filter matched."
else
echo "deploy_status_worker=false" >> "$GITHUB_OUTPUT"
echo "Status worker deploy: skipped (no relevant path changes)."
fi
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In @.github/workflows/deploy.yml around lines 100 - 107, Update the path filter
in the status worker deployment condition to include the root package.json and
the npm lockfile consumed by npm ci, alongside the existing deploy workflow and
packages/status/ paths. Keep the existing deploy_status_worker outputs and
messages unchanged.

Comment thread packages/status/email-policy.node.test.ts Outdated
Comment thread packages/status/probes.ts Outdated
Comment on lines +110 to +114
// An unauthenticated GET must produce the OAuth challenge; anything else
// (especially a 5xx) means the MCP surface is broken.
const ok =
result.response.status === 401 &&
result.response.headers.has('WWW-Authenticate')

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

Validate the required OAuth challenge.

A 401 response with WWW-Authenticate: Basic passes this probe. The MCP endpoint can then stop presenting its required OAuth challenge while the status worker reports it as operational and does not open an incident.

Validate the expected OAuth scheme and required challenge parameters. Add a test for a 401 Basic response.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@packages/status/probes.ts` around lines 110 - 114, Update the OAuth
validation in the probe’s ok calculation to require a 401 response with a
WWW-Authenticate header using the expected OAuth scheme and required challenge
parameters, rather than accepting any header value. Add a test covering a 401
response with a Basic challenge and ensure it is reported as unhealthy.

Comment thread packages/status/probes.ts Outdated
detail: result.error,
}
}
const ok = result.response.status < 500

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

Do not treat 4xx responses as healthy.

result.response.status < 500 marks 401, 403, 404, and 429 as operational. A public package-app endpoint can be inaccessible or unrouted while the status page remains green.

Keep redirects valid, but require a 2xx or 3xx response. Add a 404 probe test.

Proposed fix
-	const ok = result.response.status < 500
+	const ok =
+		result.response.status >= 200 && result.response.status < 400
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
const ok = result.response.status < 500
const ok =
result.response.status >= 200 && result.response.status < 400
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@packages/status/probes.ts` at line 136, Update the health evaluation around
result.response.status so only 2xx and 3xx responses are considered healthy,
while 4xx and 5xx responses remain unhealthy; preserve redirect validity and add
a probe test covering a 404 response.

Comment on lines +51 to +60
.banner {
margin: 1.25rem 0;
padding: 1rem 1.25rem;
border-radius: 0.75rem;
font-weight: 600;
color: #ffffff;
}
.banner.ok { background: var(--ok); }
.banner.down { background: var(--down); }
.banner.unknown { background: var(--unknown); }

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# Description: Compute WCAG contrast ratios for white text on each banner background.
set -uo pipefail

python3 - <<'PY'
def lum(hex_color):
    r, g, b = (int(hex_color[i:i+2], 16) / 255 for i in (1, 3, 5))
    def chan(c):
        return c / 12.92 if c <= 0.03928 else ((c + 0.055) / 1.055) ** 2.4
    r, g, b = chan(r), chan(g), chan(b)
    return 0.2126 * r + 0.7152 * g + 0.0722 * b

def ratio(a, b):
    la, lb = lum(a), lum(b)
    hi, lo = max(la, lb), min(la, lb)
    return (hi + 0.05) / (lo + 0.05)

for name, bg in [
    ('ok', '`#16a34a`'),
    ('down', '`#dc2626`'),
    ('unknown', '`#94a3b8`'),
    ('unknown-strong', '`#475569`'),
]:
    print(f'{name:16} {bg} vs `#ffffff` -> {ratio(bg, "`#ffffff`"):.2f}:1')
print('WCAG AA needs 4.5:1 for normal text, 3:1 for large text.')
PY

Repository: kentcdodds/kody

Length of output: 394


🏁 Script executed:

#!/bin/bash
set -euo pipefail

echo "== locate file =="
fd -a 'status-page\.ts$' . || true

echo "== relevant section =="
if [ -f packages/status/status-page.ts ]; then
  nl -ba packages/status/status-page.ts | sed -n '1,120p'
fi

echo "== status variable usages =="
rg -n "unknown:|--unknown|banner\.unknown|Status data is not available" packages/status packages -S || true

Repository: kentcdodds/kody

Length of output: 284


🏁 Script executed:

#!/bin/bash
set -euo pipefail

echo "== relevant section =="
awk '{printf "%5d\t%s\n", NR, $0}' packages/status/status-page.ts | sed -n '1,120p'

echo "== status variable usages =="
rg -n "unknown:|--unknown|banner\.unknown|Status data is not available" packages/status packages -S || true

echo "== contrast calculation =="
python3 - <<'PY'
from pathlib import Path
text = Path('packages/status/status-page.ts').read_bytes().decode()
for token in ('--ok', '--down', '--unknown', '--partial'):
    idx = text.find(token + ': ')
    if idx >= 0:
        line = text[idx:text.find('\n', idx)].strip()
        print(line)
PY

Repository: kentcdodds/kody

Length of output: 6291


Use a darker color for the unknown banner background.

The unknown status banner uses white text, but --unknown is #94a3b8 and fails WCAG AA contrast for normal text. Keep the lighter color for the status dot, and use a stronger slate shade behind the banner text.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@packages/status/status-page.ts` around lines 51 - 60, Update the
.banner.unknown background to use a darker slate color that provides WCAG AA
contrast with the existing white text, while leaving the --unknown color used by
the status dot unchanged.

Comment thread packages/status/status-page.ts
Comment thread packages/status/status-store.ts Outdated
Comment on lines +306 to +318
// Transient send errors stay unrecorded so the next tick retries;
// delivered and unconfigured-skip sends are recorded so the policy
// (pause + cap) advances.
if (!result.delivered && !result.skipped) return
this.ctx.storage.sql.exec(
`INSERT INTO notifications (kind, sent_at, day, subject, delivered)
VALUES (?, ?, ?, ?, ?)`,
decision.kind,
now,
day,
content.subject,
result.delivered ? 1 : 0,
)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🩺 Stability & Availability | 🟠 Major | 🏗️ Heavy lift

Record failed non-transient sends, or the worker resends every minute.

The code records a notification only when result.delivered or result.skipped is true. sendAlertEmail returns { delivered: false } without skipped for every API rejection, including permanent ones such as HTTP 403 from an invalid token or HTTP 422 from an unverified sender address.

For a permanent failure the sequence repeats each cron tick:

  1. lastNotifiedState stays ok, so decideStatusEmail returns incident_opened again.
  2. sendAlertEmail posts again and fails again.
  3. Nothing is written to notifications, so emailsSentToday stays 0.

The daily cap counts only recorded notifications, so it never engages. The worker then posts one alert request per minute for the whole outage. Distinguish retryable failures from permanent ones, and record permanent failures so the cap advances.

♻️ One approach: return a retryable flag and record permanent failures

In packages/status/alert-email.ts, mark 5xx and network errors as retryable:

 export type AlertEmailResult = {
 	delivered: boolean
 	skipped?: boolean
+	retryable?: boolean
 	error?: string
 }
 	} catch (error) {
 		const detail = error instanceof Error ? error.message : String(error)
 		console.warn('status-alert-email-request-failed', detail)
-		return { delivered: false, error: detail }
+		return { delivered: false, retryable: true, error: detail }
 	}
 		console.warn(
 			'status-alert-email-failed',
 			JSON.stringify({ status: response.status, detail }),
 		)
-		return { delivered: false, error: detail }
+		return { delivered: false, retryable: response.status >= 500, error: detail }
 	}

Then in packages/status/status-store.ts:

-		// Transient send errors stay unrecorded so the next tick retries;
-		// delivered and unconfigured-skip sends are recorded so the policy
-		// (pause + cap) advances.
-		if (!result.delivered && !result.skipped) return
+		// Retryable send errors stay unrecorded so the next tick retries.
+		// Delivered, skipped, and permanently failed sends are recorded so the
+		// policy (pause + cap) advances instead of resending every minute.
+		if (result.retryable) return
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@packages/status/status-store.ts` around lines 306 - 318, Update the send
result contract used by sendAlertEmail and the status-store notification flow to
distinguish retryable failures from permanent failures: mark network and 5xx
errors retryable, while treating permanent API rejections such as 403/422 as
non-retryable. In the notification recording logic around decideStatusEmail,
continue skipping retryable failures but insert permanent failures with
delivered set to false so the daily cap advances.

Comment thread packages/status/worker.ts Outdated
"PRIMARY_ORIGIN": "https://heykody.dev",
"PACKAGE_APP_ORIGIN": "https://kodyapps.dev",
"STATUS_PAGE_URL": "https://status.heykody.dev",
"ALERT_EMAIL_TO": "me@kentcdodds.com",

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔒 Security & Privacy | 🟡 Minor | ⚡ Quick win

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# Description: Check repository visibility and how status worker secrets are synced.
set -uo pipefail

gh repo view --json visibility,isPrivate 2>/dev/null || echo 'gh repo view unavailable'

fd -t f 'deploy.yml' .github
rg -n -C5 'ALERT_EMAIL|status.*secret|put secret|secret:bulk' \
  --glob '.github/**' --glob 'package.json'

Repository: kentcdodds/kody

Length of output: 224


🏁 Script executed:

#!/bin/bash
set -uo pipefail

printf '--- wrangler.jsonc ---\n'
cat -n packages/status/wrangler.jsonc | sed -n '30,55p'

printf '\n--- search ALTE_EMAIL_TO ---\n'
rg -n -C3 'ALERT_EMAIL_TO|sendAlertEmail|mailto|AlertEmail' . --glob '!node_modules' --glob '!.git'

printf '\n--- deploy workflow relevant secrets ---\n'
cat -n .github/workflows/deploy.yml | sed -n '1,240p'

printf '\n--- all workflow secret usage ---\n'
rg -n 'github\.secrets\.|secrets:|CLOUDFLARE_API_TOKEN|wrangler secrets|put\.secret|secret\.bulk|environments:' .github/workflows package.json --glob '!node_modules' --glob '!.git' || true

Repository: kentcdodds/kody

Length of output: 26472


Sensitive Data Exposure (CWE-359)

Reachability: External

Move the alert recipient out of committed config.

ALERT_EMAIL_TO stores a personal email address in a public repository. Wrangler vars are plaintext, so this value is readable from the repo and remains in git history.

Set the recipient as a Worker secret in the deploy workflow, or use a role address such as alerts@heykody.dev. Use the same secret pattern already documented for CLOUDFLARE_API_TOKEN.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@packages/status/wrangler.jsonc` at line 41, Remove the personal value from
the ALERT_EMAIL_TO entry in the committed wrangler configuration and configure
the recipient through the deploy workflow as a Worker secret, following the
existing CLOUDFLARE_API_TOKEN secret pattern; alternatively, replace it with the
approved role address alerts@heykody.dev.

Comment thread packages/worker/src/app/handlers/health-components.ts Outdated

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit db2a7a6. Configure here.

Comment thread .github/workflows/deploy.yml
@github-actions

github-actions Bot commented Aug 4, 2026 •

Copy link
Copy Markdown
Contributor

🔎 Preview deployed: https://kody-pr-1230.kody-a99.workers.dev

Worker: kody-pr-1230
D1: kody-pr-1230-db
KV: kody-pr-1230-oauth-kv

Mocks:

…end attempts, coalesced health refresh, deploy ordering
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants