Add a public status page: status.heykody.dev worker + /health/components + operator email alerts - #1230
Conversation
…, and alert email policy
📝 WalkthroughWalkthroughThis change adds an independently deployed status worker. It probes application components, stores incidents and uptime data, sends rate-limited alerts, renders status endpoints, and integrates health checks, builds, deployment, and architecture documentation. ChangesStatus page system
Estimated code review effort: 4 (Complex) | ~60 minutes Sequence Diagram(s)sequenceDiagram
participant StatusWorker
participant StatusStore
participant MainWorker
participant SQLite
participant CloudflareEmailAPI
StatusWorker->>StatusStore: Run scheduled probes
StatusStore->>MainWorker: Request health endpoints
MainWorker-->>StatusStore: Return component health
StatusStore->>SQLite: Store samples and incidents
StatusStore->>CloudflareEmailAPI: Send rate-limited alert
StatusWorker-->>StatusStore: Serve status snapshot
Possibly related PRs
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches 💡 1📝 Generate docstrings 💡
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Actionable comments posted: 10
🧹 Nitpick comments (3)
packages/status/status-page.node.test.ts (1)
43-45: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low valueAssert a literal escaped name instead of reapplying the escape rule.
Line 44 reapplies part of the
escapeHtmllogic to build the expectation. The assertion then tracks whatever the implementation does. IfescapeHtmlstopped escaping&, both sides would change together and this test would still pass.Keep the loop for presence of each component, and add one literal assertion for the name that contains
&.The static analysis warning on this line is a false positive. This is a test expectation, not an output-encoding control.
💚 Proposed refactor
- for (const component of statusComponents) { - expect(html).toContain(component.name.replaceAll('&', '&')) - } + expect(html).toContain('App & API') + expect(html).not.toContain('App & API') + for (const component of statusComponents) { + expect(html).toContain(`class="dot operational"`) + expect(html).toContain(component.id) + }Adjust the per-component assertion to whatever identifier
renderComponentemits for each component.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@packages/status/status-page.node.test.ts` around lines 43 - 45, Update the assertions in the status component rendering test to compare each component against the literal identifier emitted by renderComponent, without reapplying escapeHtml or replaceAll. Preserve the loop verifying every component, and add a separate literal assertion for the component name containing “&” to ensure the expected escaped HTML is fixed independently of the implementation.Source: Linters/SAST tools
packages/status/status-page.ts (1)
10-16: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick winDuplicated
escapeHtmlin two files of the same package. Both files define a byte-identical escaping helper. The shared root cause is a missing shared module, so any future correction to the escaping rule must be applied twice.
packages/status/status-page.ts#L10-L16: delete the localescapeHtmland import it from a new shared module, for examplepackages/status/html.ts.packages/status/email-policy.ts#L64-L70: delete the localescapeHtmland import the same shared helper.The escaping logic itself is correct in both files. Every dynamic value lands in element content or a double-quoted attribute, and
"is escaped. The static analysis suggestion to adopt DOMPurify or sanitize-html does not apply, because those libraries sanitize markup and need a DOM, while this code encodes untrusted text in a Workers runtime.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@packages/status/status-page.ts` around lines 10 - 16, Extract the existing byte-identical escapeHtml helper into a shared module, then delete the local definitions and import the shared helper in packages/status/status-page.ts#L10-L16 and packages/status/email-policy.ts#L64-L70. Preserve the current escaping behavior unchanged; both sites should use the same shared symbol.Source: Linters/SAST tools
packages/status/email-policy.ts (1)
119-129: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low valueMake the dropped footer paragraph explicit.
toHtmldrops the last paragraph with.slice(0, -1). That works only because every caller appendsfooterTextas the final paragraph. The coupling is not visible at this function boundary. If a future caller adds text after the footer, the HTML body silently loses a paragraph.Pass the body and the footer separately.
♻️ Proposed refactor
-function toHtml(text: string, footerHtml: string): string { - const paragraphs = text - .split('\n\n') - .slice(0, -1) +function toHtml(bodyText: string, footerHtml: string): string { + const paragraphs = bodyText + .split('\n\n') .map( (paragraph) => `<p>${escapeHtml(paragraph).replaceAll('\n', '<br />')}</p>`, ) .join('') return `<div>${paragraphs}${footerHtml}</div>` }Then build each case from a body string and append
footerTextonly fortext:case 'all_clear': { const subject = '[kody status] All systems operational again' - const text = `Every kody component has recovered. All systems are operational as of ${new Date(input.now).toISOString()}.\n\n${footerText}` - return { subject, text, html: toHtml(text, footerHtml) } + const body = `Every kody component has recovered. All systems are operational as of ${new Date(input.now).toISOString()}.` + return { + subject, + text: `${body}\n\n${footerText}`, + html: toHtml(body, footerHtml), + } }🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@packages/status/email-policy.ts` around lines 119 - 129, Refactor toHtml so it receives the body text and footer separately instead of relying on text.split('\n\n').slice(0, -1) to discard the final paragraph. Remove the implicit last-paragraph drop, render all supplied body paragraphs, and append footerHtml explicitly only where the text case requires it; update each caller to pass the separated values.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In @.github/workflows/deploy.yml:
- Around line 100-107: Update the path filter in the status worker deployment
condition to include the root package.json and the npm lockfile consumed by npm
ci, alongside the existing deploy workflow and packages/status/ paths. Keep the
existing deploy_status_worker outputs and messages unchanged.
In `@packages/status/email-policy.node.test.ts`:
- Around line 75-77: Rename the test describing the default input in the “no
email” case so it reflects that the healthy state has not already been
announced. Keep the decideStatusEmail(input()) assertion unchanged; only update
the test name near that case.
In `@packages/status/probes.ts`:
- Around line 110-114: Update the OAuth validation in the probe’s ok calculation
to require a 401 response with a WWW-Authenticate header using the expected
OAuth scheme and required challenge parameters, rather than accepting any header
value. Add a test covering a 401 response with a Basic challenge and ensure it
is reported as unhealthy.
- Line 136: Update the health evaluation around result.response.status so only
2xx and 3xx responses are considered healthy, while 4xx and 5xx responses remain
unhealthy; preserve redirect validity and add a probe test covering a 404
response.
In `@packages/status/status-page.ts`:
- Around line 135-142: Update the latency rendering in the component card to
include the latency value only when the component is operational, while still
hiding it when latencyMs is null. Use component.status in the existing latency
expression so down components display only their status label and preserve the
current formatting for operational components.
- Around line 51-60: Update the .banner.unknown background to use a darker slate
color that provides WCAG AA contrast with the existing white text, while leaving
the --unknown color used by the status dot unchanged.
In `@packages/status/status-store.ts`:
- Around line 306-318: Update the send result contract used by sendAlertEmail
and the status-store notification flow to distinguish retryable failures from
permanent failures: mark network and 5xx errors retryable, while treating
permanent API rejections such as 403/422 as non-retryable. In the notification
recording logic around decideStatusEmail, continue skipping retryable failures
but insert permanent failures with delivered set to false so the daily cap
advances.
In `@packages/status/worker.ts`:
- Around line 22-36: Wrap the getSnapshot calls in both the root-page and
/status.json branches with error handling so Durable Object or storage failures
return a controlled degraded response instead of propagating a runtime error.
Preserve the existing successful renderStatusPage and Response.json behavior,
and apply the same fallback semantics to both routes.
In `@packages/status/wrangler.jsonc`:
- Line 41: Remove the personal value from the ALERT_EMAIL_TO entry in the
committed wrangler configuration and configure the recipient through the deploy
workflow as a Worker secret, following the existing CLOUDFLARE_API_TOKEN secret
pattern; alternatively, replace it with the approved role address
alerts@heykody.dev.
In `@packages/worker/src/app/handlers/health-components.ts`:
- Around line 128-130: Update the cache refresh logic around
collectHealthComponents to retain an in-flight Promise<HealthComponentsReport>
and have concurrent expired-cache requests await the same promise instead of
starting duplicate probes. Set expiresAt only after the shared report resolves,
clear the in-flight state when completion allows a later refresh, and add a test
covering concurrent requests during cache expiration.
---
Nitpick comments:
In `@packages/status/email-policy.ts`:
- Around line 119-129: Refactor toHtml so it receives the body text and footer
separately instead of relying on text.split('\n\n').slice(0, -1) to discard the
final paragraph. Remove the implicit last-paragraph drop, render all supplied
body paragraphs, and append footerHtml explicitly only where the text case
requires it; update each caller to pass the separated values.
In `@packages/status/status-page.node.test.ts`:
- Around line 43-45: Update the assertions in the status component rendering
test to compare each component against the literal identifier emitted by
renderComponent, without reapplying escapeHtml or replaceAll. Preserve the loop
verifying every component, and add a separate literal assertion for the
component name containing “&” to ensure the expected escaped HTML is fixed
independently of the implementation.
In `@packages/status/status-page.ts`:
- Around line 10-16: Extract the existing byte-identical escapeHtml helper into
a shared module, then delete the local definitions and import the shared helper
in packages/status/status-page.ts#L10-L16 and
packages/status/email-policy.ts#L64-L70. Preserve the current escaping behavior
unchanged; both sites should use the same shared symbol.
🪄 Autofix
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: defaults
Review profile: CHILL
Plan: Pro Plus
Run ID: 66a94c64-2d4d-44d5-866f-4984aee61872
⛔ Files ignored due to path filters (1)
package-lock.jsonis excluded by!**/package-lock.json
📒 Files selected for processing (28)
.github/workflows/deploy.yml.github/workflows/validate.ymldocs/contributing/architecture/primitives.yamldocs/contributing/decisions/0004-status-page-separate-worker.mddocs/contributing/decisions/index.mddocs/contributing/setup-manifest.mdpackage.jsonpackages/status/alert-email.tspackages/status/email-policy.node.test.tspackages/status/email-policy.tspackages/status/incidents.node.test.tspackages/status/incidents.tspackages/status/package.jsonpackages/status/probes.node.test.tspackages/status/probes.tspackages/status/readme.mdpackages/status/status-page.node.test.tspackages/status/status-page.tspackages/status/status-store.tspackages/status/status-types.tspackages/status/tsconfig.jsonpackages/status/vitest.config.tspackages/status/worker.tspackages/status/wrangler.jsoncpackages/worker/src/app/handlers/health-components-handler.node.test.tspackages/worker/src/app/handlers/health-components.tspackages/worker/src/app/router.tspackages/worker/src/app/routes.ts
| if git diff --name-only "${DEPLOY_SHA}^" "${DEPLOY_SHA}" | grep -E \ | ||
| '^(\.github/workflows/deploy\.yml|packages/status/)'; then | ||
| echo "deploy_status_worker=true" >> "$GITHUB_OUTPUT" | ||
| echo "Status worker deploy: path filter matched." | ||
| else | ||
| echo "deploy_status_worker=false" >> "$GITHUB_OUTPUT" | ||
| echo "Status worker deploy: skipped (no relevant path changes)." | ||
| fi |
There was a problem hiding this comment.
🎯 Functional Correctness | 🟠 Major | ⚡ Quick win
Deploy when the root status deployment input changes.
Line 101 does not match root package.json. A change to package.json Line 39 changes the command executed at Line 419, but the workflow skips the status Worker deployment. Match package.json and the npm lockfile used by npm ci.
Proposed fix
- '^(\.github/workflows/deploy\.yml|packages/status/)'; then
+ '^(\.github/workflows/deploy\.yml|package(-lock)?\.json|npm-shrinkwrap\.json|packages/status/)'; then📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| if git diff --name-only "${DEPLOY_SHA}^" "${DEPLOY_SHA}" | grep -E \ | |
| '^(\.github/workflows/deploy\.yml|packages/status/)'; then | |
| echo "deploy_status_worker=true" >> "$GITHUB_OUTPUT" | |
| echo "Status worker deploy: path filter matched." | |
| else | |
| echo "deploy_status_worker=false" >> "$GITHUB_OUTPUT" | |
| echo "Status worker deploy: skipped (no relevant path changes)." | |
| fi | |
| if git diff --name-only "${DEPLOY_SHA}^" "${DEPLOY_SHA}" | grep -E \ | |
| '^(\.github/workflows/deploy\.yml|package(-lock)?\.json|npm-shrinkwrap\.json|packages/status/)'; then | |
| echo "deploy_status_worker=true" >> "$GITHUB_OUTPUT" | |
| echo "Status worker deploy: path filter matched." | |
| else | |
| echo "deploy_status_worker=false" >> "$GITHUB_OUTPUT" | |
| echo "Status worker deploy: skipped (no relevant path changes)." | |
| fi |
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In @.github/workflows/deploy.yml around lines 100 - 107, Update the path filter
in the status worker deployment condition to include the root package.json and
the npm lockfile consumed by npm ci, alongside the existing deploy workflow and
packages/status/ paths. Keep the existing deploy_status_worker outputs and
messages unchanged.
| // An unauthenticated GET must produce the OAuth challenge; anything else | ||
| // (especially a 5xx) means the MCP surface is broken. | ||
| const ok = | ||
| result.response.status === 401 && | ||
| result.response.headers.has('WWW-Authenticate') |
There was a problem hiding this comment.
🎯 Functional Correctness | 🟠 Major | ⚡ Quick win
Validate the required OAuth challenge.
A 401 response with WWW-Authenticate: Basic passes this probe. The MCP endpoint can then stop presenting its required OAuth challenge while the status worker reports it as operational and does not open an incident.
Validate the expected OAuth scheme and required challenge parameters. Add a test for a 401 Basic response.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@packages/status/probes.ts` around lines 110 - 114, Update the OAuth
validation in the probe’s ok calculation to require a 401 response with a
WWW-Authenticate header using the expected OAuth scheme and required challenge
parameters, rather than accepting any header value. Add a test covering a 401
response with a Basic challenge and ensure it is reported as unhealthy.
| detail: result.error, | ||
| } | ||
| } | ||
| const ok = result.response.status < 500 |
There was a problem hiding this comment.
🎯 Functional Correctness | 🟠 Major | ⚡ Quick win
Do not treat 4xx responses as healthy.
result.response.status < 500 marks 401, 403, 404, and 429 as operational. A public package-app endpoint can be inaccessible or unrouted while the status page remains green.
Keep redirects valid, but require a 2xx or 3xx response. Add a 404 probe test.
Proposed fix
- const ok = result.response.status < 500
+ const ok =
+ result.response.status >= 200 && result.response.status < 400📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| const ok = result.response.status < 500 | |
| const ok = | |
| result.response.status >= 200 && result.response.status < 400 |
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@packages/status/probes.ts` at line 136, Update the health evaluation around
result.response.status so only 2xx and 3xx responses are considered healthy,
while 4xx and 5xx responses remain unhealthy; preserve redirect validity and add
a probe test covering a 404 response.
| .banner { | ||
| margin: 1.25rem 0; | ||
| padding: 1rem 1.25rem; | ||
| border-radius: 0.75rem; | ||
| font-weight: 600; | ||
| color: #ffffff; | ||
| } | ||
| .banner.ok { background: var(--ok); } | ||
| .banner.down { background: var(--down); } | ||
| .banner.unknown { background: var(--unknown); } |
There was a problem hiding this comment.
📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win
🧩 Analysis chain
🏁 Script executed:
#!/bin/bash
# Description: Compute WCAG contrast ratios for white text on each banner background.
set -uo pipefail
python3 - <<'PY'
def lum(hex_color):
r, g, b = (int(hex_color[i:i+2], 16) / 255 for i in (1, 3, 5))
def chan(c):
return c / 12.92 if c <= 0.03928 else ((c + 0.055) / 1.055) ** 2.4
r, g, b = chan(r), chan(g), chan(b)
return 0.2126 * r + 0.7152 * g + 0.0722 * b
def ratio(a, b):
la, lb = lum(a), lum(b)
hi, lo = max(la, lb), min(la, lb)
return (hi + 0.05) / (lo + 0.05)
for name, bg in [
('ok', '`#16a34a`'),
('down', '`#dc2626`'),
('unknown', '`#94a3b8`'),
('unknown-strong', '`#475569`'),
]:
print(f'{name:16} {bg} vs `#ffffff` -> {ratio(bg, "`#ffffff`"):.2f}:1')
print('WCAG AA needs 4.5:1 for normal text, 3:1 for large text.')
PYRepository: kentcdodds/kody
Length of output: 394
🏁 Script executed:
#!/bin/bash
set -euo pipefail
echo "== locate file =="
fd -a 'status-page\.ts$' . || true
echo "== relevant section =="
if [ -f packages/status/status-page.ts ]; then
nl -ba packages/status/status-page.ts | sed -n '1,120p'
fi
echo "== status variable usages =="
rg -n "unknown:|--unknown|banner\.unknown|Status data is not available" packages/status packages -S || trueRepository: kentcdodds/kody
Length of output: 284
🏁 Script executed:
#!/bin/bash
set -euo pipefail
echo "== relevant section =="
awk '{printf "%5d\t%s\n", NR, $0}' packages/status/status-page.ts | sed -n '1,120p'
echo "== status variable usages =="
rg -n "unknown:|--unknown|banner\.unknown|Status data is not available" packages/status packages -S || true
echo "== contrast calculation =="
python3 - <<'PY'
from pathlib import Path
text = Path('packages/status/status-page.ts').read_bytes().decode()
for token in ('--ok', '--down', '--unknown', '--partial'):
idx = text.find(token + ': ')
if idx >= 0:
line = text[idx:text.find('\n', idx)].strip()
print(line)
PYRepository: kentcdodds/kody
Length of output: 6291
Use a darker color for the unknown banner background.
The unknown status banner uses white text, but --unknown is #94a3b8 and fails WCAG AA contrast for normal text. Keep the lighter color for the status dot, and use a stronger slate shade behind the banner text.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@packages/status/status-page.ts` around lines 51 - 60, Update the
.banner.unknown background to use a darker slate color that provides WCAG AA
contrast with the existing white text, while leaving the --unknown color used by
the status dot unchanged.
| // Transient send errors stay unrecorded so the next tick retries; | ||
| // delivered and unconfigured-skip sends are recorded so the policy | ||
| // (pause + cap) advances. | ||
| if (!result.delivered && !result.skipped) return | ||
| this.ctx.storage.sql.exec( | ||
| `INSERT INTO notifications (kind, sent_at, day, subject, delivered) | ||
| VALUES (?, ?, ?, ?, ?)`, | ||
| decision.kind, | ||
| now, | ||
| day, | ||
| content.subject, | ||
| result.delivered ? 1 : 0, | ||
| ) |
There was a problem hiding this comment.
🩺 Stability & Availability | 🟠 Major | 🏗️ Heavy lift
Record failed non-transient sends, or the worker resends every minute.
The code records a notification only when result.delivered or result.skipped is true. sendAlertEmail returns { delivered: false } without skipped for every API rejection, including permanent ones such as HTTP 403 from an invalid token or HTTP 422 from an unverified sender address.
For a permanent failure the sequence repeats each cron tick:
lastNotifiedStatestaysok, sodecideStatusEmailreturnsincident_openedagain.sendAlertEmailposts again and fails again.- Nothing is written to
notifications, soemailsSentTodaystays 0.
The daily cap counts only recorded notifications, so it never engages. The worker then posts one alert request per minute for the whole outage. Distinguish retryable failures from permanent ones, and record permanent failures so the cap advances.
♻️ One approach: return a retryable flag and record permanent failures
In packages/status/alert-email.ts, mark 5xx and network errors as retryable:
export type AlertEmailResult = {
delivered: boolean
skipped?: boolean
+ retryable?: boolean
error?: string
} } catch (error) {
const detail = error instanceof Error ? error.message : String(error)
console.warn('status-alert-email-request-failed', detail)
- return { delivered: false, error: detail }
+ return { delivered: false, retryable: true, error: detail }
} console.warn(
'status-alert-email-failed',
JSON.stringify({ status: response.status, detail }),
)
- return { delivered: false, error: detail }
+ return { delivered: false, retryable: response.status >= 500, error: detail }
}Then in packages/status/status-store.ts:
- // Transient send errors stay unrecorded so the next tick retries;
- // delivered and unconfigured-skip sends are recorded so the policy
- // (pause + cap) advances.
- if (!result.delivered && !result.skipped) return
+ // Retryable send errors stay unrecorded so the next tick retries.
+ // Delivered, skipped, and permanently failed sends are recorded so the
+ // policy (pause + cap) advances instead of resending every minute.
+ if (result.retryable) return🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@packages/status/status-store.ts` around lines 306 - 318, Update the send
result contract used by sendAlertEmail and the status-store notification flow to
distinguish retryable failures from permanent failures: mark network and 5xx
errors retryable, while treating permanent API rejections such as 403/422 as
non-retryable. In the notification recording logic around decideStatusEmail,
continue skipping retryable failures but insert permanent failures with
delivered set to false so the daily cap advances.
| "PRIMARY_ORIGIN": "https://heykody.dev", | ||
| "PACKAGE_APP_ORIGIN": "https://kodyapps.dev", | ||
| "STATUS_PAGE_URL": "https://status.heykody.dev", | ||
| "ALERT_EMAIL_TO": "me@kentcdodds.com", |
There was a problem hiding this comment.
🔒 Security & Privacy | 🟡 Minor | ⚡ Quick win
🧩 Analysis chain
🏁 Script executed:
#!/bin/bash
# Description: Check repository visibility and how status worker secrets are synced.
set -uo pipefail
gh repo view --json visibility,isPrivate 2>/dev/null || echo 'gh repo view unavailable'
fd -t f 'deploy.yml' .github
rg -n -C5 'ALERT_EMAIL|status.*secret|put secret|secret:bulk' \
--glob '.github/**' --glob 'package.json'Repository: kentcdodds/kody
Length of output: 224
🏁 Script executed:
#!/bin/bash
set -uo pipefail
printf '--- wrangler.jsonc ---\n'
cat -n packages/status/wrangler.jsonc | sed -n '30,55p'
printf '\n--- search ALTE_EMAIL_TO ---\n'
rg -n -C3 'ALERT_EMAIL_TO|sendAlertEmail|mailto|AlertEmail' . --glob '!node_modules' --glob '!.git'
printf '\n--- deploy workflow relevant secrets ---\n'
cat -n .github/workflows/deploy.yml | sed -n '1,240p'
printf '\n--- all workflow secret usage ---\n'
rg -n 'github\.secrets\.|secrets:|CLOUDFLARE_API_TOKEN|wrangler secrets|put\.secret|secret\.bulk|environments:' .github/workflows package.json --glob '!node_modules' --glob '!.git' || trueRepository: kentcdodds/kody
Length of output: 26472
Sensitive Data Exposure (CWE-359)
Reachability: External
Move the alert recipient out of committed config.
ALERT_EMAIL_TO stores a personal email address in a public repository. Wrangler vars are plaintext, so this value is readable from the repo and remains in git history.
Set the recipient as a Worker secret in the deploy workflow, or use a role address such as alerts@heykody.dev. Use the same secret pattern already documented for CLOUDFLARE_API_TOKEN.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@packages/status/wrangler.jsonc` at line 41, Remove the personal value from
the ALERT_EMAIL_TO entry in the committed wrangler configuration and configure
the recipient through the deploy workflow as a Worker secret, following the
existing CLOUDFLARE_API_TOKEN secret pattern; alternatively, replace it with the
approved role address alerts@heykody.dev.
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit db2a7a6. Configure here.
|
🔎 Preview deployed: https://kody-pr-1230.kody-a99.workers.dev Worker: Mocks:
|
…end attempts, coalesced health refresh, deploy ordering

Public status page for kody
Implements the status page per decision record
docs/contributing/decisions/0004-status-page-separate-worker.md: an independently deployed worker atstatus.heykody.devthat observes the product from the outside, so it stays up when the main worker's deploys, code, or database are broken.What's in here
1. Component-level health endpoint (main worker). New public
GET /health/componentsruns cheap per-binding checks (APP_DB,AUDIT_DB,OAUTH_KV,COMMUNITY_ASSETS) with 3s timeouts, memoizes results for 10s (with coalesced concurrent refresh) so public traffic cannot amplify load, and returns 503 when any component fails.2. Status worker (
packages/status/). Mirrors thebackup-control-planepackage shape. A cron trigger runs one probe pass per minute against public endpoints only:/health,/health/components, the unauthenticated OAuth bearer challenge on/mcp, andkodyapps.dev. A singleStatusStoreDurable Object (SQLite) keeps minute samples (24h), daily uptime rollups (90 days), incidents, and notification history — deliberately neverAPP_DB. Components open an incident after 2 consecutive failures and resolve after 2 consecutive successes. The public page (server-rendered, no client JS, 60s meta refresh) shows a banner, per-component status with 90-day uptime bars, and open/past incidents;/status.jsonexposes the same snapshot, with a controlled 503 fallback if the Durable Object itself is unavailable.3. Operator email alerts. Sent to
me@kentcdodds.comfromkody@heykody.devvia the Cloudflare Email REST API (same mechanism as existing ops alerts). Policy: one email when an outage episode starts, then silence except one reminder per 24h while unresolved, one "all is well" email once everything recovers, all under a daily cap (STATUS_ALERT_DAILY_LIMIT, default 6). Every send attempt counts toward the cap (a permanently failing token cannot retry forever), and a cap-suppressed all-clear still sends on the first uncapped tick.Deploy
New path-filtered
deploy-status-workerjob in.github/workflows/deploy.yml(mirrors the backup control plane job) that runs after the main worker deploy so probes never hit a production worker missing/health/components. It deploys with the productionCLOUDFLARE_API_TOKEN, provisions thestatus.heykody.devcustom domain via wranglerroutes, setsBUILD_COMMIT, syncs the token as a Worker secret for email sending, and healthcheckshttps://status.heykody.dev/healthfor the deployed SHA.npm run validategains astatus:builddry-run; the status package's*.node.test.tsrun in the existing node-unit project.Review notes
All CodeRabbit majors and Bugbot's deploy-ordering race are addressed (stricter
/mcpand package-apps probe validation, degraded DO fallback, capped send attempts, coalesced health refresh,needs: deploy). Skipped: widening the status deploy path filter topackage.json/lockfile (mirrors the backup-control-plane precedent; manualworkflow_dispatchforces a status deploy) and the banner color-contrast nit.Evidence
Local end-to-end test: main dev worker + status worker (probes pointed at the local main worker), forced an outage by stopping the main worker, ran probe ticks, then recovered:
package_appsstayed operational), exactly one combined incident email decision was made, and on recovery all incidents resolved with exactly one all-clear decision (sends skipped locally without the token, but recorded — seestatus-alert-email-*log lines)./health/components. Fullnpm run validategreen locally (including the newdeploy-guardrailscheck from main).System recap — adds a new primitive (high risk)
Mode: recap · Base:
main@bde1cc1e· Head:19c92607Classification: adds — new
status-pageprimitive (independently deployed worker + Durable Object);app-uiis extended with a new public route.Primitives touched
status-pagepackages/status/worker, cron prober,StatusStoreDO, alertsapp-uiGET /health/componentsroute + handlerSystem map
The status worker probes the app's public endpoints every minute and stores results in its own Durable Object; nothing in the main worker depends on it.
Legend: green = composes (wiring only) · amber = extended by this PR · red = new primitive · gray = context (unchanged, included only when an edge crosses it).
Change flow
Invariants
Status data is global operational telemetry, not user data: the status worker holds no user state, no
APP_DBaccess, and no per-user Durable Objects, so the per-user isolation invariant is not implicated./health/componentsexposes only per-binding ok/latency, never data.Summary by CodeRabbit
New Features
Documentation