Skip to content

feat(spend-alerts): optimize sweep, fix panel UX, scope org recipients - #6670

Merged
iscekic merged 21 commits into
mainfrom
kwf/spend-alerts-optimize-ux-scopes-0d5a
Sep 28, 2026
Merged

iscekic merged 21 commits into
mainfrom
kwf/spend-alerts-optimize-ux-scopes-0d5a

Conversation

@iscekic

@iscekic iscekic commented Sep 23, 2026

Copy link
Copy Markdown
Collaborator

Changelog for users

  • The spend-alerts panel now names the scope an alert watches: Account: <name> or Organization: <name>.
  • Each rule now states what it measures: this scope's rolling spend over the chosen window, or its hourly rate against its own p95 baseline.
  • The web panel no longer leaves a large empty gap below Save, and the dashboard below still does not move as settings load, fail, or swap.
  • The mobile spend-alerts screen names the scope above the 24 h spend row and keeps the rule cards in place while settings load.
  • Organization spend alerts now reach the organization owner and members holding the billing_manager role; an admin without that role no longer receives them.

Changelog for maintainers

  • The rollup is one CTE: a window aggregate feeds a guarded upsert that skips unchanged buckets, and a join returns every re-derived scope with whether it changed; the scan stops at now.
  • The sweep decides 500-scope batches with one settings+rules+state read and one aggregate, caps a run at 5,000 scopes, and bounds the aggregate's hour_start at 720 h.
  • The run serves the rollup's real delta and still-firing rules first, then rotates a window over the remainder; the re-arm walk pages on the rule-state primary key.
  • Delivery moved to its own */5 cron; a daily 0 1 * * * cron prunes spend_alert_hourly buckets older than 30 days in 10,000-row batches.
  • Migration 0258 adds a covering index on microdollar_usage (created_at, kilo_user_id, organization_id, cost) with CREATE INDEX CONCURRENTLY, so the 1.6B-row table takes no blocking lock.
  • Organization recipients are the creator, owners, and billing_manager members; who may edit settings is unchanged, and the query stays scoped to the one organization.
  • Copy changed in the web panel (hardcoded English) and mobile en.json; the email already renders scope_name and kind_label, and the push lock-screen body stays content-free.
  • The rollup previously scanned 1.64M rows and 1,509,844 buffers per tick; the run log carries the re-measured after values, and the sweep's cap and rotation math is the riskiest part.

E2E proof

[e14] ux-check: On the web panel the line under the 'Spend alerts' title reads 'Account: ' on the personal view and 'Organization: ' on the organization view. — android emulator-5554, platform android: as the org owner (e14-login-owner.log 'signed in as e2e-org-owner-spend-alerts-optimize-ux-scopes-0d5a@example.com') the personal spend-alerts screen (Profile > Preferences > Spend alerts) shows the scope row label 'Account' with value 'E2E Org Owner' (e14-personal.log/e14-personal.png) and the organization screen (Profile > account selector 'Acme Corp' > Manage organization > Spend alerts) shows 'Organization' with 'Acme Corp' plus 'Threshold crossing'/'This scope's rolling spend over the chosen window crosses your limit' and 'Hourly spike'/'This…

[e14] ux-check: On the web panel the line under the 'Spend alerts' title reads 'Account: <name>' on the personal view and 'Organization: <name>' on the organization view. — prior/e14-org.png

[e13] ux-check: hold the loading skeleton then let the settings resolve — the 'Enable spend alerts' row and the rule cards do not move — android emulator-5554: nextjs was stalled so the settings query stayed pending and the screen mounted fresh from Profile > Preferences > Spend alerts; the skeleton digest shows the placeholder blocks (e13-layout.log: [55,342][1025,573] scope group, [55,628][1025,775] enable row), then nextjs was recovered and the same mounted screen resolved to the loaded form ('Account'/'E2E Org Owner', 'Enable spend alerts', 'Threshold crossing', 'Hourly spike'). Screenshots e13-skeleton.png and e13-loaded-enabled.png are captured for the visual reviewer, which owns the pixel judgement of the vertical…

[e13] ux-check: hold the loading skeleton then let the settings resolve — the 'Enable spend alerts' row and the rule cards do not move — prior/e13-loaded-enabled.png

[e8] needs:fault: mobile spend-alerts load error on the device — with the spendAlerts.get request failed the screen shows its error copy and a Retry, and the screen stays put — android emulator-5554 (platform android): state settings reported STATE HIT (e8-state.log) and nextjs was killed ('fault.sh: nextjs killed (port 7100 refuses; pids 787117 )', e8-fault-down.log); the scene reached the Spend alerts screen and logged 'SCENE e8 OK' with 'android.widget.TextView Couldn't load spend alerts. tappable [345,1235][736,1281]' and 'android.widget.Button Retry tappable [461,1318][619,1434]' under the screen header 'android.view.View Spend alerts tappable [111,102][1044,167]' with the tab bar still present, so the screen stayed put (capture e8.png for the visual reviewer)…

[e8] needs:fault: mobile spend-alerts load error on the device — with the spendAlerts.get request failed the screen shows its error copy and a Retry, and the screen stays put — scripted-shard1/e8.png

[e7] mobile spend-alerts screen on the device: the empty state, then enabling the alerts shows both rule cards — each with its push toggle — whose descriptions say what is measured and a scope line naming… — SCENE e7 OK (e7-scene.log): the empty state showed only 'Enable spend alerts'; after enabling the digest showed the scope row 'Account'/'e2e-mobile-spend-alerts-optimize-ux-scopes-0d5a-android', 'Threshold crossing' + "This scope's rolling spend over the chosen window crosses your limit", 'Hourly spike' + "This scope's hourly rate is far above its own p95 baseline" and the 'Threshold crossing and Push'/'Hourly spike and Push' toggles, Save showed 'Spend alert settings saved' and the run ended on 'Spend alerts are off. Turn them on to alert this scope's billing contacts.'; DB after (e7.log)…

[e7] mobile spend-alerts screen on the device: the empty state, then enabling the alerts shows both rule cards — each with its push toggle — whose descriptions say what is measured and a scope line naming… — scripted-shard1/e7.png

[e14] ux-check: On the web panel the line under the 'Spend alerts' title reads 'Account: ' on the personal view and 'Organization: ' on the organization view.

[e14] ux-check: On the web panel the line under the 'Spend alerts' title reads 'Account: <name>' on the personal view and 'Organization: <name>' on the organization view. — prior/e14-personal.png

[e13] ux-check: hold the loading skeleton then let the settings resolve — the 'Enable spend alerts' row and the rule cards do not move

[e13] ux-check: hold the loading skeleton then let the settings resolve — the 'Enable spend alerts' row and the rule cards do not move — prior/e13-skeleton.png

E2E proof — log excerpts

[e8] ux-check: On the web personal spend view and the organization usage-details -> pass :: jev read the digest: pass (confidence 0.99)
[e9] ux-check: The generic lock-screen push body contains no dollar figure. -> pass :: jev read the digest: pass (confidence 0.99)
/home/igor_kilocode_ai/.local/share/kwf/sections/spend-alerts-optimize-ux-scopes-0d5a/e2e-mobile-app/scripted-e8.log
android.widget.TextView Preferences tappable [171,697][953,743]
android.widget.TextView Appearance, notifications, thinking, and screen behavior tappable [171,747][953,784]
android.widget.Button Tutorial tappable [37,840][1043,974]
android.widget.TextView Tutorial tappable [171,884][953,930]
android.widget.TextView LINKED ACCOUNTS tappable [37,1029][1045,1073]
android.widget.TextView Email tappable [171,1129][1017,1175]
android.widget.TextView e2e-mobile-spend-alerts-optimize-ux-scopes-0d5a-android@example.com tappable [171,1179][1017,1216]
android.widget.TextView Test Account tappable [171,1300][1017,1346]
android.widget.TextView e2e-mobile-spend-alerts-optimize-ux-scopes-0d5a-android@example.com tappable [171,1350][1017,1387]
android.widget.Button Feedback tappable [37,1470][1043,1603]
android.widget.TextView Feedback tappable [170,1513][1016,1559]
android.widget.Button Privacy choices tappable [37,1631][1043,1765]
android.widget.TextView Privacy choices tappable [170,1674][1016,1720]
android.widget.Button Sign out tappable [37,1793][1043,1927]
android.widget.TextView Sign out tappable [170,1836][1016,1882]
android.widget.Button Delete Account tappable [37,1954][1043,2088]
android.widget.TextView Delete Account tappable [170,1997][1016,2043]
android.widget.TextView v1.0.12 (1) tappable [37,2115][1045,2152]
android.widget.Button Home, tab, 1 of 3 tappable [0,2195][360,2337]
android.widget.TextView HOME tappable [13,2281][347,2320]
android.widget.Button Agents, tab, 2 of 3 tappable [360,2195][720,2337]
android.widget.TextView AGENTS tappable [373,2281][707,2320]
android.widget.Button Profile, tab, 3 of 3 tappable [720,2195][1080,2337]
android.widget.TextView PROFILE tappable [733,2281][1067,2320]
/home/igor_kilocode_ai/.local/share/kwf/sections/spend-alerts-optimize-ux-scopes-0d5a/e2e-mobile-app/scripted-e9.log
android.widget.TextView LIVE NOW tappable [36,282][537,319]
android.widget.Button See all tappable [555,281][1043,320]
android.widget.TextView SEE ALL tappable [897,281][1043,320]
android.widget.TextView Nothing running right now tappable [348,428][732,474]
android.widget.Button New coding task tappable [37,602][1043,718]
android.widget.TextView New coding task tappable [450,636][696,682]
android.widget.Button New task from a picture tappable [37,736][1043,852]
android.widget.TextView New task from a picture tappable [396,770][749,816]
android.widget.TextView EXPLORE tappable [36,887][1044,924]
android.widget.Button Code Reviewer, Automatic PR reviews tappable [37,943][1043,1087]
android.widget.TextView Code Reviewer tappable [171,971][953,1017]
android.widget.TextView Automatic PR reviews tappable [171,1021][953,1058]
android.widget.Button Security Agent, Find and remediate vulnerabilities tappable [37,1105][1043,1249]
android.widget.TextView Security Agent tappable [171,1133][953,1179]
android.widget.TextView Find and remediate vulnerabilities tappable [171,1183][953,1220]
android.widget.Button PR Review, Review pull requests on mobile tappable [37,1268][1043,1410]
android.widget.TextView PR Review tappable [171,1296][953,1342]
android.widget.TextView Review pull requests on mobile tappable [171,1346][953,1383]
android.widget.Button Home, tab, 1 of 3 tappable [0,2195][360,2337]
android.widget.TextView HOME tappable [13,2281][347,2320]
android.widget.Button Agents, tab, 2 of 3 tappable [360,2195][720,2337]
android.widget.TextView AGENTS tappable [373,2281][707,2320]
android.widget.Button Profile, tab, 3 of 3 tappable [720,2195][1080,2337]
android.widget.TextView PROFILE tappable [733,2281][1067,2320]
  • proved live: mobile spend-alerts screen on the device: the empty state, then enabling the alerts shows both rule cards — each with its push toggle — whose descriptions say what is measured and a scope line naming… — SCENE e7 OK (e7-scene.log): the empty state showed only 'Enable spend alerts'; after enabling the digest showed the scope row 'Account'/'e2e-mobile-spend-alerts-optimize-ux-scopes-0d5a-android', 'Threshold crossing' + "This scope's rolling spend over the chosen window crosses your limit", 'Hourly spike' + "This scope's hourly rate is far above its own p95 baseline" and the 'Threshold crossing and Push'/'Hourly spike and Push' toggles, Save showed 'Spend alert settings saved' and the run ended on 'Spend alerts are off. Turn them on to alert this scope's billing contacts.'; DB after (e7.log)… mobile spend-alerts screen on the device: the empty state, then enabling the alerts shows both rule cards — each with its push toggle — whose descriptions say what is measured and a scope line naming… — e7.png
  • proved live: needs:fault: mobile spend-alerts load error on the device — with the spendAlerts.get request failed the screen shows its error copy and a Retry, and the screen stays put — android emulator-5554 (platform android): state settings reported STATE HIT (e8-state.log) and nextjs was killed ('fault.sh: nextjs killed (port 7100 refuses; pids 787117 )', e8-fault-down.log); the scene reached the Spend alerts screen and logged 'SCENE e8 OK' with 'android.widget.TextView Couldn't load spend alerts. tappable [345,1235][736,1281]' and 'android.widget.Button Retry tappable [461,1318][619,1434]' under the screen header 'android.view.View Spend alerts tappable [111,102][1044,167]' with the tab bar still present, so the screen stayed put (capture e8.png for the visual reviewer)… needs:fault: mobile spend-alerts load error on the device — with the spendAlerts.get request failed the screen shows its error copy and a Retry, and the screen stays put — e8.png
  • proved live: needs:fault: mobile spend-alerts save failure on the device — with the spendAlerts.save request failed, Save shows the save-error copy with a Retry and the edited draft is kept — android emulator-5554 (platform android): state settings STATE HIT (e10-state5.log); with nextjs up the form loaded ('SCENE e10 OK', e10-load.log: 'android.widget.TextView Account tappable [83,425][206,471]' and 'Email or push an alert to the billing contacts when this scope's own spend crosses a limit or spikes above its usual hourly rate. tappable [55,249][1025,341]', capture e10-loaded.png), then nextjs was killed ('fault.sh: nextjs killed (port 7100 refuses; pids 1691028 )', e10-fault-down.log) and Save logged 'SCENE e10 OK' with 'android.widget.TextView Couldn't save spend alerts.… needs:fault: mobile spend-alerts save failure on the device — with the spendAlerts.save request failed, Save shows the save-error copy with a Retry and the edited draft is kept — e10-loaded.png
  • proved live: ux-check: On the web panel the line under the 'Spend alerts' title reads 'Account: ' on the personal view and 'Organization: ' on the organization view. — android emulator-5554, platform android: as the org owner (e14-login-owner.log 'signed in as e2e-org-owner-spend-alerts-optimize-ux-scopes-0d5a@example.com') the personal spend-alerts screen (Profile > Preferences > Spend alerts) shows the scope row label 'Account' with value 'E2E Org Owner' (e14-personal.log/e14-personal.png) and the organization screen (Profile > account selector 'Acme Corp' > Manage organization > Spend alerts) shows 'Organization' with 'Acme Corp' plus 'Threshold crossing'/'This scope's rolling spend over the chosen window crosses your limit' and 'Hourly spike'/'This… ux-check: On the web panel the line under the 'Spend alerts' title reads 'Account: <name>' on the personal view and 'Organization: <name>' on the organization view. — e14-org.png
  • proved live from the UI tree: ux-check: The generic lock-screen push body contains no dollar figure. — jev read the digest: pass (confidence 0.99) (scripted-e9.log)

[e14] ux-check: On the web panel the line under the 'Spend alerts' title reads 'Account: <name>' on the personal view and 'Organization: <name>' on the organization view. — prior/e14-org.png

[e13] ux-check: hold the loading skeleton then let the settings resolve — the 'Enable spend alerts' row and the rule cards do not move — prior/e13-loaded-enabled.png

[e7] mobile spend-alerts screen on the device: the empty state, then enabling the alerts shows both rule cards — each with its push toggle — whose descriptions say what is measured and a scope line naming… — prior/e7.png

[e14] ux-check: On the web panel the line under the 'Spend alerts' title reads 'Account: <name>' on the personal view and 'Organization: <name>' on the organization view. — prior/e14-personal.png

[e13] ux-check: hold the loading skeleton then let the settings resolve — the 'Enable spend alerts' row and the rule cards do not move — prior/e13-skeleton.png

[e8] needs:fault: mobile spend-alerts load error on the device — with the spendAlerts.get request failed the screen shows its error copy and a Retry, and the screen stays put — prior/e8.png

/home/igor_kilocode_ai/.local/share/kwf/sections/spend-alerts-optimize-ux-scopes-0d5a/e2e-web/web-e2e.log
apps/web/.env.local: line 87: PRIVATE: command not found
Running 1 test using 1 worker
  ✓  1 [chromium] › tests/setup-smoke/profile.spec.ts:10:7 › local setup smoke › signs in with fake auth and renders the profile page (7.8s)
  1 passed (10.5s)
Owner request

Surface: backend (spend-alert sweep and delivery) + web spend view + mobile spend-alerts screen

Optimize the spend-alert feature, fix its two visible UX defects, and settle the
alert scopes. One item, three parts, one PR.

Every number in Part 1 is measured on the EU read replica on 2026-09-23
~13:20-13:28 UTC. Re-measure before and after; the PR body must carry both.

Part 1 — optimization and bugs

Measured today

  • microdollar_usage: 1,598,971,136 rows, 332 GB table, 738 GB with indexes.
    Not partitioned. ~486k rows written per hour (1,456,624 rows in a 3-hour
    window; 489,355 in one hour; 10,506,318 in 24 hours).
  • One cron tick re-derives every one of those 3-hour rows: 1,638,671 rows
    scanned, 2,260 ms, 1,509,844 shared buffers (~11.5 GB) per tick.
    At 12 ticks/hour that is ~27 s of DB time and ~138 GB of buffer traffic per
    hour, waking 1.5M pages every 5 minutes.
  • Candidate scopes per tick (distinct scopes with usage in the 3-hour window):
    17,700. Each one costs one sequential awaited round trip.
  • The skip path (17,670 of the 17,700 have no settings row) measures
    0.028 ms execution but 0.186 ms planning — planning dominates 6:1, and the
    loop is one round trip per scope.
  • Upserts per tick: my count over the window is 37,237; the 13:20:03 tick wrote
    38,915 rows. Per hour that is ~447k row versions to store ~17,700 new
    buckets. The upsert has no guard, so it rewrites a bucket whose sum did not
    change (sweep.ts:381).
  • The true 5-minute delta is 3,820 scopes (1-minute: 1,926; 15-minute:
    6,041). So ~78% of the per-tick decision work and ~90% of the written row
    versions are work on data that did not change.
  • Config: cron */5 (apps/web/vercel.json:136), maxDuration = 300,
    delivery limit 50 per tick (600/hour).
  • Only 30 scopes have settings (29 enabled: 28 personal, 2 org). 7 rules are
    in the firing state. So today the feature is a 17,700-row scan per tick to
    serve 30 configured scopes.

Bugs and risks to fix

  1. The decision loop cannot finish at scale, and it starves delivery. It is
    one sequential awaited round trip per candidate (sweep.ts:608, 615-617),
    with no budget, inside a 300 s route. The delivery drain runs after the
    sweep in the same request (route.ts), so a slow or killed sweep delivers
    nothing that tick either. Make the per-scope work set-based — one aggregate
    with scope_key = ANY($1) GROUP BY scope_key, plus one joined
    settings+rules+state read for the batch — and cap the work per run. Drain
    the outbox separately from the sweep so one cannot starve the other.
  2. The rollup rewrites unchanged buckets. Add
    WHERE spend_alert_hourly.cost_microdollars IS DISTINCT FROM EXCLUDED.cost_microdollars
    to the ON CONFLICT DO UPDATE (sweep.ts:381). That makes RETURNING
    return the real delta (3,820 scopes instead of 17,700) and cuts the hourly
    row-version churn from ~447k to ~46k. The re-arm pass still clears a firing
    rule, so the one-alert guarantee holds.
  3. A per-scope read grows with table age. The aggregate's only WHERE clause
    is scope_key = ${scopeKey} (sweep.ts:508); the 24 h, 7 d, 14 d and 30 d
    ranges are aggregate FILTERs, which Postgres cannot push into an index
    qual. It therefore reads every bucket the scope ever stored. Measured today:
    41 rows, 0.258 ms — small only because the table is 2 days old. After a year
    it is ~8,760 rows per scope per tick. Add the hour_start lower bound to the
    WHERE clause.
  4. The re-arm walk is quadratic. spend_alert_rule_state has a primary key
    on rule_id only, and scope_key lives in spend_alert_settings, so every
    500-row page (FIRING_SCOPE_PAGE_SIZE) is a full scan, join and sort of the
    whole firing set: O(F²/500). Make the re-arm set directly indexable.
  5. Nothing is ever pruned. No retention job exists for
    spend_alert_hourly (383,010 rows now, 71,735 distinct scopes, 82 MB total
    for a 39 MB table). The longest window a rule can use is 720 h, so a bucket
    older than 30 days can never change a decision. Add a retention window, or
    state in the PR body why the history is kept.
  6. microdollar_usage.created_at contains future timestamps. The 13:20:03
    tick wrote buckets covering hour_start 10:00 through 23:00 while now was
    13:20. A bucket for a future hour breaks the "current partial hour"
    assumption the anomaly rule and the baseline both rely on. Find the producer,
    then decide how the rollup treats a future-dated row.
  7. The rollup's scan cost. Consider idx_created_at with INCLUDE (kilo_user_id, organization_id, cost) so the 1.6M-row scan becomes an
    index-only scan instead of 1.5M heap fetches. A watermark-driven incremental
    rollup with a periodic full-window repair is the larger option. Follow the
    repo's migration convention for a large table; do not take a blocking lock.
  8. Keep the cron's scopesTouched summary and add per-phase timings, so the
    next regression is visible in the logs without a manual measurement.

Do not change the correctness properties that already hold: the re-derivation
is idempotent, the dedupe key collapses a repeated episode, the delivery claim
fences on attempt_count, and a rule never fires with no outbox row.

Part 2 — the two UX defects

The owner's screenshot shows the spend-alerts settings screen: Spend alerts
title, an Enable spend alerts row, a Threshold crossing card (Limit USD,
Window) and an Hourly spike card (Spike multiplier), each with Email and Push
toggles, and a Save button bottom-right. Two defects:

A. Excessive empty space below Save. Cause found on web: the panel reserves
a fixed worst-case height,
SPEND_ALERTS_PANEL_SLOT_CLASS = 'min-h-[66rem] @min-[310px]:min-h-[58rem] @min-[380px]:min-h-[55rem] @min-[580px]:min-h-[52rem]'
(spendAlertsPanelState.ts:123), so a loaded form that is shorter than the
reservation shows dead space. Measure the real worst-case ready-form height per
band and cut the reservation to it, or reserve the height only while the
settings load. Keep the property the reservation exists for: the dashboard
below must not move when the form arrives, fails, or swaps. Check the mobile
screen (spend-alerts-screen.tsx, which renders no bottom spacer today) for the
same defect at its own widths, and fix it there if it has it.

B. The copy does not say what the alert is for. A reader cannot tell
whether an alert watches their own account or an organization.

  • Web copy is hardcoded English in SpendAlertsPanel.tsx: Spend alerts
    (lines 142, 230), Threshold crossing / Hourly spike (line 337),
    Rolling spend crosses your limit / An hour is far above this scope's usual rate (lines 341-342).
  • Mobile copy is the spendAlerts block in
    apps/mobile/src/i18n/locales/en.json:3217-3221.

Make the scope explicit in the panel: name the owner the settings belong to
(this account, or the organization's name), and make each rule's description
say what is measured (the scope's rolling spend over the chosen window; the
scope's hourly rate against its own p95 baseline). The alert itself must carry
the same clarity: the email already has scope_name and kind_label
(email.ts:419-426) — check that the rendered template states the scope and the
kind plainly. The push lock-screen body is deliberately content-free
(generic.body.spendAlert = "Your spend needs attention",
push-presentation.ts:341-350); keep spend figures off the lock screen, but the
title or body may name the scope and the alert kind. State in the PR body which
copy you changed and which you deliberately left.

i18n rule: add and edit copy in en.json only, reuse an existing key when the
copy matches, and run pnpm check:i18n. The web panel has no i18n today —
do not invent a translation layer for it in this item.

Part 3 — the alert scopes

The owner wants three scopes to work. Some already do; confirm each with proof
and close the gaps.

  • Personal account. Exists: no organizationId means the caller's own
    scope (spend-alert-router.ts:213-224), reached from
    preferences-screen.tsx:111 on mobile and the personal usage view on web.
  • An organization. Exists: organizationId selects the organization scope
    (spend-alert-router.ts:242-250), reached from hub-screen.tsx:212 on mobile
    and the organization usage-details view on web.
  • An entire organization, delivered to the owner and billing managers only.
    Today the recipient set is every membership whose role is in
    ORGANIZATION_BILLING_ROLES, which is owner, admin and billing_manager
    (packages/app-shared/src/organizations/roles.ts:18-22), read in
    settings.ts:authorizedBillingContacts. Change the recipients of an
    organization-scope alert to the organization owner and members holding the
    billing_manager role. Do not change who may edit the settings. Confirm the
    owner is included even when their membership row is not one of those roles.

Proof

  • Part 1: the before and after of every measured number above — the rollup's
    execution time and buffers from EXPLAIN (ANALYZE, BUFFERS), rows scanned and
    rows touched per tick, candidate count per tick, and the query count per tick.
    Quote the decisive log lines. State the tick interval you observed while
    measuring.
  • Part 2: a screenshot of the settings screen after the padding fix, at the
    widths where the reservation applied, and a screenshot of the copy change on
    each surface the copy exists on (the web spend view and the mobile
    spend-alerts screen). # platform is not narrowed, so prove the mobile screen
    on the platform the section proves. A screenshot of the rendered alert email
    and of the push copy.
  • Part 3: proof that each of the three scopes configures, saves and alerts, and
    proof of the recipient set for an organization scope (owner and billing
    managers receive; an admin without the billing role does not).

Repo rules that bind this change

  • The owner did not grant # allow-e2e-code. Keep fixtures and test-only
    runtime support out of the product diff.
  • apps/mobile/AGENTS.md: haptics for commits and outcomes only; new copy goes
    to en.json only.
  • A migration on a 1.6B-row table must not take a blocking lock.

[e9] mobile spend-alerts permission denied on the device — device account is a non-billing org member (db.sh: role 'member'); the org-scope screen shows 'Access denied' and 'You don't have permission to manage spend alerts.' with no settings controls (e9-permission.log); the org hub renders no 'Spend alerts' row for that role (e9-hub.log: 0 occurrences), so the described hub-row route does not exist for a member and the screen was reached by the app's own spend-alert deep link (e9-open.log: result ok, mode deeplink) — plan-route discrepancy, [pre-existing], consistent with the other billing rows hidden for the same member.

[e9] mobile spend-alerts permission denied on the device — e9-hub.png

[e15] ux-check: mobile spend-alerts summary-group scope row and rule-card descriptions (personal and organization) — e15-personal.txt (+e15-personal.png): first summary-group row is 'Account'/'E2E Org Owner' with both rule descriptions; e15-org.txt (+e15-org.png): first row is 'Organization'/'Acme Corp' with the same two descriptions (e15.log); both visited screens captured. No UX-DEFECT observed.

[e15] ux-check: mobile spend-alerts summary-group scope row and rule-card descriptions (personal and organization) — e15-org.png

[e15] ux-check: mobile spend-alerts summary-group scope row and rule-card descriptions (personal and organization)

[e15] ux-check: mobile spend-alerts summary-group scope row and rule-card descriptions (personal and organization) — e15-personal.png

[e9] mobile spend-alerts permission denied on the device

[e9] mobile spend-alerts permission denied on the device — e9-permission.png

[e18] ux-check: configure and trigger a threshold crossing for an organization whose created_by_kilo_user_id is null and whose only owner holds the owner membership role — the owner receives the alert — android emulator-5554: signed in on the device as the org owner (e18-login.log) and switched to the org in Profile > account selector, then enabled spend alerts and saved a 24 h / $1 threshold from the org spend-alerts screen (e18-recipient.log: org:b0622be7-6122-4b60-bbcc-e7056e527516|true|threshold|true|1000000|24|true|false; screen e18-org-configured.png). The org row is 'b0622be7-6122-4b60-bbcc-e7056e527516|Acme Corp|NULL' and its only owner d839e489-… holds role 'owner'. The sweep tick fired 1 alert and the delivery row it enqueued named the owner (e18-recipient.log: recipients…

[e18] ux-check: configure and trigger a threshold crossing for an organization whose created_by_kilo_user_id is null and whose only owner holds the owner membership role — the owner receives the alert — e18-org-configured.png

[e10] needs:fault: mobile spend-alerts save failure on the device — with the spendAlerts.save request failed, Save shows the save-error copy with a Retry and the edited draft is kept — android emulator-5554 (platform android): state settings STATE HIT (e10-state5.log); with nextjs up the form loaded ('SCENE e10 OK', e10-load.log: 'android.widget.TextView Account tappable [83,425][206,471]' and 'Email or push an alert to the billing contacts when this scope's own spend crosses a limit or spikes above its usual hourly rate. tappable [55,249][1025,341]', capture e10-loaded.png), then nextjs was killed ('fault.sh: nextjs killed (port 7100 refuses; pids 1691028 )', e10-fault-down.log) and Save logged 'SCENE e10 OK' with 'android.widget.TextView Couldn't save spend alerts.…

[e10] needs:fault: mobile spend-alerts save failure on the device — with the spendAlerts.save request failed, Save shows the save-error copy with a Retry and the edited draft is kept — e10-save-error.png

[e16] rendered spend-alert email names the scope and the kind — drain cron rendered fresh emails from db.sh-seeded delivery rows: personal shows 'Your account' + 'Spend threshold', org shows 'Acme Corp' + 'Spend threshold', org anomaly shows 'Acme Corp' + 'Hourly spike' (e16-email-.log); rows ended 'sent' (e16-deliveries.log); PNGs e16-email-.png captured for the visual reviewer; observation: an anomaly delivery with a null payload.thresholdMicrodollars fails 'spend_alert_delivery_missing_payload' (e16-anomaly-null-threshold.log), so the 'Hourly spike' email may not deliver when the sweep emits a null threshold.

[e16] rendered spend-alert email names the scope and the kind — e16-email-personal.png

Follow-ups (not changed here)

  • not proved live: needs:seed: mobile spend-alerts permission denied on the device — signed in as a member of an organization whose role is not a billing role, the organization hub's Spend alerts row opens a screen with no settings and only the permission message (no capture cited it)
  • not proved live: ux-check: Configure and trigger a threshold crossing for an organization whose created_by_kilo_user_id is null and whose only owner holds the owner membership role: the owner receives the alert. (no capture cited it)
  • not proved live: ux-check: On the mobile spend-alerts screen (personal and organization), the first row of the summary group reads 'Account'/'Organization' with the scope's name, and each rule card's description names the scope and what is measured ('This scope's rolling spend over the chosen window crosses your limit' / 'This scope's hourly rate is far above its own p95 baseline'). (no capture cited it)
  • not proved live: ux-check: On the web panel, hold the loading skeleton then let the settings resolve: the vertical position of the 'Enable spend alerts' row and the rule cards does not move between the skeleton and the loaded form. (no capture cited it)
  • not proved live: ux-check: On the web personal spend view and the organization usage-details view, after the settings load, the gap between the Save button and the bottom edge of the panel card is at most a few pixels at card widths ~385px, ~441px, and ~786px (no large empty band below Save). (no capture cited it)
  • not proved live: ux-check: Open the rendered spend-alert email: it names the scope ('Your account' or the organization name) and the kind ('Spend threshold' or 'Hourly spike'). (no capture cited it)
  • not proved live: iOS: not run — the diff forks on no platform, so Android proves both

…ucket retention job (kwf spend-alerts-optimize-ux-scopes-0d5a/s3)
…nd name the scope and the measures (kwf spend-alerts-optimize-ux-scopes-0d5a/s5)
…nly (kwf spend-alerts-optimize-ux-scopes-0d5a/s4)
…elta, phase timings (kwf spend-alerts-optimize-ux-scopes-0d5a/s1)
…measures, and check for dead space (kwf spend-alerts-optimize-ux-scopes-0d5a/s6)
…ot move the dashboard (kwf spend-alerts-optimize-ux-scopes-0d5a/c1)
…u (kwf spend-alerts-optimize-ux-scopes-0d5a/ux1)
…s (kwf spend-alerts-optimize-ux-scopes-0d5a/ux2)
@iscekic
iscekic marked this pull request as draft September 23, 2026 22:58
Comment thread packages/db/src/migrations/meta/_journal.json Outdated
Comment thread apps/web/src/lib/spend-alerts/sweep.ts
@kilo-code-bot

kilo-code-bot Bot commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Code Review Summary

Status: No Issues Found | Recommendation: Merge

Executive Summary

The incremental commit is test-only hygiene (organization cleanup moved from afterAll to afterEach in delivery.test.ts); it is correct, afterAll is no longer referenced, and every prior finding is resolved at current HEAD.

Files Reviewed (1 file)
  • apps/web/src/lib/spend-alerts/delivery.test.ts

Prior Findings Re-checked

  • Migration-number collision on packages/db/src/migrations/meta/_journal.json — resolved: regenerated as 0263_spend_alerts_rollup_covering_index with a newer when than every earlier entry.
  • apps/web/src/lib/spend-alerts/sweep.ts:922 remainder coverage bound — resolved: ROLLUP_WINDOW_TICKS/SKIP_PATH_ROTATION_TICKS now derive and log the bound.
  • apps/web/src/lib/spend-alerts/settings.ts:531 removed organization creator recipients — resolved: creator branch requires a current membership and the drain re-resolves contacts.
Previous Review Summaries (3 snapshots, latest commit 598179b)

Current summary above is authoritative. Previous snapshots are kept for context only.

Previous review (commit 598179b)

Status: No Issues Found | Recommendation: Merge

Executive Summary

The incremental commit resolves the prior open finding — a removed organization creator no longer receives organization spend alerts — by requiring a current membership row for the creator branch and re-resolving organization recipients at drain time; no new issues were found in the changed code.

Files Reviewed (4 files)
  • apps/web/src/lib/spend-alerts/settings.ts
  • apps/web/src/lib/spend-alerts/settings.test.ts
  • apps/web/src/lib/spend-alerts/delivery.ts
  • apps/web/src/lib/spend-alerts/delivery.test.ts

Previous review (commit 11c20cb)

Status: No Issues Found | Recommendation: Merge

Executive Summary

The incremental commits resolve the prior migration-number collision (the rollup covering index is regenerated as migration 0263_spend_alerts_rollup_covering_index) and the sweep remainder-window bound (the rotation constants are now derived from LOOKBACK_HOURS/TICK_INTERVAL_MS and an overflow warning is logged); no new issues were found in the changed code.

Files Reviewed (5 files)
  • apps/web/src/lib/spend-alerts/sweep.ts
  • apps/web/src/lib/spend-alerts/sweep.test.ts
  • packages/db/src/migrations/0263_spend_alerts_rollup_covering_index.sql
  • packages/db/src/migrations/meta/_journal.json
  • packages/db/src/spend-alerts-rollup-index-migration.test.ts

Previous review (commit f41e511)

Status: 2 Issues Found | Recommendation: Address before merge

Executive Summary

The spend-alert sweep, retention, recipient, and UX changes are sound; the merge-blocking risk is that the new migration reuses number 0258, which the target branch already uses, so the covering index would be silently skipped after rebase.

Overview

Severity Count
CRITICAL 1
WARNING 0
SUGGESTION 1
Issue Details (click to expand)

CRITICAL

File Line Issue
packages/db/src/migrations/meta/_journal.json 1815 Appends migration 0258_spend_alerts_rollup_covering_index (idx 258, when 1790173354576) while main (base c09a1708) already has 0258_github_connection_role (idx 258) and 0259_github_connection_role_indexes (idx 259). After rebase the new entry's folderMillis is older than the newest applied row, so Drizzle skips it and the covering index is never built; the merged journal also fails migration-journal.test.ts.

SUGGESTION

File Line Issue
apps/web/src/lib/spend-alerts/sweep.ts 910 The remainder floor cap (2500) bounds coverage to ceil(rederived.length / 2500) ticks, so the documented "reached within the rollup window" guarantee only holds below roughly 90,000 candidate scopes.
Files Reviewed (26 files)
  • apps/mobile/src/components/organization/spend-alerts-screen.mounted.test.tsx
  • apps/mobile/src/components/organization/spend-alerts-screen.tsx
  • apps/mobile/src/i18n/locales/en.json
  • apps/web/src/app/api/cron/dispatch-spend-alerts/route.test.ts
  • apps/web/src/app/api/cron/dispatch-spend-alerts/route.ts
  • apps/web/src/app/api/cron/drain-spend-alert-deliveries/route.test.ts
  • apps/web/src/app/api/cron/drain-spend-alert-deliveries/route.ts
  • apps/web/src/app/api/cron/prune-spend-alert-hourly/route.test.ts
  • apps/web/src/app/api/cron/prune-spend-alert-hourly/route.ts
  • apps/web/src/components/spend-alerts/SpendAlertsPanel.tsx
  • apps/web/src/components/spend-alerts/spendAlertsPanelState.test.ts
  • apps/web/src/components/spend-alerts/spendAlertsPanelState.ts
  • apps/web/src/lib/spend-alerts/retention.test.ts
  • apps/web/src/lib/spend-alerts/retention.ts
  • apps/web/src/lib/spend-alerts/settings.test.ts
  • apps/web/src/lib/spend-alerts/settings.ts
  • apps/web/src/lib/spend-alerts/sweep.test.ts
  • apps/web/src/lib/spend-alerts/sweep.ts
  • apps/web/vercel.json
  • packages/app-shared/src/organizations/roles.test.ts
  • packages/app-shared/src/organizations/roles.ts
  • packages/db/src/migrations/0258_spend_alerts_rollup_covering_index.sql
  • packages/db/src/migrations/meta/0258_snapshot.json (generated)
  • packages/db/src/migrations/meta/_journal.json
  • packages/db/src/schema.ts
  • packages/db/src/spend-alerts-rollup-index-migration.test.ts

Fix these issues in Kilo Cloud


Reviewed by deepseek-v4.1-flash · Input: 0 · Output: 0 · Cached: 0

Review guidance: REVIEW.md from base branch main

The branch carried 0258_spend_alerts_rollup_covering_index.sql with idx 258,
which collides with main's own 0258_github_connection_role. Merge origin/main
and regenerate the migration with the drizzle CLI.

- Drop the branch's 0258 migration file and its snapshot.
- Take main's migration folder verbatim (0254-0259).
- Regenerate as 0260_spend_alerts_rollup_covering_index.sql (journal idx 260).
- Restore the COMMIT;/BEGIN; boundary the transactional migrator needs for
  CREATE INDEX CONCURRENTLY, which spend-alerts-rollup-index-migration.test.ts
  asserts. The regenerated SQL matches the previous file byte for byte.

Guards: migration-journal.test.ts and spend-alerts-rollup-index-migration.test.ts
both pass.
… window

The run's cap defers the re-derived remainder, and the rotation only finishes inside the LOOKBACK_HOURS window while the remainder fits under MAX_REMAINDER_FLOOR x the window's ticks. Above that a deferred scope stops being re-derived before its turn, so a crossing it would have alerted on is never decided. Derive the window's tick count and the rotation target from LOOKBACK_HOURS and the cron interval instead of the bare literals, state the bound the guarantee actually has, and log a warning with the shortfall when the rotation cannot meet it.
Assert the per-tick phase log reports remainderTicks against rollupWindowTicks, and that a remainder needing exactly the whole window reports 36 of 36 without warning.
…rating

main has since appended its own 0260-0262 entries, so the branch's generated SQL, snapshot and journal entry are removed and will be regenerated at the next free index from the merged schema.
pnpm drizzle-kit generate --name spend_alerts_rollup_covering_index on the merged tree emits the concurrent btree over (created_at, kilo_user_id, organization_id, cost) and nothing else; the COMMIT;/BEGIN; migrator boundaries are appended per packages/db/AGENTS.md.
@iscekic iscekic added the human-ready The PR is ready for human review. label Sep 25, 2026
@iscekic
iscekic marked this pull request as ready for review September 25, 2026 13:49
@iscekic iscekic added human-ready The PR is ready for human review. merge-by-human the merge bot routed this PR to a human and removed human-ready The PR is ready for human review. labels Sep 25, 2026
Comment thread apps/web/src/lib/spend-alerts/settings.ts
A removed organization creator stayed eligible for the organization's
spend alerts. The creator branch read `created_by_kilo_user_id`, which is
history and is never cleared, and treated it as an owner even when the
creator no longer had a membership, which is the only thing
`ensureOrganizationAccess` reads as access.

Require a membership row of the organization for the creator branch, so
the role stays unrestricted (a creator whose membership says `member`
still receives the alert) but a removed creator does not.

Re-check the organization's current contacts when the drain sends a
queued alert as well: the outbox row carries the recipients resolved at
enqueue time and may be sent an hour later, so the snapshot is now
narrowed to the contacts that still hold access. A personal scope is
unchanged.
@iscekic
iscekic marked this pull request as draft September 25, 2026 19:09
@iscekic
iscekic marked this pull request as ready for review September 25, 2026 19:09
The organization cleanup of this suite ran in an afterAll hook. The jest worker teardown closes the database pool in its own afterAll, and that hook runs first. The delete therefore hit a closed pool: "Failed query: delete from organizations where id in ($1, $2)", cause "Cannot use a pool after calling end on the pool". All 14 tests passed, but the suite failed on that hook.

Move the cleanup into the existing afterEach, which runs before the pool closes. No dependent rows need removal first: the membership rows have no foreign key to organizations, and the hook already deletes the suite's own delivery rows.
@iscekic iscekic removed the human-ready The PR is ready for human review. label Sep 27, 2026
@iscekic
iscekic requested a review from pandemicsyn September 28, 2026 13:02
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

merge-by-human the merge bot routed this PR to a human

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants