Skip to content

fix: reconcile usage_events.hive_credit_delta with the ledger (#1180) - #1194

Merged
sakibsadmanshajib merged 3 commits into
mainfrom
fix/usage-ledger-reconciliation
Aug 25, 2026
Merged

sakibsadmanshajib merged 3 commits into
mainfrom
fix/usage-ledger-reconciliation

Conversation

@sakibsadmanshajib

@sakibsadmanshajib sakibsadmanshajib commented Aug 25, 2026 •

Copy link
Copy Markdown
Owner

Summary

Fixes #1180. usage_events.hive_credit_delta and credit_ledger_entries.credits_delta disagreed for the same request. Investigation found two independent, confirmed-live root causes, not one:

Root cause A (the primary defect, not a rescale/rounding artifact): a genuine duplicate write. Every reservation-backed request writes TWO usage_events rows with event_type='completed' for the same attempt:

  • Control-plane's own accounting.finalizeLocked (apps/control-plane/internal/accounting/service.go) writes the authoritative row during FinalizeReservation, in-process. hive_credit_delta = -actualCredits, the same figure charged to the ledger in the same call.
  • edge-api's recordCompletedEvent (apps/edge-api/internal/inference/orchestrator.go, stream.go) then makes a separate, unconditional HTTP POST to /internal/usage/events that writes a second row for the identical attempt. Its hive_credit_delta is set to usage.TotalTokens — a raw token count, not a credit amount, not even negative.

Live measurement on the demo box (2026-08-25, full history, read-only bounded queries): 576 duplicate pairs, zero triples. 576/576 have the earlier row's delta <= 0 (ledger-matching); 575/576 have the later row's delta equal to input_tokens + output_tokens exactly. The ordering is deterministic by the code path, not a race: edge-api's POST only fires after the synchronous FinalizeReservation call returns.

Root cause B (a real stale-unit bug, distinct from A): a missed backfill. usage_events.hive_credit_delta was omitted from the column list in 20260823_40_credit_unit_rescale_billion.sql (D-046, factor 10000). credit_ledger_entries.credits_delta was backfilled by that migration; usage_events.hive_credit_delta was not. Every authoritative row written before the rescale boundary is off from the ledger by exactly 10000x. Confirmed live (masked attempt id): -12 in usage_events vs -120000 in the ledger for the identical attempt, -12 * 10000 = -120000 exactly.

Which surface is authoritative: credit_ledger_entries, confirmed from code (it's what actually posts against the account balance) rather than assumed. usage_events's correct row is a mirror of the same actualCredits value, not an independent computation.

Does anything bill from the wrong row: no live billing/invoicing path reads usage_events.hive_credit_delta. GetSpendSummary and GetUsageSummary.total_credits_spent filter WHERE hive_credit_delta < 0, so the always-positive wrong duplicate was silently excluded from those two figures by accident of the sign filter — not corrupted, but GetUsageSummary.request_count (unfiltered COUNT(*)) was inflated up to 2x, and the raw ListEvents listing showed two contradictory rows per request.

Fix

Chose to make the two surfaces agree (not merely document the estimate), since the authoritative value is cheaply recoverable:

  1. supabase/migrations/20260825_03_usage_events_completed_dedup_and_rescale_backfill.sql:
    • merges duplicate 'completed' rows per attempt (keep the earlier/authoritative row, fold the later row's real token counts in, stamp merged ids for forensics)
    • backfills the missed rescale (root cause B) using the exact flag convention 20260823_40 already established (credit_unit: legacy-1usd-100k-credits)
    • adds a partial unique index ux_usage_events_completed_attempt so a future duplicate POST folds instead of inserting
  2. apps/control-plane/internal/usage/repository.go: RecordEvent's INSERT gained the matching ON CONFLICT ... DO UPDATE, deliberately never touching hive_credit_delta/event_type/status. edge-api is unchanged (out of this fix's scope) — its redundant POST now folds harmlessly.
  3. New live guard test apps/control-plane/internal/accounting/usage_ledger_reconciliation_live_test.go: reproduces both writes against a real DB and asserts the surviving row's hive_credit_delta equals the ledger's credits_delta. Verified failing before the fix, passing after.
  4. Two existing test fixtures (repository_live_test.go) inserted multiple 'completed' events against one shared attempt id to simulate multiple requests — a shape that never occurs in production and that the new constraint correctly rejects. Fixed to seed one real attempt per event (what the tests were actually meant to exercise).

Known residual gaps, out of scope (edge-api not in this task's allowlist), noted for the record:

  • edge-api never populates FinalizeReservationInput.InputTokens/OutputTokens despite Console overview and analytics read zero after live chat traffic on the same workspace #856 already wiring the field end to end; this fix's merge compensates without needing it fixed, but it's a one-line edge-api bug worth its own ticket.
  • The reservation.ID == "" fallback path (no reservation ever created) still writes a lone wrong-value row with nothing to reconcile against, since nothing was charged.
  • Relation to Cache token counts never reach api_key_usage_rollups #1174 (cache tokens never reach api_key_usage_rollups): same family of defect, different table/path. The usage_events row this fix produces does carry cache tokens correctly; api_key_usage_rollups is the surface still missing them.

Review round 2 (independent review, DO NOT BLOCK, confirmed by reproduction)

Reviewer confirmed the merge logic, backfill idempotency, and index safety by
reproducing rather than reading, and confirmed edge-api's write path only
logs RecordUsageEvent errors and never propagates them to the HTTP
response, so the migrate-then-deploy window is not customer-facing. One
finding required a fix before merge.

Finding: the backfill under-applied. The first version bounded the
rescale backfill on usage_events.created_at < credit_unit_rescale.applied_at.
20260823_40 documents a deploy-gap race of its own: the old binary keeps
writing old-unit rows for a window after that migration's COMMIT (until the
container recreate), so a straggler row can carry created_at > applied_at
while still holding a stale-unit value. A timestamp bound skips exactly
those rows, permanently — usage_events carries no positive new-unit stamp,
so nothing could ever later tell a missed straggler apart from a
legitimately small delta.

Fix chosen: option 3, reconcile against the ledger. credit_ledger_entries
was already established as authoritative and was correctly backfilled by
20260823_40. The rescale step now reconciles directly against it instead of
inferring the unit from a timestamp: for every surviving 'completed' row,
if its hive_credit_delta disagrees with the matching
credit_ledger_entries.credits_delta (same account, same attempt,
entry_type = 'usage_charge') by exactly a factor of 10000, the ledger's
value is copied onto it. A disagreement that is not exactly 10000x is left
untouched — that would be a different, unexplained mismatch this migration
has no evidence about, and "fixing" it would be a guess. This needed no new
column and no timestamp heuristic (option 1/2 both would have), uses truth
already established, and has no dependency on created_at or
credit_unit_rescale.applied_at at all, so it catches every stale-unit row
including deploy-gap stragglers a timestamp bound structurally cannot see.

Verified by reproduction, not just code-reading: seeded a fabricated
straggler (old-unit usage_events row, created_at 90s after the
rescale marker's applied_at, matching ledger entry in new units) on a
throwaway Postgres 17. Before: hive_credit_delta = -72 vs ledger
-720000, a 10000x miss the old bound would never have touched (its
created_at is after applied_at). After running the revised migration:
hive_credit_delta = -720000, internal_metadata stamped
{"credit_unit": "legacy-1usd-100k-credits", "ledger_reconciled_from": -72}.
Re-ran the file a second time: UPDATE 0 on the reconciliation step, value
unchanged — idempotent.

Second finding: triple-row guard (accepted, cheap guard over a fix). The
dedup merge is a many-to-one join; on an (unobserved, live evidence is 0
triples) 3+-row shape Postgres would pick an arbitrary source row for the
folded token columns. Added a step-0 check that raises and rolls back the
whole transaction if any attempt has 3+ 'completed' rows, before any
mutation. Verified by reproduction: seeded a genuine triple, ran the
migration, got a loud ERROR naming the attempt count and the reason, and
confirmed via a fresh count(*) that all three rows were left untouched
(transaction rolled back, no partial state, index also not created).

Rollback asymmetry (documented in the migration header, not fixed — no
action needed per reviewer):
this migration has no down-migration. Rolling
it back while control-plane's binary stays on the version that assumes
ux_usage_events_completed_attempt exists (the ON CONFLICT target in
usage/repository.go's RecordEvent) breaks every 'completed'
usage_events write — Postgres rejects an ON CONFLICT clause naming a
missing index. That fails loudly as errors, not as silent money corruption,
which is why it's acceptable, but a migration-only rollback must never be
attempted without rolling the binary back first.

Test plan

  • go build ./apps/control-plane/... and go vet on touched packages: clean
  • go test ./apps/control-plane/internal/accounting/... ./apps/control-plane/internal/usage/... ./apps/control-plane/internal/ledger/... against real Postgres 17 (scripts/ci-throwaway-db.sh, 107/107 migrations applied): all green, including the new guard test
  • Guard test verified RED before the fix (fails on the unique-violation / wrong-merge), GREEN after
  • Migration applies cleanly against the full migration chain from scratch
  • Migration idempotency verified by literal replay against fabricated dirty data matching the real live shape (not just code inspection): first run merges + backfills, second run changes 0 rows
  • gofmt -l clean on every touched file
  • Deploy-gap straggler (old-unit row, created_at after the rescale marker) reproduced and confirmed corrected by the ledger reconciliation, not the old timestamp bound
  • Fabricated 3-row shape reproduced; migration raises and rolls back with zero rows mutated

🤖 Generated with Claude Code

…1180)

Two independent bugs made the two money surfaces disagree for the same
request. First, every reservation-backed request wrote TWO usage_events
rows with event_type='completed' for one attempt: control-plane's own
finalizeLocked writes the authoritative row (hive_credit_delta matches
the ledger charge by construction), and edge-api's separate, unconditional
POST to /internal/usage/events writes a second row whose hive_credit_delta
is a raw token count, not a credit figure at all. Confirmed live: 576
duplicate pairs across full history, 576/576 with the same shape. Second,
usage_events.hive_credit_delta was missing from the 20260823_40 credit
unit rescale migration's column list, so every authoritative row written
before that migration is off from the ledger by exactly the 10000x
rescale factor, a stale-unit bug distinct from the duplicate-row bug.

The migration merges duplicate completed rows (keeping the earlier,
authoritative one and folding in the later row's real token counts),
backfills the missed rescale using the same flag convention 20260823_40
already established, and adds a partial unique index so a future
duplicate POST from edge-api folds into the existing row instead of
inserting a second one. usage/repository.go's RecordEvent gained the
matching ON CONFLICT DO UPDATE, deliberately never touching
hive_credit_delta so the ledger-matching value always wins. A new live
guard test proves the two surfaces agree after both writes land; two
existing test fixtures that inserted multiple completed events against
one shared attempt id (a shape that never occurs in production, and that
the new constraint correctly rejects) were fixed to seed one real attempt
per event.

Buglog entry (to be appended to main separately, per repo convention):
{"error_message": "usage_events.hive_credit_delta does not match credit_ledger_entries.credits_delta for the same request", "root_cause": "edge-api writes a redundant, unconditional second usage_events row per completed request with hive_credit_delta set to a raw token count instead of a credit amount, duplicating control-plane's own authoritative write from finalizeLocked; separately, the 20260823_40 credit unit rescale migration omitted usage_events.hive_credit_delta from its column list, leaving pre-rescale rows off by the 10000x factor", "fix": "supabase/migrations/20260825_03_usage_events_completed_dedup_and_rescale_backfill.sql merges duplicate completed rows and backfills the missed rescale; usage/repository.go RecordEvent gained ON CONFLICT DO UPDATE on a new partial unique index that folds a future duplicate write's token counts without touching hive_credit_delta", "tags": ["billing", "ledger", "usage-events", "credit-rescale", "duplicate-write"]}

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@coderabbitai

coderabbitai Bot commented Aug 25, 2026 •

Copy link
Copy Markdown

Warning

Review limit reached

Next included review available in 28 minutes.

View limit details

Limit details: You’ve used the included review currently available.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: c7e15d4f-0b87-46da-9f4f-599094c85fc5

📥 Commits

Reviewing files that changed from the base of the PR and between 4c87c3e and c901724.

📒 Files selected for processing (4)
  • apps/control-plane/internal/accounting/usage_ledger_reconciliation_live_test.go
  • apps/control-plane/internal/usage/repository.go
  • apps/control-plane/internal/usage/repository_live_test.go
  • supabase/migrations/20260825_03_usage_events_completed_dedup_and_rescale_backfill.sql

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Comment thread apps/control-plane/internal/usage/repository.go
Comment thread apps/control-plane/internal/usage/repository.go
@sakibsadmanshajib

Copy link
Copy Markdown
Owner Author

Independent review summary, PR #1194

Read the pushed diff directly (gh pr diff 1194), built and go vet'd the touched Go packages in the toolchain container (clean), applied the full 107-migration chain against a throwaway Postgres 17 (clean, matches the repo's own CI method), and reproduced the migration's dedup/rescale/idempotency behavior against seeded fixtures shaped like the claimed live data (a normal pair, a pair needing rescale, a synthetic triple, and a deploy-gap straggler). Findings below are evidence-based, not speculation, and are posted inline at the exact lines.

Verdict: do not block, ship with two follow-ups tracked before or shortly after merge.

The core diagnosis (two independent root causes, duplicate write plus missed rescale column) is correct and well-evidenced. The merge is safe on the pair shape that is actually live today (576/576, confirmed pattern), the rescale backfill cannot double-apply (confirmed empirically: replay changes 0 rows), and the new unique index does not break the live serving path (confirmed: recordCompletedEvent/recordErrorEvent only log a RecordUsageEvent error, never propagate it into the HTTP response, so even the narrow migrate-then-deploy window is not a customer-facing outage).

Two real, reproduced gaps, both non-blocking for money correctness but worth fixing:

  1. Dedup for 3+ duplicate rows silently drops data (MEDIUM). UPDATE ... FROM _usage_events_completed_dupes is a many-to-one join; Postgres uses exactly one arbitrary matching source row when there are multiple. Reproduced: a synthetic triple folded only one dupe's token counts, dropping the other's, while the forensic merged_duplicate_completed_event_ids metadata correctly lists both ids, so the audit trail overstates what actually merged. hive_credit_delta (money) is untouched by this bug. Zero triples exist in the live sample per the PR, so this is inert today, but the script itself is wrong for that shape.

  2. Rescale backfill cannot catch deploy-gap stragglers (HIGH, but likely small blast radius). The created_at < applied_at bound is the wrong discriminator for exactly the race 20260823_40 documents about itself: an old-unit row can commit after applied_at if the old binary is still serving. That migration's own fix is a flag-only detector with no created_at bound. This migration requires both the flag AND the boundary, so any usage_events straggler from that exact window (2026-08-23, real production event) is skipped, permanently and undetectably, since nothing stamps credit_unit on new-code writes to this table either. Reproduced: a seeded straggler row stayed wrong across two runs of the migration. Worth running the sibling migration's own flag-only detector query against usage_events before merge to see whether this is zero rows (like the dedup case) or a real number.

Also verified, not blocking:

  • Rollback asymmetry: reverting only repository.go is safe; reverting only the migration (dropping the index) while the new code is deployed breaks the ON CONFLICT arbiter for every 'completed'/'reconciled' write, including the authoritative one, though tracing finalizeLocked shows the ledger charge has already posted by that point, so this fails as errors and a usage_events gap, not as money corruption. State this if these two changes can ever land or roll back independently.
  • The ORDER BY created_at ASC keeper-selection has no tiebreaker (low severity, cheap one-line fix: add , id).
  • The residual gaps the PR already discloses (edge-api never populating FinalizeReservationInput.InputTokens/OutputTokens, the reservation.ID == "" fallback path) are correctly scoped out; neither blocks this fix.

Nothing here found a way this migration corrupts the ledger-matching hive_credit_delta value on any pair shape I could construct. The two gaps above cost token-count fidelity and audit completeness, not billing correctness.

sakibsadmanshajib and others added 2 commits August 25, 2026 18:45
… timestamp bound

Review of #1194 reproduced a deploy-gap straggler the created_at < applied_at
bound permanently misses: 20260823_40 itself documents that the old binary
keeps writing old-unit rows for a window after the rescale migration's
COMMIT, so a straggler can carry created_at > applied_at while still holding
a stale-unit value, and usage_events has no positive new-unit stamp that
could later distinguish it from a genuinely small delta.

Replaced the timestamp-bound backfill with reconciliation against
credit_ledger_entries, which 20260823_40 did correctly backfill and which
this migration already treats as authoritative: any surviving 'completed'
row whose hive_credit_delta disagrees with the matching ledger charge by
exactly a factor of 10000 gets the ledger's value copied onto it. No
dependency on created_at or credit_unit_rescale.applied_at, so it catches
every stale-unit row including deploy-gap stragglers, with no heuristic
beyond the exact-factor discriminator that also guards against
"correcting" some other, unrelated disagreement this migration has no
evidence about.

Also added a step-0 guard that RAISEs and rolls back the whole transaction
if any attempt has three or more completed rows, before the dedup UPDATE
runs. The pairwise merge assumes at most two (576 pairs, 0 triples measured
live); a triple would let Postgres pick an arbitrary source row for the
folded token columns, which should stop the migration for a human to look
at rather than guess silently.

Verified against a real Postgres 17: full 107-migration chain still applies
cleanly; a fabricated deploy-gap straggler (old-unit value, created_at after
the rescale marker) is now caught and corrected to the ledger's figure;
rerunning the file a second time changes 0 rows (idempotent); a fabricated
triple raises and rolls back with all three rows left untouched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Review comment: ORDER BY created_at ASC alone leaves the keeper choice
unspecified if two completed rows for one attempt ever share an identical
timestamp. The two writers are always separated by a real HTTP round trip
today, so this has never fired, but the tiebreaker is free and removes the
ambiguity entirely rather than leaving it theoretically open.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@sakibsadmanshajib
sakibsadmanshajib merged commit 84b858b into main Aug 25, 2026
28 checks passed
@sakibsadmanshajib
sakibsadmanshajib deleted the fix/usage-ledger-reconciliation branch August 25, 2026 23:14
sakibsadmanshajib added a commit that referenced this pull request Aug 26, 2026
…rphaned request_attempts, issue #1102) (#1201)

## Summary

PR #1194's migration
`20260825_03_usage_events_completed_dedup_and_rescale_backfill.sql` has
failed on every deploy since it merged. `deploy-demo-box` run
32912034013 (and one run since) hit:

```
psql:.../20260825_03_...sql:227: ERROR:  insert or update on table "usage_events" violates foreign key constraint "usage_events_request_attempt_id_fkey"
DETAIL:  Key (request_attempt_id)=(819cc2ad-d657-4419-9731-349b43675ecf) is not present in table "request_attempts".
```

Nothing has reached the box since, including PR #1193 (Cowork composer
mode).

## Root cause

The live database already holds `usage_events` rows whose
`request_attempt_id` no longer has a matching `request_attempts` row.
This is **issue #1102**, not a new defect: a retention purge deletes
`request_attempts` rows without going through the FK's own `ON DELETE
CASCADE` trigger (it bypasses it rather than the FK being
misconfigured). Live count 2026-08-25: 483 orphaned `usage_events` rows
spanning 2026-04-01 through 2026-08-18 — an ongoing, ordinary state of
this table, not a one-off.

The migration was validated against a throwaway database with none of
these orphans, so the defect never showed up before merge. Step 2 (the
ledger-reconciliation UPDATE) is what trips it: reproducibly isolated
via a rolled-back replay against the live data (confirmed twice,
byte-identical error and row).

## Was the live database left in a bad state?

No. Verified directly against the live box before writing any fix:
- The migration wraps `BEGIN`/`COMMIT` around the whole file and every
failed deploy attempt rolled back cleanly: the 604 duplicate
`'completed'` pairs Step 1 processes were still fully present afterward
(unmerged), and the Step 3 unique index
(`ux_usage_events_completed_attempt`) did not exist.
- The file was never recorded in `public.hive_schema_migrations` (empty
result on every check), so it was safe to amend in place rather than
ship as a new migration.

## Fix

Step 2's `WHERE` clause now requires the row's `request_attempt_id` to
still have a live `request_attempts` parent:

```sql
AND EXISTS (
  SELECT 1 FROM public.request_attempts ra WHERE ra.id = ue.request_attempt_id
)
```

A row with no live parent is skipped, not silently reconciled — there is
nothing left to reconcile it against with confidence either way. Step
1's dedup UPDATE/DELETE needed no change: it was proven safe against the
same live orphaned data in two independent rolled-back replays before
this fix was written.

Fixing the retention purge itself is issue #1102's job, tracked
separately (extend the purge to cascade/null the referencing rows, or an
owner decision to drop constraint enforcement). Out of scope here.

## Verification

- Isolated the failing statement on live production data via rolled-back
transactions (`BEGIN; ...; ROLLBACK;`), never committing a probe.
- Confirmed the fixed Step 2 (with the `EXISTS` guard) completes the
full Step 1 + Step 2 sequence against the real orphaned live data
without error, in a rolled-back replay.
- Added `TestMigrationSurvivesOrphanedRequestAttempt`: builds a real
reservation + duplicate `'completed'` write + stale pre-rescale credit
delta through the actual services, deletes its `request_attempts` row
while bypassing the cascade trigger (reproducing issue #1102's exact
live shape), then executes the real on-disk migration file via the
Postgres simple query protocol (same execution path `psql -f` uses)
against a from-scratch, fully-migrated local Postgres 17 test database.
Asserts no error and that the orphaned row is left untouched, not
reconciled.
- Negative-controlled the new test twice: against the original unpatched
migration it fails, first because the guard predicate text is missing
(static check), and — before that check was tightened to a precise
string — because the orphaned row silently got reconciled instead of
skipped. Restored the fixed file and reran; the full
`internal/accounting` suite (30 tests) and the full `apps/control-plane`
short suite (55 packages) pass.

## Test plan

- [x] `go build ./apps/control-plane/...`
- [x] `go vet ./apps/control-plane/...`
- [x] `go test ./apps/control-plane/internal/accounting/... -v` against
a from-scratch, fully-migrated local Postgres 17 (all 30 tests pass,
including the new one and its #1180 sibling)
- [x] `go test -short ./apps/control-plane/...` (55 packages, all pass)
- [x] Fixed Step 1 + Step 2 sequence replayed against real live orphaned
data on the demo box in a rolled-back transaction, reaches completion
with no FK error
- [x] Negative control: new test fails against the original unpatched
migration file
- [ ] Live deploy: merging should unblock `deploy-demo-box`, confirm the
triggered run succeeds

Buglog entry included in the commit message body per `.wolf/`'s
buglog-lands-on-main convention; will be appended to `main` in a
separate buglog-only PR once this merges.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

usage_events.hive_credit_delta does not match credit_ledger_entries.credits_delta

1 participant