Repository navigation
fix(proxy): requeue daily spend rows when the commit fails without the Redis buffer - #41878
Conversation
…e Redis buffer With the Redis transaction buffer off, each daily spend queue (user, team, org, end user, agent) was drained into a dict and handed to the bulk upsert. When the upsert raised after its retries, the drained dict was discarded and the exception escaped update_spend, so those rows never reached the daily rollup tables and /user/daily/activity stayed short forever while /spend/logs had every request. Each daily queue now flushes through one helper that puts the uncommitted remainder back on the queue for the next tick and moves on to the next table, the same shape the window-spend step already used.
🤖 Devin AI EngineerI'll be helping with this pull request! Here's what you should know: ✅ I will automatically:
Note: I can only respond to comments from users who have write access to this repository. ⚙️ Control Options:
|
|
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
|
bugbot run |
…stead of requeueing them
|
bugbot run |
…e-sent, requeue the unsent ones
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit c6c8aed. Configure here.
TLDR
Problem this solves:
/user/daily/activity,/tag/daily/activity, and the Usage page stay short;/spend/logshas every requestHow it solves it:
Decisions made on the way:
User Flow
Before: an admin reconciling per-key usage on a proxy that runs without Redis finds the daily rollups permanently short of the spend logs after one failed rollup write, for every endpoint
"model": "gpt-5.4-nano", 6 calls per key over a minute; every call returns 200 with a usage blocktotal_costcovering all 6 of its callsbreakdown.api_keys.<hash>.metricsshowsapi_requests: 3and aspendabout half of the report'stotal_costbreakdown.entities.<tag>.metricsshowsapi_requests: 3as wellAfter: the same blip delays the rollups by one write cycle instead of losing them, so the reports agree again on their own
"model": "gpt-5.4-nano", 6 calls per key over a minute; every call returns 200 with a usage blocktotal_costcovering all 6 of its callsbreakdown.api_keys.<hash>.metricsshowsapi_requests: 6and aspendequal to the report'stotal_costbreakdown.entities.<tag>.metricsshowsapi_requests: 6with the samespendRelevant issues
Affected release
Linear ticket
Resolves LIT-6026
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Both legs ran the same way, one after the other on one machine, differing only in the commit the proxies were booted from: two proxy processes per leg, each
--num_workers 2, on two random ports sharing one fresh Postgres database, no Redis (so the in-memory daily queues are the only buffer),proxy_batch_write_at: 10(so the tag rollup job ticks every 23s), one internal user with one virtual key per endpoint, each key carrying one tag, requests alternating across the two instances. Real OpenAI spend ongpt-5.4-nanoSetup, identical for both legs
The flow in order: 2 requests per endpoint, wait 15s, compare (control); take the tables away, 3 requests per endpoint, wait 30s, bring them back (the blip); wait 25s, 1 request per endpoint, wait 30s, compare. Every key ends the run with 6 requests in the spend logs. The
tag_*columns read the key's own tag row (breakdown.entities.<tag>) from the tag rollupBefore (88e150b)
Control: 2 requests per key, both rollups match the spend logs on every endpoint
The
tag_*columns in this control snapshot summed the key's metrics across every tag row of the tag rollup, and each request carries three tags (its key's tag plus the twoUser-Agent: curltags the proxy adds), hence exactly 3x: 6 tag requests for 2 calls,tag_spend= 3 xrollup_spend. The join was switched to the key's own tag row before the final comparison below and stayed that way for the after legAfter the blip: both rollups stay at 3 of 6 per key on every endpoint, the 3 calls made during the blip never land
This comparison is the read taken after the flow ended (the same three GETs, nothing else sent, the join on the key's own tag row); the read right after the flow's last request showed the same totals, so the rows are gone for good
Before leg, every request and response as returned (ports 57321 and 54801)
/v1/responsesand/v1/messagesids are shortened to their first 32 characters here; everything else is verbatimAfter (c6c8aed)
Control: 2 requests per key, the daily rollup matches the spend logs on every endpoint
The control read lands 15s after the last request and the tag job ticks every 23s, so the tag job had not run once at that read (0 tag requests). Every call is there in the final read below (6 of 6 per key on both rollups); nothing was taken away during the control phase
After the blip: both rollups reach 6 of 6 per key on every endpoint, the 3 calls made during the blip land once the tables are back
The daily rollup total (0.000184, 18 requests) equals the spend log total, and the tag rollup total is exactly 3x it (54 = 18 requests x 3 tags per request), so nothing was lost and nothing was counted twice. The proxy logs carry 12
Re-queued N rowslines across the two instances during the blip (a missing table is Postgres error 42P01, a failure that proves the write never applied, so it requeues) and nodroppedline. The same run at the earlier tips 3edbf60 and b3cf45e ended 6 of 6 per key as wellAfter leg, every request and response as returned (ports 45152 and 47423)
/v1/responsesand/v1/messagesids are shortened to their first 32 characters here; everything else is verbatimObservations from the run
User-Agenttags included; pre-existing, PR leaves it alone/v1/messageson an OpenAI model returns aresp_id; pre-existing, PR leaves it aloneVerdict: PASS (before=fail at 88e150b, after=succeeds at c6c8aed)
/live-pr-risk at c6c8aed: a daily rollup row Postgres refuses for good, without and with the Redis buffer
Scenario: a
NOT VALIDCHECK constraint onLiteLLM_DailyUserSpend.api_keymakes the chat key's daily rollup row fail with SQLSTATE 23514 on every upsert (the shape of a NUL byte in an id or any other permanent rejection) while the responses and messages keys stay writable. One chat and one responses call in the first write cycle, one messages call in the next, a read 40s later with the constraint still on, the constraint dropped, one more chat call, a final read 15s later. Same two-instance x two-worker topology as the QA legsNo Redis buffer, head (the after leg's instances :45152/:47423, right after the run above)
One
droppedline, no re-send on later cycles, the other keys and the tag rollup untouched, and the chat key's next call lands. The merge base drops the failed batch on every failure on this path, so it is a superset of this behavior and was not re-run for this classRedis transaction buffer on (
general_settings.use_redis_transaction_buffer: true, both instances pushing to one Redis), merge base 88e150b (:48967/:45189) against head c6c8aed (:48547/:52324), same scenario at the same timeWhat changed on the Redis-buffered path: the merge base re-sends the whole uncommitted set every cycle, so while one row is refused no daily user rollup row lands for any key (responses and messages sat at 2 above until the constraint was dropped), and a rejection that never clears freezes the Usage page for every key until a restart drops the buffer; the head drops the refused row's batch once (one row here: the two instances' cycles put the chat and responses rows in different batches) and every other row keeps landing. Same rule as the no-Redis path, same trade: the refused batch is gone from the rollups and named in one error line, and the spend logs still carry it. The per-entity spend increments and the window path on the Redis-buffered side were not changed, and the tag rollup kept every request on both sides (
tag_requestsabove, on its own 23s cadence)Type
🐛 Bug Fix
Caveats (if any)
Low
droppedline from the commit is that tick's only record. A second line there would repeat the firsttest_bedrock_invoke_messages_with_all_beta_headersfor both Bedrock params (main pipeline 89818 job 2188120; red since fix(anthropic): register thinking-binding-controls-2026-08-01 in beta headers config #41203 registeredthinking-binding-controls-2026-08-01as a Bedrock pass-through, tracked in LIT-8150),test_models_by_provider(main pipeline 89818 job 2188086; red since feat(proxy): add Amazon Transcribe pass-through with completion-time job pricing #41515 added thetranscribecost map entry without registering the provider), and the two integration retry tests (main pipelines 89785 jobs 2186412 and 2186414 and 89818 jobs 2188082 and 2188083, whose fix fix(proxy): refuse config-owned keys on POST /config/update #41868 reached main at 00:23Z on 2026-09-19, after main's last scheduled run). All three breaks landed on main before this branch's merge base 88e150b, and merging main in would clear only the integration pair, so the branch stays on its merge baseFinal Attestation
Note
Medium Risk
Changes billing/usage rollup persistence on the no-Redis path; failures are requeued (with a known ambiguous post-send double-count tradeoff) instead of dropped, which affects Usage and daily activity APIs.
Overview
Fixes permanent loss of daily rollup spend when the proxy commits without the Redis transaction buffer and a database write fails. Previously, flushed in-memory rows were dropped and the error could stop the rest of the tick’s daily flushes.
All daily entity queues (user, team, org, end user, agent, and tag via its separate job) now go through
_flush_daily_spend_queue: flush aggregated transactions, call the existing per-table commit, and on failure log, re-queue the batch withadd_update, and continue—matching the window-spend error handling already in the same path. Redis-buffered commit flows are unchanged.Adds regression tests that a failed user daily upsert requeues while team still commits, and that failed tag commits requeue for the next scheduler tick.
Reviewed by Cursor Bugbot for commit 3edbf60. Bugbot is set up for automated code reviews on this repo. Configure here.