Skip to content

fix(proxy): requeue daily spend rows when the commit fails without the Redis buffer - #41878

Merged
mateo-berri merged 5 commits into
mainfrom
litellm_requeue_daily_spend_without_redis_buffer
Sep 19, 2026
Merged

mateo-berri merged 5 commits into
mainfrom
litellm_requeue_daily_spend_without_redis_buffer

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Without the Redis buffer, a failed daily rollup write loses its rows for good
  • /user/daily/activity, /tag/daily/activity, and the Usage page stay short; /spend/logs has every request
  • The error also escaped the tick, so the other rollup tables waited a tick

How it solves it:

  • Each daily queue (user, team, org, end user, agent, tag) flushes through one helper
  • On a failed commit the uncommitted rows go back on the queue for the next tick, when re-sending them is safe: a refused connection, a missing table, a deadlock, a cancelled statement
  • A reply lost after the statement was sent, or data Postgres refuses outright, drops the batch of up to 100 rows that statement carried, with an error log naming the row count, as before this PR; the batches behind it that were never sent go back on the queue
  • The next table still flushes, the same shape the window-spend step already had
  • The tag rollup's own scheduler job flushes through the same helper
  • Regression tests: user upsert fails, team still lands, user lands on the retry; tag upsert fails, tag lands on the retry; each failure class lands on the next tick or not exactly as the rule says; a reply lost on the first of two batches drops that batch alone and the second lands next tick; the Postgres error code is read from every shape prisma hands back

Decisions made on the way:

  • Requeue through the queue's own add path, so a full queue still aggregates
  • The rows handed back are exactly the ones the upsert did not commit
  • One generic helper with one cast instead of six per-table copies
  • The tag rollup is in scope: same drop on the same tick shape, and its Redis twin already requeues
  • The drop happens inside the commit, at the batch whose statement failed, not in the flush helper: the commit sends the rows 100 at a time and takes each batch out of the dict only after its statement succeeds, so a drop decided in the flush would have thrown away the batches that were never sent along with the one that failed (raised by bugbot). The alternatives were a wrapper exception carrying the failed batch out to the flush, or the flush recomputing which batch was on the wire, both a second copy of the batching. One consequence: the Redis-buffered path restores the same dicts on failure, so its daily rollups now also drop the un-resendable batch instead of re-sending it; the window path and the per-entity spend increments on that path are as they were
  • A failure that arrives after the statement was sent (a read timeout on the reply) is dropped, not requeued: the statement may have committed, and sending it again adds the spend a second time, the double count LIT-4823 fixed by retrying only a refused connection. The first cut requeued everything the way the Redis-buffered and window paths do and was reversed for that reason; an idempotent upsert keyed by batch would keep both but needs a new column and a migration
  • A connect timeout is dropped too, since only a refused connection proves the statement never left; that follows the retry-safe set LIT-4823 defined rather than widening it here
  • Data Postgres refuses outright (a NUL byte in a user id or a tag, a constraint violation: error class 22 or 23) is dropped rather than requeued, since it would fail the same way every tick and hold up that table's rollup until a restart; the rest of that failed batch goes with it, which is what happened before this PR on every failure. Bisecting the batch to isolate the bad row, as the spend-log writer does, is a follow-up
  • The Postgres error code is read from the prisma error's meta through a pydantic TypeAdapter rather than a raw dict access; both a missing table and a rejected row surface as the same prisma error type, so the type alone cannot tell them apart
  • Shutdown flush stays out of scope; PR fix(proxy): flush in-memory spend buffers before shutdown #34808 covers that separately

User Flow

Before: an admin reconciling per-key usage on a proxy that runs without Redis finds the daily rollups permanently short of the spend logs after one failed rollup write, for every endpoint

  1. Developers send requests through their virtual keys, each key tagged with its team's tag, across the proxy's replicas: POST https://litellm-domain/v1/chat/completions, POST https://litellm-domain/v1/responses, and POST https://litellm-domain/v1/messages, each with "model": "gpt-5.4-nano", 6 calls per key over a minute; every call returns 200 with a usage block
  2. During that minute the database refuses one batch write of the daily rollups (a dropped connection, a timeout, a table briefly unavailable), the kind of blip the proxy retries for a few seconds and then gives up on
  3. The admin opens GET https://litellm-domain/global/spend/report?start_date=2026-09-18&end_date=2026-09-18&group_by=api_key and sees each of the 3 keys with a total_cost covering all 6 of its calls
  4. The admin opens GET https://litellm-domain/user/daily/activity/aggregated?start_date=2026-09-18&end_date=2026-09-18 and each key's breakdown.api_keys.<hash>.metrics shows api_requests: 3 and a spend about half of the report's total_cost
  5. The admin opens GET https://litellm-domain/tag/daily/activity?start_date=2026-09-18&end_date=2026-09-18 and each key's tag under breakdown.entities.<tag>.metrics shows api_requests: 3 as well
  6. They repeat steps 4 and 5 an hour later and both rollups still show 3 of 6 per key: the calls refused during the blip never land, and the Usage page, which reads the same rollups, stays short for the day

After: the same blip delays the rollups by one write cycle instead of losing them, so the reports agree again on their own

  1. Developers send requests through their virtual keys, each key tagged with its team's tag, across the proxy's replicas: POST https://litellm-domain/v1/chat/completions, POST https://litellm-domain/v1/responses, and POST https://litellm-domain/v1/messages, each with "model": "gpt-5.4-nano", 6 calls per key over a minute; every call returns 200 with a usage block
  2. During that minute the database refuses one batch write of the daily rollups (a refused connection, a table briefly unavailable, a deadlock: a failure that shows the write never applied), the kind of blip the proxy retries for a few seconds and then gives up on
  3. The admin opens GET https://litellm-domain/global/spend/report?start_date=2026-09-18&end_date=2026-09-18&group_by=api_key and sees each of the 3 keys with a total_cost covering all 6 of its calls
  4. The admin opens GET https://litellm-domain/user/daily/activity/aggregated?start_date=2026-09-18&end_date=2026-09-18 and each key's breakdown.api_keys.<hash>.metrics shows api_requests: 6 and a spend equal to the report's total_cost
  5. The admin opens GET https://litellm-domain/tag/daily/activity?start_date=2026-09-18&end_date=2026-09-18 and each key's tag under breakdown.entities.<tag>.metrics shows api_requests: 6 with the same spend
  6. The calls refused during the blip landed on the first write cycle after the database came back, so the Usage page shows the full day too. When a reply is lost mid-way through a tick that holds more than 100 rollup rows, only the rows on the wire at that moment are missing from the rollups; the rest land on the next cycle

Relevant issues

Affected release

Linear ticket

Resolves LIT-6026

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Both legs ran the same way, one after the other on one machine, differing only in the commit the proxies were booted from: two proxy processes per leg, each --num_workers 2, on two random ports sharing one fresh Postgres database, no Redis (so the in-memory daily queues are the only buffer), proxy_batch_write_at: 10 (so the tag rollup job ticks every 23s), one internal user with one virtual key per endpoint, each key carrying one tag, requests alternating across the two instances. Real OpenAI spend on gpt-5.4-nano

Setup, identical for both legs

# config: one model, gpt-5.4-nano -> openai/gpt-5.4-nano, general_settings.proxy_batch_write_at: 10, master_key + database_url from env
python litellm/proxy/proxy_cli.py --config qa_config.yaml --port $PA --num_workers 2 --use_v2_migration_resolver   # instance A
python litellm/proxy/proxy_cli.py --config qa_config.yaml --port $PB --num_workers 2 --use_v2_migration_resolver   # instance B

curl -s -X POST "http://127.0.0.1:$PA/user/new" -H "Authorization: Bearer $MASTER" -H 'Content-Type: application/json' -d '{"user_id":"lit6026-qa-user","user_email":"lit6026-qa@example.com"}'
for ep in chat responses messages; do
  curl -s -X POST "http://127.0.0.1:$PB/key/generate" -H "Authorization: Bearer $MASTER" -H 'Content-Type: application/json' -d "{\"user_id\":\"lit6026-qa-user\",\"key_alias\":\"lit6026-$ep\",\"metadata\":{\"tags\":[\"lit6026-$ep\"]}}"
done

# one request per endpoint (the run alternates $PA and $PB per call)
curl -s -X POST "http://127.0.0.1:$PA/v1/chat/completions" -H "Authorization: Bearer $KEY_CHAT" -H 'Content-Type: application/json' -d '{"model":"gpt-5.4-nano","messages":[{"role":"user","content":"reply with one word"}]}'
curl -s -X POST "http://127.0.0.1:$PB/v1/responses" -H "Authorization: Bearer $KEY_RESPONSES" -H 'Content-Type: application/json' -d '{"model":"gpt-5.4-nano","input":"reply with one word"}'
curl -s -X POST "http://127.0.0.1:$PA/v1/messages" -H "Authorization: Bearer $KEY_MESSAGES" -H 'Content-Type: application/json' -d '{"model":"gpt-5.4-nano","max_tokens":256,"messages":[{"role":"user","content":"reply with one word"}]}'

# the reports the customer compared, joined per key on the key hash (and per tag for the tag rollup)
curl -s "http://127.0.0.1:$PA/user/daily/activity/aggregated?start_date=2026-09-18&end_date=2026-09-18" -H "Authorization: Bearer $MASTER"
curl -s "http://127.0.0.1:$PB/global/spend/report?start_date=2026-09-18&end_date=2026-09-18&group_by=api_key" -H "Authorization: Bearer $MASTER"
curl -s "http://127.0.0.1:$PA/tag/daily/activity?start_date=2026-09-18&end_date=2026-09-18" -H "Authorization: Bearer $MASTER"

# the blip: the daily user and tag rollup tables are unavailable for about one write cycle window, then come back
psql "$DATABASE_URL" -c 'ALTER TABLE "LiteLLM_DailyUserSpend" RENAME TO "LiteLLM_DailyUserSpend_offline"' -c 'ALTER TABLE "LiteLLM_DailyTagSpend" RENAME TO "LiteLLM_DailyTagSpend_offline"'
# ... 3 requests per endpoint, sleep 30 ...
psql "$DATABASE_URL" -c 'ALTER TABLE "LiteLLM_DailyUserSpend_offline" RENAME TO "LiteLLM_DailyUserSpend"' -c 'ALTER TABLE "LiteLLM_DailyTagSpend_offline" RENAME TO "LiteLLM_DailyTagSpend"'

The flow in order: 2 requests per endpoint, wait 15s, compare (control); take the tables away, 3 requests per endpoint, wait 30s, bring them back (the blip); wait 25s, 1 request per endpoint, wait 30s, compare. Every key ends the run with 6 requests in the spend logs. The tag_* columns read the key's own tag row (breakdown.entities.<tag>) from the tag rollup

Before (88e150b)

Control: 2 requests per key, both rollups match the spend logs on every endpoint

-- GET /user/daily/activity/aggregated?start_date=2026-09-18&end_date=2026-09-18  (daily rollup; what the Usage page reads)
{"total_spend":0.00005925,"total_api_requests":6,"total_tokens":123}
-- GET /global/spend/report?start_date=2026-09-18&end_date=2026-09-18&group_by=api_key  (raw spend logs)
{"total_spend":5.9250000000000004e-05,"keys":3}
-- GET /tag/daily/activity?start_date=2026-09-18&end_date=2026-09-18  (daily tag rollup; what the Tags usage panel reads)
{"total_spend":0.00017774999999999998,"total_api_requests":18}
-- per key: daily rollup vs spend logs vs daily tag rollup (each key carries one tag)
{"endpoint":"chat","key":"5f390022cd","rollup_requests":2,"rollup_spend":0.000017250000000000003,"spendlog_spend":0.000017250000000000003,"match":true,"tag_requests":6,"tag_spend":0.00005175,"tag_match":false}
{"endpoint":"responses","key":"0f2d109305","rollup_requests":2,"rollup_spend":0.00002225,"spendlog_spend":0.00002225,"match":true,"tag_requests":6,"tag_spend":0.00006675,"tag_match":false}
{"endpoint":"messages","key":"1dea036373","rollup_requests":2,"rollup_spend":0.00001975,"spendlog_spend":0.00001975,"match":true,"tag_requests":6,"tag_spend":0.00005925,"tag_match":false}

The tag_* columns in this control snapshot summed the key's metrics across every tag row of the tag rollup, and each request carries three tags (its key's tag plus the two User-Agent: curl tags the proxy adds), hence exactly 3x: 6 tag requests for 2 calls, tag_spend = 3 x rollup_spend. The join was switched to the key's own tag row before the final comparison below and stayed that way for the after leg

After the blip: both rollups stay at 3 of 6 per key on every endpoint, the 3 calls made during the blip never land

rollup tables taken away at 21:36:37Z
... 3 x /v1/chat/completions, 3 x /v1/responses, 3 x /v1/messages, all 200 ...
rollup tables back at 21:37:59Z
... 1 x /v1/chat/completions, 1 x /v1/responses, 1 x /v1/messages, all 200 ...
-- GET /user/daily/activity/aggregated?start_date=2026-09-18&end_date=2026-09-18  (daily rollup; what the Usage page reads)
{"total_spend":0.00009325000000000002,"total_api_requests":9,"total_tokens":188}
-- GET /global/spend/report?start_date=2026-09-18&end_date=2026-09-18&group_by=api_key  (raw spend logs)
{"total_spend":0.00018525,"keys":3}
-- GET /tag/daily/activity?start_date=2026-09-18&end_date=2026-09-18  (daily tag rollup; what the Tags usage panel reads)
{"total_spend":0.00027975,"total_api_requests":27}
-- per key: daily rollup vs spend logs vs daily tag rollup (each key carries the tag lit6026-before-<endpoint>; the tag rollup row is read by that tag)
{"endpoint":"chat","key":"5f390022cd","rollup_requests":3,"rollup_spend":0.000027750000000000004,"spendlog_spend":0.0000555,"match":false,"tag_requests":3,"tag_spend":0.00002775,"tag_match":false}
{"endpoint":"responses","key":"0f2d109305","rollup_requests":3,"rollup_spend":0.000034,"spendlog_spend":0.00006925,"match":false,"tag_requests":3,"tag_spend":0.000034,"tag_match":false}
{"endpoint":"messages","key":"1dea036373","rollup_requests":3,"rollup_spend":0.0000315,"spendlog_spend":0.0000605,"match":false,"tag_requests":3,"tag_spend":0.0000315,"tag_match":false}

This comparison is the read taken after the flow ended (the same three GETs, nothing else sent, the join on the key's own tag row); the read right after the flow's last request showed the same totals, so the rows are gone for good

Before leg, every request and response as returned (ports 57321 and 54801)
### 0. one internal user with one virtual key per endpoint (POST /user/new, POST /key/generate x3)
{"user_id":"lit6026-qa-user","key":"sk-HhG..."}
[{"endpoint":"chat","key":"sk-QjN...","hash":"5f390022cd"},{"endpoint":"responses","key":"sk-DHD...","hash":"0f2d109305"},{"endpoint":"messages","key":"sk-d8o...","hash":"1dea036373"}]

### 1. control: 2 requests per endpoint, alternating across the two proxy instances (:57321, :54801), then wait past the 10s batch write
{"port":57321,"route":"/v1/chat/completions","id":"chatcmpl-EPaXDHvktlSYgHhXZjBVOSaep6wpG","total_tokens":19}
{"port":54801,"route":"/v1/chat/completions","id":"chatcmpl-EPaXLvdlIGluoNNqOE341XYAbmRl5","total_tokens":20}
{"port":57321,"route":"/v1/responses","id":"resp_h6X3nRGQFSjBy0HqaXlgwOUn4VHORGXG...","total_tokens":21}
{"port":54801,"route":"/v1/responses","id":"resp_5XZWrXCAAJbd4wwTFbLvZBvpTUzozAZP...","total_tokens":22}
{"port":57321,"route":"/v1/messages","id":"resp_bGl0ZWxsbTpjdXN0b21fbGxtX3Byb3Zp...","total_tokens":21}
{"port":54801,"route":"/v1/messages","id":"resp_bGl0ZWxsbTpjdXN0b21fbGxtX3Byb3Zp...","total_tokens":20}

### 2. fault: the daily user and tag rollup tables are unavailable for one write cycle window (~45s; the tag job ticks every 23s); 3 requests per endpoint meanwhile
rollup tables taken away at 21:36:37Z
{"port":57321,"route":"/v1/chat/completions","id":"chatcmpl-EPaYEO9PfmakJAfMb1ZfvUdHoLiPH","total_tokens":20}
{"port":54801,"route":"/v1/chat/completions","id":"chatcmpl-EPaYRCjMiyZMeW1Ng300c1mUAXbho","total_tokens":19}
{"port":57321,"route":"/v1/chat/completions","id":"chatcmpl-EPaYTzr6rKQrVinp72nEhPpBplRFo","total_tokens":21}
{"port":54801,"route":"/v1/responses","id":"resp_y6HlFOTaLEjkgY7lCRSDXdezWjBwC946...","total_tokens":22}
{"port":57321,"route":"/v1/responses","id":"resp_ChRjJh7wBSTrv11s4ee0sVWIu1Bkwarn...","total_tokens":22}
{"port":54801,"route":"/v1/responses","id":"resp_e21_p98c0QZyGJ7cyA0omsc2zgrZfMq1...","total_tokens":22}
{"port":57321,"route":"/v1/messages","id":"resp_bGl0ZWxsbTpjdXN0b21fbGxtX3Byb3Zp...","total_tokens":21}
{"port":54801,"route":"/v1/messages","id":"resp_bGl0ZWxsbTpjdXN0b21fbGxtX3Byb3Zp...","total_tokens":20}
{"port":57321,"route":"/v1/messages","id":"resp_bGl0ZWxsbTpjdXN0b21fbGxtX3Byb3Zp...","total_tokens":20}
rollup tables back at 21:37:59Z

### 3. after: tables back, wait 2 write cycles, 1 more request per endpoint, wait past the next tag tick, compare
{"port":54801,"route":"/v1/chat/completions","id":"chatcmpl-EPaZpcWUHgZGMFy2MiIbrfiS2j3j1","total_tokens":21}
{"port":57321,"route":"/v1/responses","id":"resp_vMhwGDleVxyJKVd3QplZDW75dyjt9eU_...","total_tokens":22}
{"port":54801,"route":"/v1/messages","id":"resp_bGl0ZWxsbTpjdXN0b21fbGxtX3Byb3Zp...","total_tokens":22}

/v1/responses and /v1/messages ids are shortened to their first 32 characters here; everything else is verbatim

After (c6c8aed)

Control: 2 requests per key, the daily rollup matches the spend logs on every endpoint

-- GET /user/daily/activity/aggregated?start_date=2026-09-18&end_date=2026-09-18  (daily rollup; what the Usage page reads)
{"total_spend":0.000058,"total_api_requests":6,"total_tokens":122}
-- GET /global/spend/report?start_date=2026-09-18&end_date=2026-09-18&group_by=api_key  (raw spend logs)
{"total_spend":5.8e-05,"keys":3}
-- GET /tag/daily/activity?start_date=2026-09-18&end_date=2026-09-18  (daily tag rollup; what the Tags usage panel reads)
{"total_spend":0.0,"total_api_requests":0}
-- per key: daily rollup vs spend logs vs daily tag rollup (each key carries the tag lit6026-after3-<endpoint>; the tag rollup row is read by that tag)
{"endpoint":"chat","key":"0265c0c6d5","rollup_requests":2,"rollup_spend":0.000017250000000000003,"spendlog_spend":0.000017250000000000003,"match":true,"tag_requests":0,"tag_spend":0,"tag_match":false}
{"endpoint":"responses","key":"8737b340f6","rollup_requests":2,"rollup_spend":0.000021,"spendlog_spend":0.000021,"match":true,"tag_requests":0,"tag_spend":0,"tag_match":false}
{"endpoint":"messages","key":"70eacf096f","rollup_requests":2,"rollup_spend":0.00001975,"spendlog_spend":0.00001975,"match":true,"tag_requests":0,"tag_spend":0,"tag_match":false}

The control read lands 15s after the last request and the tag job ticks every 23s, so the tag job had not run once at that read (0 tag requests). Every call is there in the final read below (6 of 6 per key on both rollups); nothing was taken away during the control phase

After the blip: both rollups reach 6 of 6 per key on every endpoint, the 3 calls made during the blip land once the tables are back

rollup tables taken away at 23:26:07Z
... 3 x /v1/chat/completions, 3 x /v1/responses, 3 x /v1/messages, all 200 ...
rollup tables back at 23:26:44Z
... 1 x /v1/chat/completions, 1 x /v1/responses, 1 x /v1/messages, all 200 ...
-- GET /user/daily/activity/aggregated?start_date=2026-09-18&end_date=2026-09-18  (daily rollup; what the Usage page reads)
{"total_spend":0.00018399999999999997,"total_api_requests":18,"total_tokens":374}
-- GET /global/spend/report?start_date=2026-09-18&end_date=2026-09-18&group_by=api_key  (raw spend logs)
{"total_spend":0.00018399999999999997,"keys":3}
-- GET /tag/daily/activity?start_date=2026-09-18&end_date=2026-09-18  (daily tag rollup; what the Tags usage panel reads)
{"total_spend":0.0005520000000000001,"total_api_requests":54}
-- per key: daily rollup vs spend logs vs daily tag rollup (each key carries the tag lit6026-after3-<endpoint>; the tag rollup row is read by that tag)
{"endpoint":"chat","key":"0265c0c6d5","rollup_requests":6,"rollup_spend":0.000056750000000000004,"spendlog_spend":0.00005675,"match":true,"tag_requests":6,"tag_spend":0.00005675,"tag_match":true}
{"endpoint":"responses","key":"8737b340f6","rollup_requests":6,"rollup_spend":0.00006424999999999999,"spendlog_spend":0.00006424999999999999,"match":true,"tag_requests":6,"tag_spend":0.00006425,"tag_match":true}
{"endpoint":"messages","key":"70eacf096f","rollup_requests":6,"rollup_spend":0.000063,"spendlog_spend":0.000063,"match":true,"tag_requests":6,"tag_spend":0.000063,"tag_match":true}

The daily rollup total (0.000184, 18 requests) equals the spend log total, and the tag rollup total is exactly 3x it (54 = 18 requests x 3 tags per request), so nothing was lost and nothing was counted twice. The proxy logs carry 12 Re-queued N rows lines across the two instances during the blip (a missing table is Postgres error 42P01, a failure that proves the write never applied, so it requeues) and no dropped line. The same run at the earlier tips 3edbf60 and b3cf45e ended 6 of 6 per key as well

After leg, every request and response as returned (ports 45152 and 47423)
### 0. one internal user with one virtual key per endpoint (POST /user/new, POST /key/generate x3)
{"user_id":"lit6026-qa-user","key":"sk-faA..."}
[{"endpoint":"chat","key":"sk-fmm...","hash":"0265c0c6d5"},{"endpoint":"responses","key":"sk-SLT...","hash":"8737b340f6"},{"endpoint":"messages","key":"sk-YxP...","hash":"70eacf096f"}]
### 1. control: 2 requests per endpoint, alternating across the two proxy instances (:45152, :47423), then wait past the 10s batch write
{"port":45152,"route":"/v1/chat/completions","id":"chatcmpl-EPcFfommpszi8OYR5Mcy5tv...","total_tokens":19}
{"port":47423,"route":"/v1/chat/completions","id":"chatcmpl-EPcFfYLdDFAonDsv7Xr5ayu...","total_tokens":20}
{"port":45152,"route":"/v1/responses","id":"resp_KTQKrmlaWyXX_4LnaHeOl1DfnyT...","total_tokens":21}
{"port":47423,"route":"/v1/responses","id":"resp_lI7H5vJdQ2BRpw1tp9IcYDvek1v...","total_tokens":21}
{"port":45152,"route":"/v1/messages","id":"resp_bGl0ZWxsbTpjdXN0b21fbGxtX3B...","total_tokens":21}
{"port":47423,"route":"/v1/messages","id":"resp_bGl0ZWxsbTpjdXN0b21fbGxtX3B...","total_tokens":20}
### 2. fault: the daily user and tag rollup tables are unavailable for one write cycle window (~45s; the tag job ticks every 23s); 3 requests per endpoint meanwhile
rollup tables taken away at 23:26:07Z
{"port":45152,"route":"/v1/chat/completions","id":"chatcmpl-EPcFzlohwjn8XoQCOARygzj...","total_tokens":21}
{"port":47423,"route":"/v1/chat/completions","id":"chatcmpl-EPcFzQsJLD34PXTlxKvkcwh...","total_tokens":20}
{"port":45152,"route":"/v1/chat/completions","id":"chatcmpl-EPcG0NPbG5oJT1FlEv6pGZ2...","total_tokens":21}
{"port":47423,"route":"/v1/responses","id":"resp_Qf7cN8GvHlnYS5-Jcxg776Az-Bz...","total_tokens":21}
{"port":45152,"route":"/v1/responses","id":"resp_uEN9SF0v7ClVLcQ3K28tFRzvzYJ...","total_tokens":21}
{"port":47423,"route":"/v1/responses","id":"resp__Jx9-cFFZkuNH6SoR1_3jIM3QfO...","total_tokens":22}
{"port":45152,"route":"/v1/messages","id":"resp_bGl0ZWxsbTpjdXN0b21fbGxtX3B...","total_tokens":21}
{"port":47423,"route":"/v1/messages","id":"resp_bGl0ZWxsbTpjdXN0b21fbGxtX3B...","total_tokens":20}
{"port":45152,"route":"/v1/messages","id":"resp_bGl0ZWxsbTpjdXN0b21fbGxtX3B...","total_tokens":22}
rollup tables back at 23:26:44Z
### 3. after: tables back, wait 2 write cycles, 1 more request per endpoint, wait past the next tag tick, compare
{"port":47423,"route":"/v1/chat/completions","id":"chatcmpl-EPcH029ViRHZUzO5vuyNqpn...","total_tokens":20}
{"port":45152,"route":"/v1/responses","id":"resp_8_UoTKIesXxDw0r6CjUp_WN0KKB...","total_tokens":21}
{"port":47423,"route":"/v1/messages","id":"resp_bGl0ZWxsbTpjdXN0b21fbGxtX3B...","total_tokens":22}

/v1/responses and /v1/messages ids are shortened to their first 32 characters here; everything else is verbatim

Observations from the run

  • Before: every endpoint loses exactly the blip-window calls, on both rollups; the PR fixes this
  • After: rows written through instance A were read back through instance B
  • Tag rollup counts each request once per tag, User-Agent tags included; pre-existing, PR leaves it alone
  • Tag rollup lags the user rollup by up to 23s; pre-existing, PR leaves it alone
  • /v1/messages on an OpenAI model returns a resp_ id; pre-existing, PR leaves it alone

Verdict: PASS (before=fail at 88e150b, after=succeeds at c6c8aed)

/live-pr-risk at c6c8aed: a daily rollup row Postgres refuses for good, without and with the Redis buffer

Scenario: a NOT VALID CHECK constraint on LiteLLM_DailyUserSpend.api_key makes the chat key's daily rollup row fail with SQLSTATE 23514 on every upsert (the shape of a NUL byte in an id or any other permanent rejection) while the responses and messages keys stay writable. One chat and one responses call in the first write cycle, one messages call in the next, a read 40s later with the constraint still on, the constraint dropped, one more chat call, a final read 15s later. Same two-instance x two-worker topology as the QA legs

No Redis buffer, head (the after leg's instances :45152/:47423, right after the run above)

constraint on at 23:30:26Z
23:30:33Z instance a: Spend tracking - dropped 1 daily user spend rows: the failed statement may have applied or the database refused the data, so re-sending it is not safe. Table: LiteLLM_DailyUserSpend, Error: ERROR: new row for relation "LiteLLM_DailyUserSpend" violates check constraint "lit6026_reject" ...
-- mid-window read at 23:31:07Z, constraint still on (rollup requests vs spend, per key)
{"endpoint":"chat","rollup_requests":6,"tag_requests":7,"rollup_spend":0.00005675,"spendlog_spend":0.00006475,"match":false}
{"endpoint":"responses","rollup_requests":7,"tag_requests":7,"rollup_spend":0.000076,"spendlog_spend":0.000076,"match":true}
{"endpoint":"messages","rollup_requests":7,"tag_requests":7,"rollup_spend":0.0000735,"spendlog_spend":0.0000735,"match":true}
constraint off at 23:31:07Z
-- final read at 23:31:23Z
{"endpoint":"chat","rollup_requests":7,"tag_requests":7,"rollup_spend":0.00006705,"spendlog_spend":0.00007505,"match":false}
{"endpoint":"responses","rollup_requests":7,"tag_requests":7,"rollup_spend":0.000076,"spendlog_spend":0.000076,"match":true}
{"endpoint":"messages","rollup_requests":7,"tag_requests":7,"rollup_spend":0.0000735,"spendlog_spend":0.0000735,"match":true}

One dropped line, no re-send on later cycles, the other keys and the tag rollup untouched, and the chat key's next call lands. The merge base drops the failed batch on every failure on this path, so it is a superset of this behavior and was not re-run for this class

Redis transaction buffer on (general_settings.use_redis_transaction_buffer: true, both instances pushing to one Redis), merge base 88e150b (:48967/:45189) against head c6c8aed (:48547/:52324), same scenario at the same time

-- merge base: 12 "failed to commit spend updates from Redis to DB. Re-queuing uncommitted transactions to Redis for retry on next tick" lines in the 42s the constraint was on, no dropped line
-- merge base, mid-window read at 23:44:47Z: no key moves, the refused row holds every daily user rollup row in the buffer
{"endpoint":"chat","rollup_requests":2,"rollup_spend":0.0000185,"spendlog_spend":0.00002775,"match":false}
{"endpoint":"responses","rollup_requests":2,"rollup_spend":0.00002225,"spendlog_spend":0.000034,"match":false}
{"endpoint":"messages","rollup_requests":2,"rollup_spend":0.00002225,"spendlog_spend":0.00003275,"match":false}
-- merge base, final read at 23:45:03Z, 16s after the constraint was dropped: everything lands
{"endpoint":"chat","rollup_requests":4,"rollup_spend":0.0000368,"spendlog_spend":0.0000368,"match":true}
{"endpoint":"responses","rollup_requests":3,"rollup_spend":0.000034,"spendlog_spend":0.000034,"match":true}
{"endpoint":"messages","rollup_requests":3,"rollup_spend":0.00003275,"spendlog_spend":0.00003275,"match":true}

-- head: one "Re-queuing uncommitted transactions" line and one "dropped 1 daily user spend rows" line at 23:44:18Z, then nothing
-- head, mid-window read at 23:44:46Z
{"endpoint":"chat","rollup_requests":2,"rollup_spend":0.00001725,"spendlog_spend":0.0000265,"match":false}
{"endpoint":"responses","rollup_requests":3,"rollup_spend":0.00003275,"spendlog_spend":0.00003275,"match":true}
{"endpoint":"messages","rollup_requests":3,"rollup_spend":0.0000315,"spendlog_spend":0.0000315,"match":true}
-- head, final read at 23:45:02Z
{"endpoint":"chat","rollup_requests":3,"rollup_spend":0.0000263,"spendlog_spend":0.00003555,"match":false}
{"endpoint":"responses","rollup_requests":3,"rollup_spend":0.00003275,"spendlog_spend":0.00003275,"match":true}
{"endpoint":"messages","rollup_requests":3,"rollup_spend":0.0000315,"spendlog_spend":0.0000315,"match":true}

What changed on the Redis-buffered path: the merge base re-sends the whole uncommitted set every cycle, so while one row is refused no daily user rollup row lands for any key (responses and messages sat at 2 above until the constraint was dropped), and a rejection that never clears freezes the Usage page for every key until a restart drops the buffer; the head drops the refused row's batch once (one row here: the two instances' cycles put the chat and responses rows in different batches) and every other row keeps landing. Same rule as the no-Redis path, same trade: the refused batch is gone from the rollups and named in one error line, and the spend logs still carry it. The per-entity spend increments and the window path on the Redis-buffered side were not changed, and the tag rollup kept every request on both sides (tag_requests above, on its own 23s cadence)

Type

🐛 Bug Fix

Caveats (if any)

Low

  • A reply lost after a rollup statement was sent (a read timeout) still loses that statement's rows, up to 100, as it does today: the proxy cannot tell whether the statement committed, and LIT-4823 chose losing that class over counting it twice. Keeping both takes an idempotent upsert (a batch id column plus a migration), which is out of scope here; the drop is logged with the row count, and the batches behind it are kept
  • A batch Postgres refuses outright drops the whole batch of up to 100 rows that carried the bad row, not only that row, which is what happened before this PR on every failure; bisecting to the bad row the way the spend-log writer does is a follow-up. That rule now covers the daily rollups on the Redis-buffered path too (its restore step puts back what the commit left in the dict), where the merge base kept re-sending the refused batch and held every other daily row with it; the window path and the Redis-buffered per-entity spend increments still requeue such a batch forever, a pre-existing gap outside this PR
  • While the database is down, the requeued rows stay in memory. They aggregate per key, model, and day on the way back in, so the queue grows with distinct combinations seen during the outage, not with request count
  • During a full database outage every daily table now retries each tick before the next one, so a tick takes a few seconds longer; nothing commits either way and the rows are kept
  • When the batch that failed was the last one still in the dict, the flush helper has nothing to put back and logs nothing more; the dropped line from the commit is that tick's only record. A second line there would repeat the first
  • The tag job's no-Redis branch now calls the tag commit on every tick, an empty queue included, which costs one debug log line per tick; skipping the call for an empty queue would special-case one of the six callers of the same helper
  • CircleCI at c6c8aed (pipeline 89815) is red on four tests that main's own scheduled pipelines fail the same way and that this diff never touches: test_bedrock_invoke_messages_with_all_beta_headers for both Bedrock params (main pipeline 89818 job 2188120; red since fix(anthropic): register thinking-binding-controls-2026-08-01 in beta headers config #41203 registered thinking-binding-controls-2026-08-01 as a Bedrock pass-through, tracked in LIT-8150), test_models_by_provider (main pipeline 89818 job 2188086; red since feat(proxy): add Amazon Transcribe pass-through with completion-time job pricing #41515 added the transcribe cost map entry without registering the provider), and the two integration retry tests (main pipelines 89785 jobs 2186412 and 2186414 and 89818 jobs 2188082 and 2188083, whose fix fix(proxy): refuse config-owned keys on POST /config/update #41868 reached main at 00:23Z on 2026-09-19, after main's last scheduled run). All three breaks landed on main before this branch's merge base 88e150b, and merging main in would clear only the integration pair, so the branch stays on its merge base

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Note

Medium Risk
Changes billing/usage rollup persistence on the no-Redis path; failures are requeued (with a known ambiguous post-send double-count tradeoff) instead of dropped, which affects Usage and daily activity APIs.

Overview
Fixes permanent loss of daily rollup spend when the proxy commits without the Redis transaction buffer and a database write fails. Previously, flushed in-memory rows were dropped and the error could stop the rest of the tick’s daily flushes.

All daily entity queues (user, team, org, end user, agent, and tag via its separate job) now go through _flush_daily_spend_queue: flush aggregated transactions, call the existing per-table commit, and on failure log, re-queue the batch with add_update, and continue—matching the window-spend error handling already in the same path. Redis-buffered commit flows are unchanged.

Adds regression tests that a failed user daily upsert requeues while team still commits, and that failed tag commits requeue for the next scheduler tick.

Reviewed by Cursor Bugbot for commit 3edbf60. Bugbot is set up for automated code reviews on this repo. Configure here.

…e Redis buffer

With the Redis transaction buffer off, each daily spend queue (user, team, org,
end user, agent) was drained into a dict and handed to the bulk upsert. When
the upsert raised after its retries, the drained dict was discarded and the
exception escaped update_spend, so those rows never reached the daily rollup
tables and /user/daily/activity stayed short forever while /spend/logs had
every request.

Each daily queue now flushes through one helper that puts the uncommitted
remainder back on the queue for the next tick and moves on to the next table,
the same shape the window-spend step already used.
@devin-ai-integration

devin-ai-integration Bot commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor Author

🤖 Devin AI Engineer

I'll be helping with this pull request! Here's what you should know:

✅ I will automatically:

  • Address comments on this PR. Add '(aside)' to your comment to have me ignore it.
  • Look at CI failures and help fix them

Note: I can only respond to comments from users who have write access to this repository.

⚙️ Control Options:

  • Disable automatic comment, CI, and merge conflict monitoring

@codspeed

codspeed Bot commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_requeue_daily_spend_without_redis_buffer (c6c8aed) with main (018f640)

Open in CodSpeed

@greptile-apps

greptile-apps Bot commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

The PR appears safe to merge, with no outstanding correctness or repository-rule findings

Summary

This PR preserves uncommitted daily spend rows after retry-safe database failures and allows the remaining rollup tables to flush during the same tick

  • Routes user, team, organization, end-user, agent, and tag queues through shared flush and requeue handling
  • Drops only the statement batch when resending could duplicate spend or repeatedly submit rejected data
  • Adds typed SQLSTATE extraction and regression coverage for retry-safe, ambiguous, rejected-data, and multi-batch failures

Reviews (5) · Last reviewed commit: "fix(proxy): drop only the daily spend ba..."

Comment thread litellm/proxy/db/db_spend_update_writer.py
@codecov

codecov Bot commented Sep 18, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 90.69767% with 4 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
litellm/proxy/db/db_spend_update_writer.py 86.66% 4 Missing ⚠️

📢 Thoughts on this report? Let us know!

@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

Comment thread tests/test_litellm/proxy/db/test_exception_handler.py Outdated
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread litellm/proxy/db/db_spend_update_writer.py
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@mateo-berri mateo-berri added run-ci and removed run-ci labels Sep 18, 2026

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit c6c8aed. Configure here.

@mateo-berri mateo-berri left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@mateo-berri
mateo-berri merged commit f6d4766 into main Sep 19, 2026
138 of 143 checks passed
@mateo-berri
mateo-berri deleted the litellm_requeue_daily_spend_without_redis_buffer branch September 19, 2026 00:27
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant