Repository navigation
test(integration): regression tests for August cost tracking and budgeting bugs - #42622
Conversation
…ricing as a deployment override (Pylon #6870) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…s budget_reset_at (Pylon #6913) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…and a later completion still succeeds (Pylon #6966) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…loyment absent from the cost map (Pylon #7014) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…SON (Pylon #7180) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
… team spend in one page (Pylon #7224) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…end report and daily activity agree (Pylon #7268) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
… capped by the team organization budget (Pylon #7291) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ervation from the spend counter (Pylon #7295) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…d counts output and error file failures (Pylon #7341) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…so newer batches are costed (Pylon #7342) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ts membership row (Pylon #7363) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…kens in spend logs (Pylon #7519) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…l definitions (Pylon #7524) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…replaced by default_team_params (Pylon #7536) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…s it (Pylon #7577) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ut leaking pricing fields upstream (Pylon #7587) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…log for an Azure Model Router alias (Pylon #7636) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ls terminal usage (Pylon #7685) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…s (Pylon #7738) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
… and error file failures (Pylon #7928) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…age (Pylon #7958) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…get so a completion still reaches the provider (Pylon #7307) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ing budget before provider (Pylon #7691) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ider response headers (Pylon #7775) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…n tests Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…tests Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".
|
|
|
|
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
… callback batch accumulation Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
This comment has been minimized.
This comment has been minimized.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
| yield | ||
| finally: | ||
| cleared: Final = gateway.request("PATCH", "/update/default_team_settings", {}) | ||
| assert cleared.status_code == 200, cleared.text |
There was a problem hiding this comment.
Team defaults leak across shared proxy
Medium Severity
_default_team_budget_duration patches /update/default_team_settings on the shared proxy and persists it to the config DB. Cleanup sends {}, which stores {models: []} rather than restoring the original unset default_team_params, so later team-create tests on this shard (or after a reload) inherit that leftover default.
Reviewed by Cursor Bugbot for commit 2d2bdf8. Configure here.
There was a problem hiding this comment.
Fixed in 240b274: the test now runs on an owned proxy with the default set in yaml, so nothing is persisted
…stic Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ion_tests_accounting
…ion_tests_accounting Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ion_tests_accounting Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
These tests are not vacuous. Most exercise the real proxy and verify independent outcomes such as response cost plus persisted spend, budget decisions plus upstream dispatch, or agreement between daily activity and global spend reports. Several also assert the exact request received by the scripted upstream, so a canned success response cannot make them pass. The documented red/green mutation runs further support that they detect regressions. There is moderate intentional coupling: some tests assert internal details such as Redis key names, Overall: high-signal regression coverage with moderate intentional coupling, not overfit or vacuous. If brittleness appears, relax implementation-name/storage-shape assertions while preserving the independent accounting invariants and upstream-dispatch checks. |
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit 865c3c2. Configure here.
| batch_id: Final = string_value(JSON_OBJECT.validate_json(batch_response.content)["id"]) | ||
| retrieval: Final = gateway.request("GET", f"/v1/batches/{batch_id}", key=key) | ||
| assert retrieval.status_code == 200, retrieval.text | ||
| assert retrieval.json()["status"] == "completed", retrieval.text |
There was a problem hiding this comment.
Batch status asserted without polling
Low Severity
GET /v1/batches/{id} is asserted completed on the first response. The same lifecycle in test_batch_observability.py polls for that status. A retrieve that still reflects the create-time validating row fails immediately even though spend is polled for 70s afterward.
Additional Locations (1)
Reviewed by Cursor Bugbot for commit 865c3c2. Configure here.


TLDR
Problem this solves:
How it solves it:
User Flow
Before: an engineer refactors spend tracking and reintroduces one of the August regressions, for example OCR annotation pages billed at the plain page rate
x-litellm-response-costheader and https://litellm-domain/ui/?page=logs both under-report the spendAfter: the same refactor turns the accounting integration shard red before merge
integration-accountingshard fails ontest_ocr_annotation_pages_are_billed_at_annotation_cost_per_page, showing the expected and actual costRelevant issues
Regression coverage for August customer bug reports, one Pylon ticket per test, listed in the table below. No open issue is fixed by this PR
Affected release
Linear ticket
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
This PR ships tests only, so the proof is the mutation check tests/AGENTS.md asks for: each test was run with its fix commit reverted (red) and with the fix restored (green), through the same real proxy, Postgres, Redis and scripted upstream the CircleCI shard uses. The harness is deterministic by design and does not call real providers. The accounting shard passed locally on the tip with two order seeds (30 passed each), plus management (21 passed) and extensions (21 passed)
Before (fix commit reverted on top of df7db2c)
Per test, the fix sha in the table was reverted (
git revert --no-commit <sha>, or a worktree at the fix parent where the revert conflicted), the proxy restarted, and the new node run. Every node failed on the behavioral assertion. Output is in the fold-outs belowAfter (df7db2c)
Fix restored, same node rerun, passed. The shards:
python tests/integration/run.py accountingreports 30 passed withINTEGRATION_ORDER_SEED=1and again withINTEGRATION_ORDER_SEED=2,management21 passed,extensions21 passedTests and their red/green evidence
a969319fb5tests/integration/pricing/test_configured_prices.py::test_saving_echoed_model_info_does_not_freeze_cost_map_price_into_deployment96b9437bddtests/integration/management/test_budget_updates.py::test_shortening_budget_duration_moves_reset_at_onto_the_new_schedule599daea985tests/integration/spend/test_cache_and_quota.py::test_repeated_count_tokens_on_budgeted_key_does_not_reserve_budget_or_block_later_completion9bfe593241tests/integration/pricing/test_configured_prices.py::test_cost_estimate_reports_configured_prices_for_model_absent_from_cost_map3ea1c16b0dtests/integration/management/test_team_member_budget_cache.py::test_team_member_default_budget_lands_in_redis_after_first_member_call9ec0145986tests/integration/spend/test_team_daily_activity_aggregated.py::test_aggregated_team_activity_reports_the_whole_range_team_spend_in_one_pagef6d4766ebetests/integration/spend/test_daily_rollup_retry.py::test_failed_daily_user_rollup_commit_is_retried_so_spend_report_and_daily_activity_agree579f41d57ftests/integration/spend/test_org_budget_cli_session_token.py::test_cli_session_token_without_org_id_charges_and_caps_the_team_organization258fe3e4batests/integration/spend/test_passthrough_budget_reservation.py::test_repeated_gemini_passthrough_calls_stay_served_while_key_spend_is_below_max_budget599daea985tests/integration/spend/test_cache_and_quota.py::test_in_flight_count_tokens_does_not_reserve_key_budget_away_from_a_completion20cfccaf5ftests/integration/spend/test_batch_observability.py::test_batch_retrieval_row_sums_reasoning_tokens_and_counts_output_and_error_file_failures2b63919f67tests/integration/spend/test_batch_poll_starvation.py::test_batches_gone_at_provider_do_not_starve_a_newer_batch_out_of_cost_polling006080ea6dtests/integration/spend/test_team_member_spend.py::test_member_added_without_any_budget_is_charged_on_its_membership_rowd6afe728aatests/integration/spend/test_failed_dispatch_tokens.py::test_provider_500_after_dispatch_records_estimated_prompt_tokens_on_failure_row853fed824etests/integration/observability/test_guardrail_effects.py::test_bedrock_passthrough_converse_guardrail_ignores_denied_term_in_tool_definition262ed530f8tests/integration/management/test_team_budget_duration_defaults.py::test_team_new_explicit_null_budget_duration_is_not_replaced_by_defaulta502c728actests/integration/management/test_organization_budget_clear.py::test_patch_organization_update_with_null_tpm_limit_clears_it_and_keeps_sibling_limits6b7adf011etests/integration/pricing/test_service_tier_pricing.py::test_ultrafast_service_tier_bills_ultrafast_rates_and_keeps_pricing_off_the_wire3db6c5ab18tests/integration/spend/test_model_router_selected_model.py::test_model_router_alias_without_router_in_name_keeps_selected_model_in_response_and_spend_log99703a30f0tests/integration/spend/test_disconnected_bedrock_messages_stream_billing.py::test_client_disconnect_mid_bedrock_messages_stream_still_bills_terminal_usage3122600e21tests/integration/pricing/test_databricks_cache_pricing.py::test_databricks_cached_prompt_tokens_bill_at_cache_rates_not_input_ratec19d49d919tests/integration/observability/test_callback_delivery.py::test_streamed_responses_success_callback_carries_provider_apim_request_id20cfccaf5ftests/integration/spend/test_batch_completion_accounting.py::test_completed_batch_spend_row_records_reasoning_tokens_and_error_file_failuresb86791b4b9tests/integration/pricing/test_ocr_page_pricing.py::test_ocr_annotation_pages_are_billed_at_annotation_cost_per_pagePylon 6870: PATCH /model/{id}/update (and POST /model/new) must not persist server-derived pricing keys carried in the request's mod [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 6913: POST /budget/update with a JSON body containing an existing budget_id and a changed budget_duration, without budget_rese [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 6966: The customer-facing regression is on authenticated POST /v1/messages/count_tokens, with an Anthropic-native body such as [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 7014: POST /cost/estimate with a model name that is a router model_group for an on-prem/self-hosted deployment (e.g. nvidia) w [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 7180: Proxy runs with enable_redis_auth_cache: true and a redis cache backend (the integration proxy_config.yaml already enabl [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 7224: The customer-facing team usage flow queried GET /team/daily/activity with start_date, end_date, page_size and page repea [...]
Red method: worktree_at_fix_parent_then_cherry_pick
Red (fix reverted):
Green (fix present):
Pylon 7268: A successful POST /v1/chat/completions through a virtual key can persist its raw spend log while a daily user rollup com [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 7291: Drive a real proxy request with POST /v1/chat/completions using Authorization: Bearer produced by th [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 7295: A caller sends a successful native Gemini passthrough request as POST /gemini/v1beta/models/{model}:generateContent with [...]
Red method: worktree_at_fix_parent_then_cherry_pick
Red (fix reverted):
Green (fix present):
Pylon 7307: For a budgeted proxy key with a small max_budget, POST /v1/messages/count_tokens with an Anthropic Messages request such [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 7341: Drive the proxy batch flow with POST /v1/files using a multipart JSONL batch input, POST /v1/batches with the returned i [...]
Red method: worktree_at_fix_parent_then_cherry_pick
Red (fix reverted):
Green (fix present):
Pylon 7342: A customer creates a managed batch with POST /v1/batches, supplying input_file_id, endpoint /v1/chat/completions, comple [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 7363: Regression flow: an admin POSTs /team/member_add with {"team_id":"","member":{"user_id":"","role":"use [...]
Red method: worktree_at_fix_parent_then_cherry_pick
Red (fix reverted):
Green (fix present):
Pylon 7519: For a dispatched provider failure, exercise POST /v1/chat/completions with a configured model whose scripted upstream re [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 7524: POST /bedrock/model/{modelId}/converse (and converse-stream) through the proxy with a pre_call guardrail enabled (defaul [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 7536: The regression is on authenticated POST /team/new. With litellm.default_team_params.budget_duration set to 30d, sending [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 7577: The customer sent PATCH /organization/update with Authorization: Bearer and JSON {"organization_id":""," [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 7587: A proxy caller sends POST /v1/chat/completions with a registered model, messages, and service_tier=ultrafast. When the d [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 7636: For POST /v1/chat/completions using an Azure AI Model Router deployment exposed through an arbitrary model-group alias s [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 7685: For a proxy POST /v1/messages request with stream=true routed to a Bedrock Anthropic Claude model, a client that stops r [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 7738: A caller sends POST /v1/chat/completions with an OpenAI-compatible body such as {"model":"databricks/databricks-claude-s [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 7775: For a proxy POST to /v1/responses with stream=true, the provider receives the Responses request and returns a 200 SSE st [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 7928: PR 37208 fixes completed batch accounting on the proxy batch lifecycle: after POST /v1/files and POST /v1/batches, GET / [...]
Red method: worktree_at_fix_parent_then_cherry_pick
Red (fix reverted):
Green (fix present):
Pylon 7958: POST /v1/ocr with a response whose `usage_info` includes `pages_processed_annotation` should bill those pages at `annota [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Commit df7db2c answers the review: the budget reset test asserts the new reset_at falls inside the next daily window from the update time instead of matching a second control budget, and the callback delivery poll keeps drained batches in a Final list with a mutable-ok reason since drain() consumes the queue; both were rerun red (fix reverted) and green. In CircleCI, pipeline 90108 (f54bc74) ran all eight integration shards green, including accounting (30 passed), management and extensions. Pipeline 90117 ran 767cb7d green on every integration shard
Commit 2d2bdf8 merges main (contracts.json union plus main's appended test in test_callback_delivery.py), pipeline 90122 ran it green on the integration workflow. Commit 240b274 answers the follow-up Bugbot review: the budget reset test starts from a 10d duration, since 30d is a monthly reset that can share a midnight with the 1d reset at month end, and the team default test moves to an owned proxy with
default_team_params.budget_durationset in yaml andSTORE_MODEL_IN_DB=False, so it no longer writes default team settings into the shared config; both tests were rerun red (fix reverted) and greenType
✅ Test
Caveats (if any)
Low
Link to Devin session: https://app.devin.ai/sessions/d330510cfc8f463c823bd79c40b7e5e2
Open in Devin Desktop: https://app.devin.ai/desktop/session/d330510cfc8f463c823bd79c40b7e5e2?variant=devin
Note
Low Risk
Test-only changes plus harness/config tweaks for integration CI; no production proxy or billing logic is modified in this diff.
Overview
Adds 24 end-to-end integration tests (one per previously fixed customer bug) that drive the real proxy against scripted upstreams and assert
x-litellm-response-cost, Postgres spend/budget rows, Redis cache, callbacks, or 422 budget decisions. New suites sit undertests/integration/spend,pricing,management, andobservability, each wired intocontracts.jsonvia@pytest.mark.coversfor shard contract checks.The shared
wiretest server gainsReply.headers,pause_between_chunks, and chunked flush behavior so tests can simulate slow Bedrock event streams (client disconnect billing) and SSE/v1/responseswith provider headers in generic-API callbacks.proxy_config.yamlsetsproxy_batch_polling_interval: 1to keep batch cost-polling tests deterministic.A few existing integration tests are tightened (callback batch draining, cache partition helper) and
owned_proxyis used where tests must not mutate shared default team params.Reviewed by Cursor Bugbot for commit 865c3c2. Bugbot is set up for automated code reviews on this repo. Configure here.