Skip to content

test(integration): endpoint, breakdown component and failure support in the cost harness - #41999

Merged
kerry-berri merged 9 commits into
mainfrom
litellm_cost_shard_harness_extensions
Sep 21, 2026
Merged

kerry-berri merged 9 commits into
mainfrom
litellm_cost_shard_harness_extensions

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Every cost case can only post to /v1/chat/completions
  • Cache, reasoning and tool costs are only checked through their sum
  • Upstream errors cannot be scripted, so zero-spend failure rows are untested
  • Per-component x-litellm-response-cost-* headers are never asserted

How it solves it:

  • Optional endpoint on each case, defaulting to chat completions
  • Optional cache_read_cost, cache_creation_cost, reasoning_cost, tool_usage_cost expectations
  • Each set component is checked on the stored breakdown and the response header
  • status on JSON scripted responses plus a failure expected variant
  • Failure cases assert proxy status, zero cost header and a zero-spend failure row
  • Data validation rejects inconsistent components or status combinations
  • 3 new cases (native /v1/responses, upstream 500, upstream 429) and 8 existing cases now carry component expectations
  • Fireworks fallback cache-read case updated for the 50% default that landed in 1b305cd

User Flow

Before: a maintainer wants to pin the cache-read component of a Claude call and the zero-spend row of a failed call, and the suite has nowhere to put either

  1. They add cache_read_cost to an entry in tests/integration/cost_calculation/cost_tracking_cases.json
  2. The suite rejects the file at import with a Pydantic "extra fields not permitted" error
  3. They add a case whose scripted upstream answers 500 and the test fails at assert response.is_success before looking at the spend row

After: both expectations are plain data in the same file

  1. They add "cache_read_cost": 0.00405504 next to input_cost on the Bedrock cache-read case
  2. The run posts the case, then compares the stored breakdown and x-litellm-response-cost-cache-read against that value, and x-litellm-response-cost-input against input_cost minus the cache components
  3. They add a case with "response": {"content_type": "application/json", "status": 500, "body": {...}} and "expected": {"failure": {"status": 500}}
  4. The run asserts the proxy answered 500, no non-zero cost header, and a status = failure spend row with spend = 0
  5. They add "endpoint": "/v1/responses" to a case and the run posts to that path instead of chat completions

Relevant issues

Affected release

Linear ticket

Resolves LIT-8185

Pre-Submission checklist

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Screenshots / Proof of Fix

The deliverable is the integration shard itself, run through tests/integration/run.py against a live proxy, Postgres, Redis and the scripted upstream

Before (659cef0)

  1. python tests/integration/run.py --group cost on the merge base: 363 cases, none can post to another endpoint, script a non-200 upstream, or name a breakdown component

After (9bd648b)

  1. Targeted run of the new and edited cases against the live harness: 10 passed, 356 deselected in 13.59s, then deepseek-v4p1-flash-fallback_cache_read_at_half_input_rate alone: 1 passed, 365 deselected in 2.85s
  2. Full cost group at e42f3f1: 1 failed, 365 passed in 856.31s, the failure being the Fireworks case fixed in the last commit

Type

✅ Test

Caveats (if any)

Low

  • Streaming responses send headers before the cost is known, so the cost header stays JSON-only
  • Provider prefix additions (azure_ai, openrouter, deepseek, ...) land with their cases in LIT-8189

Link to Devin session: https://app.devin.ai/sessions/5425681f498543a19de84adcf8daeed8
Open in Devin Desktop: https://app.devin.ai/desktop/session/5425681f498543a19de84adcf8daeed8?variant=devin

kerry and others added 2 commits September 19, 2026 18:44
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration
devin-ai-integration Bot requested a review from a team September 19, 2026 19:13
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

@greptile-apps

greptile-apps Bot commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

The PR appears safe to merge; no actionable new issues remain, and the prior validation finding was resolved.

Summary

This PR expands the cost-tracking integration harness to exercise multiple endpoints, detailed cost components, and failed upstream requests.

  • Adds configurable endpoint and scripted JSON response status support.
  • Validates component-cost expectations against spend-log breakdowns and response headers.
  • Verifies failed requests produce the expected proxy status and a zero-spend failure row.
  • Adds native Responses API, upstream 500, and upstream 429 fixtures.
  • The follow-up change constrains both scripted and expected failure statuses to the HTTP 400–599 range.

Reviews (2) · Last reviewed commit: "test(integration): reject out-of-range f..."

Comment thread tests/integration/cost_calculation/cost_tracking_case.py
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

Pushed dccb1b5 adding the failure status range check Greptile asked for. Rebutted the upstream-equals-proxy status equality since the proxy remaps some codes

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit dccb1b5. Configure here.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@codecov

codecov Bot commented Sep 19, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

kerry and others added 3 commits September 19, 2026 19:52
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…pectations

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…essages

test(integration): native responses and messages cost cases
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@kerry-berri
kerry-berri merged commit 00c41e8 into main Sep 21, 2026
93 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant