Repository navigation
test(e2e): tolerate provider-side flakes on five full-suite cells - #42628
Conversation
Mistral OCR retries a provider-relayed 429 with backoff, the Vertex vision probe turns reasoning off so its 32 tokens go to the answer, the Vertex cache cell spaces eight never-seen prefixes 15s apart around Google's nondeterministic minimum-token rejection and prices the cached tokens instead of prompt_tokens, and the Azure content-policy cell resends the jailbreak prompt while Azure skips its filter
|
I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".
|
|
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
The gateway already retries a provider 429 three times per call and the Mistral key is shared across pipelines, so a throttle can hold across all four attempts of the OCR cell. After the bounded retries the cell now accepts the gateway's faithful relay of the provider's 429 (throttling_error, code 429) as its second expected outcome; the gateway's own 429 and any other error still fail the cell at once.
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit fc88955. Configure here.
…lake tolerance to rc/1.103.0 (#43400) * test(e2e): tolerate provider-side flakes on five full-suite cells (#42628) * test(e2e): tolerate provider-side flakes on five full-suite cells Mistral OCR retries a provider-relayed 429 with backoff, the Vertex vision probe turns reasoning off so its 32 tokens go to the answer, the Vertex cache cell spaces eight never-seen prefixes 15s apart around Google's nondeterministic minimum-token rejection and prices the cached tokens instead of prompt_tokens, and the Azure content-policy cell resends the jailbreak prompt while Azure skips its filter * test(e2e): shorten the new helper docstrings * test(e2e): accept a relayed provider 429 on the rust OCR cells The gateway already retries a provider 429 three times per call and the Mistral key is shared across pipelines, so a throttle can hold across all four attempts of the OCR cell. After the bounded retries the cell now accepts the gateway's faithful relay of the provider's 429 (throttling_error, code 429) as its second expected outcome; the gateway's own 429 and any other error still fail the cell at once. * test(e2e): drop the harness unit tests, the live cells cover the helpers --------- Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com> (cherry picked from commit 41ca465) * fix(streaming): keep litellm Usage on text-completion usage chunks (#43047) * fix(streaming): keep litellm Usage on text-completion usage chunks * fix(streaming): convert provider usage to litellm Usage instead of dropping it (cherry picked from commit be35b22) --------- Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
TLDR
Problem this solves:
How it solves it:
reasoning_effort: "none"so the 32 tokens go to the answercontent_filterincomplete was a one-offUser Flow
Before: a contributor's litellm PR gets five red cells on the litellm-e2e-pr full suite for provider behavior their change never touched
"model": "rust-ocr-mistral-<id>"and the dummy.pdf document URL; it gets 429MistralException - (429, ... "Rate limit exceeded" ... "code":"1300")followed byNo fallback model group found, and the cell goes red"max_tokens": 32; the 200 carries"content": null,"finish_reason": "length"and every completion token counted underreasoning_tokens, and the cell goes red for not naming the catcache_controlback to back to POST http://litellm-domain/v1/chat/completions; Vertex answers each with 400The cached content is of 1 tokens. The minimum token count to start explicit caching is 1024., and the cell goes red with "explicit context caching did not engage""cached_tokens": 13109next to"prompt_tokens": 5; GET http://litellm-domain/spend/logs?request_id= shows"spend": 0.00046827, the cell compares it to the uncached price of 5 tokens (0.0000015) and goes red with "cache-read pricing was not applied"azure/gpt-5.4-nanodeployment expecting a 400; Azure answers 200 with"prompt_filter_results":[{"prompt_index":0,"content_filter_results":{}}]and a refusal written in the message text, and the cell goes redAfter: the same PR's full suite stays green unless the gateway itself misbehaves
"type":"throttling_error","code":"429", message startinglitellm.RateLimitError) and prints that it did; a 429 from the gateway's own rate limiters still fails the cell at once"reasoning_effort": "none"added; the 200 carries"content": "Cat"with"finish_reason": "stop"cached_tokensof 1024 or more is the cold call under testprompt_tokenscontent_filterit asserts on and, on the second leg withcontent_policy_fallbacksset, the 200 served by gpt-5.5Relevant issues
No GitHub issue. Red on litellm-e2e-pr builds 747, 752, 754 and 755 (the full-suite runs for #42518 and #42503). LIT-8379 tracks Vertex reporting
prompt_tokensinclusive of the cached prefix on some calls and exclusive on others, which this PR works around in the test and does not fixAffected release
Linear ticket
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Shared setup: a local copy of the
.github/e2e-stacktwo-gateway stack (nginx LB at http://127.0.0.1:20057, control plane at http://127.0.0.1:25724, Postgres, Valkey, Keycloak, Jaeger),tests/e2e/.envcarrying the Vertex, Mistral, Azure and OpenAI credentials, and the five cells run one by one withpytest -p no:cacheprovideron their node ids. Buildkite lines come from the full-suite runs on the e2e runner. Real provider calls throughoutBefore (238f434)
The local Before ran at 238f434, the main tip when the branch was cut; the merge base 7688f56 carries byte-identical files for these five cells. Buildkite builds 747, 752, 754 and 755 are the litellm-e2e-pr full suite for #42518 and #42503
rust-ocr mistral
tests/e2e/llm_translation/test_ocr_rust_e2e.py::TestRustOcrGateway::test_rust_ocr_response[mistral]:litellm.RateLimitError: RateLimitError: MistralException - (429, ... "Rate limit exceeded" ... "code":"1300" ...)No fallback model group found for lookup_groups=rust-ocr-mistral-... LiteLLM Retried: 3 times, LiteLLM Max Retries: 3vertex vision
curl -s -X POST http://127.0.0.1:20057/v1/chat/completions -H "Authorization: Bearer $KEY" -H "Content-Type: application/json" -d '{"model":"e2e-vertex-vision-proof","max_tokens":32,"messages":[{"role":"user","content":[{"type":"text","text":"What animal is in this image? Answer in one word."},{"type":"image_url","image_url":{"url":"data:image/jpeg;base64,<tests/e2e/llm_translation/fixtures/cat.jpg>"}}]}]}'{"content": null, "finish_reason": "length", "usage": {"completion_tokens": 28, "prompt_tokens": 1302, "total_tokens": 1330, "completion_tokens_details": {"reasoning_tokens": 28, "text_tokens": 0}, "prompt_tokens_details": {"text_tokens": 12, "image_tokens": 1290}}}(build 747 saw 29 reasoning tokens of 32)vision response did not describe the imagevertex prompt caching
system_instruction, the "." placeholder user turn,ttl300s) posted straight tohttps://us-central1-aiplatform.googleapis.com/v1/projects/<project>/locations/us-central1/cachedContentswith fresh content each time, one POST every ~3s for 102s:oxxoxooxxxxoooxxxxxxxxxxxxxxxoxxoxxxwhereois 200 withtotalTokenCountnear 22k andxis 400The cached content is of 1 tokens. The minimum token count to start explicit caching is 1024.; 9 of 36 succeeded, with failure bursts of 45s and longerAssertionError: e2e-vertex-cache-2a70f31b97a9: 3 never-seen prompts marked with cache_control were each either rejected by Vertex's minimum-token check or served with fewer than 1024 cached tokens on their first call; explicit context caching did not engagecached_tokens=13109,prompt_tokens=5,text_tokens=-13104,completion_tokens=30, spend row0.00046827; the cell asserted0.00046827 < 1.5e-06(5 prompt tokens at 3e-7) and failed withcache-read pricing was not applied. The cost map prices gemini-2.5-flash at 3e-7 input, 3e-8 cache read and 2.5e-6 output, and 13109 × 3e-8 + 30 × 2.5e-6 = 0.00046827 exactly, so the cache-read rate was applied and the baseline was wrongprompt_tokensexcluded the cached prefix)responses vertex
tests/e2e/llm_translation/test_responses_e2e.py::TestResponses::test_responses_vertex_returns_completion:status incomplete, reason content_filteron "reply with one word"; the only occurrence, 28-day pass rate above 0.946azure content policy
tests/e2e/router/test_reliability_fallbacks_e2e.py::TestReliabilityFallbacks::test_content_policy_routes_to_fallback: the primary leg got 200 fromazure/gpt-5.4-nanowith"prompt_filter_results":[{"prompt_index":0,"content_filter_results":{}}]and the messageI can't comply with requests to ignore previous instructions or claim to be "DAN"...; the cell asserted a 400 and failed. Azure skipped its prompt filter on 23 of 145 runs in 28 daysAfter (fc88955)
Commits since the runs labeled 0d3be46 below: e18fb87 trims two docstrings, 14db6db changes only the OCR cell and the harness helper it calls, and fc88955 only deletes the two harness unit-test files, so the four other cells' local runs stand as recorded
rust-ocr mistral
1 passed, 4 deselected in 54.12s, no throttle in this run; when Mistral throttles, the run log showse2e-http: provider rate limit relayed by the proxy; retry N/3 in Nsbefore each retry and, if all four attempts are throttled,mistral/mistral-ocr-latest: the provider's rate limit held across 4 attempts; the gateway relayed it as throttling_error 429before the passvertex vision
"reasoning_effort":"none"added to the body{"content": "Cat", "finish_reason": "stop", "usage": {"completion_tokens": 1, "prompt_tokens": 1302, "total_tokens": 1303, "completion_tokens_details": {"text_tokens": 1}, "prompt_tokens_details": {"text_tokens": 12, "image_tokens": 1290}}}vertex prompt caching
cache_control: vertex did not engage the cache on cold attempt N/8; pausing 15.0s before the next never-seen prefixresponses vertex
azure content policy
e2e: azure answered without running its prompt filter; sending the jailbreak prompt again (N/2)before the resendFive cells together at 0d3be46 against the local stack:
5 passed in 103.39sType
✅ Test
Caveats (if any)
Low
oxxoxooxxxxoooxxxxxxxxxxxxxxxoxxoxxx), a burst past the budget is a Vertex incident that reds the suite anyway, and a bigger budget costs another full-suite cycle for a case no run has shownspend < prompt_tokens * input_ratecheck never caught that case either (0.00046827 spend against a 0.0039342 threshold), the new one still catches a cache read billed at the full rate (0.0040077 is not below 0.0039327), and LIT-8379 owns normalizingprompt_tokensagainstcached_tokenson the product side, which this suite-only PR must not touchmisc / Run testsis red ontests/test_litellm/test_utils.py::test_wrapper_async_does_not_fire_failure_hook_for_post_success_error(DID NOT RAISE), a file this tests/e2e-only PR never touches; main's own Unit Tests runs at b0ac23d (https://github.com/BerriAI/litellm/actions/runs/35807522965) and the three before it fail the same way after fix(utils): isolate callback errors in async_post_call_success_deployment_hook #42535 (944f44d); test(utils): raise the post-success hook error from a guardrail in the failure-hook regression #42646 fixed it on main at 5fa1257 (Unit Tests green, https://github.com/BerriAI/litellm/actions/runs/35809354378), and the shard is not one of main's required status contexts, so it does not gate this PRcontent_filter_results; a changed Azure shape shows up in the assertion text, never as a silent passlitellm.RateLimitErrorwiththrottling_errorand429: the relayed Mistral 429, and a router cooldown (No deployments available for selected model), which only follows that same provider throttle; a 5xx, a timeout, or a 429 from the proxy's own rate limiters still fails it, because those reach the client throughproxy_exception_from_http_exceptionwith the limiter text as the message and nolitellm.RateLimitErrormarker (litellm/proxy/common_request_processing.py:595); that rejection is proven by code read, no live gateway-429 leg ran, and tightening the marker to the provider signature (MistralException -) would cost a full-suite cycle to exclude a cooldown the cell would tolerate anywayrelayed_provider_rate_limitraises a pydanticValidationError(a traceback, not the cell's assertion text) if a 429 body carries the marker in a shape other than{"error": {"message", "type", "code"}};ProxyException.to_dictis the one writer of that body today, so a failure-as-value rewrite would add code for a shape nothing producestolerate_provider_rate_limit,_first_engaged_cold_call,_chat_once_azure_runs_its_filter) run only when the provider misbehaves, so a healthy suite run never exercises them; the harness unit tests that did were dropped at fc88955 because tests/e2e holds live-stack cells only, and build 784's OCR cell passed in 89s with no retry line in its log, so no run in this PR is known to have taken the tolerant branch; that branch is proven by code readquota_management/ratelimit/test_model_group_alias_rate_limit_e2e.pycells (400Invalid model name passed in model=e2e-alias-rl-target) and errored collectinglogging/test_langsmith_batch_serialization_e2e.py(import litellmin a runner without it); both files landed on main on 2026-09-22 (fix(proxy): share model rate-limit buckets between a model_group_alias and its target #42516, fix(langsmith): json.dumps with default=str so non-serializable metadata does not crash batch flush #42424), this PR leaves them alone, and they are tracked on their own. Build 784's tally was 2 failed, 1406 passed, 57 skipped, 1 error in 46:32, and those three are its only redsQA runbook
tests/e2e/llm_translation/test_ocr_rust_e2e.py::TestRustOcrGateway::test_rust_ocr_response[mistral] - a Mistral OCR deployment returns the parsed document even when Mistral throttles the shared key a few times in a row
{"model_name": "rust-ocr-mistral-qa", "litellm_params": {"model": "mistral/mistral-ocr-latest", "api_key": "os.environ/MISTRAL_API_KEY"}}(needs MISTRAL_API_KEY){"model": "rust-ocr-mistral-qa", "document": {"type": "document_url", "document_url": "<the dummy.pdf fixture URL in the test>"}}and expect 200 with apageslist carrying markdown"type":"throttling_error","code":"429") which it accepts and prints, while a 429 from a key withrpm_limit: 1fails at oncetests/e2e/llm_translation/test_chat_completions_regression_e2e.py::TestVertexChatCompletions::test_vertex_chat_vision_describes_image - a 32-token vision request over Vertex names the cat instead of spending the budget on reasoning
{"model_name": "vertex-vision-qa", "litellm_params": {"model": "vertex_ai/gemini-2.5-flash", "vertex_project": "<project>", "vertex_location": "us-central1", "vertex_credentials": "<service account JSON>"}}(needs VERTEXAI_PROJECT and VERTEXAI_CREDENTIALS)"max_tokens": 32,"reasoning_effort": "none"; expect 200 with "cat" inchoices[0].message.contentand"finish_reason": "stop"reasoning_effortand expect"content": null,"finish_reason": "length"andreasoning_tokensnear 32, which is the failure this guardstests/e2e/llm_translation/test_cache_control.py::TestCacheControl::test_vertex_prompt_caching_reads_cache - the first call with a never-seen cache_control system prompt is served from Vertex's explicit cache and billed at the cache-read rate
{"model": "vertex-vision-qa", "messages": [{"role": "system", "content": [{"type": "text", "text": "<a fresh text of about 22k tokens>", "cache_control": {"type": "ephemeral", "ttl": "300s"}}]}, {"role": "user", "content": [{"type": "text", "text": "Reply with one word."}]}]}The cached content is of 1 tokenswait 15s and send again with a new fresh text, up to eight times; on the 200 expectusage.prompt_tokens_details.cached_tokensof at least 1024 and at least 90% ofprompt_tokensspendbelowcached_tokenstimes theinput_cost_per_tokenfrom GET http://localhost:4000/model/info (13109 cached tokens: below 0.0039327; a row priced at the cache-read rate lands near 0.00047)tests/e2e/router/test_reliability_fallbacks_e2e.py::TestReliabilityFallbacks::test_content_policy_routes_to_fallback - Azure's content filter refuses the jailbreak with a 400, and with
content_policy_fallbacksset the same prompt is served by gpt-5.5azure/gpt-5.4-nanodeployment withapi_key,api_baseandapi_version(needs the Azure OpenAI credentials)content_filterin the body; a 200 carrying"prompt_filter_results":[{"prompt_index":0,"content_filter_results":{}}]means Azure skipped its filter, so send again (the cell does this up to twice)content_policy_fallbacks: [{"<primary>": ["gpt-5.5"]}]and expect 200 served by gpt-5.5 (thex-litellm-model-idheader andmodelin the body)Decisions
litellm.RateLimitError, four attempts, 5s geometric backoff honoring Retry-After; no fallback group and no second key, since the throttle comes from five suites sharing one keyreasoning_effort: "none"withmax_tokenskept at 32, so the cell still proves a tight budget answers, instead of raising the budget until reasoning fitsprompt_tokenstocached_tokenstimes the input rate, and thecached <= prompt_tokensself-consistency assertion was dropped because Vertex's exclusive counts (LIT-8379) would keep the cell red about half the timecontent_filterincomplete in 28 days is below the change threshold; the prompt gets swapped if it repeatsProxyErrorDetail) that lives ine2e_http.py, sincemodels.pyalready imports from itFinal Attestation
The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR
fc88955 passes /live-pr-risk