Repository navigation
fix(passthrough): log upstream 4xx/5xx error bodies and carry them into the failure hook - #42695
Conversation
…from logs and spend row Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…to the failure hook Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…m body text Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".
|
|
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…before logging Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…d-routes cast Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ad stays bounded Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
UI proof at a947ba7: the Logs page detail for a Gemini passthrough 404 now shows the upstream body in the error message |
…lity Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…pstream_error_body_logging Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> # Conflicts: # tests/integration/contracts.json
…e logger Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…gh error tests Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
bugbot run |
…ails Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
1 similar comment
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit c3c5676. Configure here.
…th (#44267) The failure spend-log row from #42695 is written for pass-through routes that run as LLM API routes, which a config route only does with auth: true. The test omitted auth and passed only while config wins (#41779) registered config entries through the typed model, where auth defaults to true. #43962 restored the pre-config-wins registration, so the route lost that status and the row was never written. Set auth: true on the route so the test covers the logging it was written for without depending on that side effect (cherry picked from commit 1d9cd9b)
…th (#44269) The failure spend-log row from #42695 is written for pass-through routes that run as LLM API routes, which a config route only does with auth: true. The test omitted auth and passed only while config wins (#41779) registered config entries through the typed model, where auth defaults to true. #43962 restored the pre-config-wins registration, so the route lost that status and the row was never written. Set auth: true on the route so the test covers the logging it was written for without depending on that side effect (cherry picked from commit 1d9cd9b)
…th (#44265) The failure spend-log row from #42695 is written for pass-through routes that run as LLM API routes, which a config route only does with auth: true. The test omitted auth and passed only while config wins (#41779) registered config entries through the typed model, where auth defaults to true. #43962 restored the pre-config-wins registration, so the route lost that status and the row was never written. Set auth: true on the route so the test covers the logging it was written for without depending on that side effect
TLDR
Problem this solves:
--detailed_debugHow it solves it:
redacted-by-litellmwhen message logging is turned off (globallitellm.turn_off_message_loggingor the key/teamturn_off_message_loggingcallback var), so an echoed prompt never lands in logs or callbacksUser Flow
Before: an operator debugging a failing Vertex passthrough call sees only a status code and cannot tell what the provider rejected
NOT_FOUNDbody saying the publisher model was not found--detailed_debugand find only the uvicorn access linePOST ... 404 Not Found, no provider message anywhere404: Upstream passthrough request failed with status 404After: the same request leaves the provider's own explanation in the log and in the request's error details
NOT_FOUNDbody, unchangedpass_through_endpoint: upstream POST https://aiplatform.googleapis.com/v1/projects/.../claude-nope-9:streamRawPredict returned 404: {"error": {"code": 404, "message": "Publisher model ... was not found ...", "status": "NOT_FOUND"}}, with any?key=query removed from the URL404: Upstream passthrough request failed with status 404: {"error": {... "was not found ..."}}Relevant issues
Supersedes #39687, which targeted the retired staging branch and had a merge conflict; the fix here was re-derived from the ticket on current main with the query-string redaction and the normalization guard added
Affected release
Linear ticket
Resolves LIT-6644
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Screenshots / Proof of Fix
Real Google APIs (Vertex AI and Gemini), no mocks. Proxy launched with
python -m litellm.proxy.proxy_cli --config <cfg> --port <port> --num_workers 1. Config:model_list: [], a generatedmaster_key($KEY),database_url: os.environ/DATABASE_URL,store_model_in_db: true,proxy_batch_write_at: 1, anddefault_vertex_configwithvertex_project: os.environ/VERTEXAI_PROJECT,vertex_location: global,vertex_credentials: os.environ/VERTEXAI_CREDENTIALS.$PROJECTis the Vertex project idBefore (40ec84c)
Vertex passthrough streamRawPredict, nonexistent model
curl -s -w "\nHTTP %{http_code}\n" http://localhost:4011/vertex_ai/v1/projects/$PROJECT/locations/global/publishers/anthropic/models/claude-nope-9:streamRawPredict -H "Authorization: Bearer $KEY" -H "content-type: application/json" -d '{"anthropic_version":"vertex-2023-10-16","max_tokens":16,"messages":[{"role":"user","content":"hi"}],"stream":true}'{ "error": { "code": 404, "message": "Publisher model `projects/$PROJECT/locations/global/publishers/anthropic/models/claude-nope-9` was not found or your project does not have access to it. Ensure you are using a valid model name and that the model is available in the specified region. For more information, see: https://docs.cloud.google.com/gemini-enterprise-agent-platform/resources/locations.", "status": "NOT_FOUND" } } HTTP 404grep -c "was not found" /tmp/before.log-> 0,grep "pass_through_endpoint: upstream" /tmp/before.log-> no output. Only the uvicorn access line mentions the 404Gemini passthrough 404
curl -s -w "\nHTTP %{http_code}\n" http://localhost:4011/gemini/v1beta/models/gemini-nope-9:generateContent -H "Authorization: Bearer $KEY" -H "x-goog-api-key: $KEY" -H "content-type: application/json" -d '{"contents":[{"role":"user","parts":[{"text":"hi"}]}]}'{ "error": { "code": 404, "message": "models/gemini-nope-9 is not found for API version v1beta, or is not supported for generateContent. Call ModelService.ListModels to see the list of available models and their supported methods.", "status": "NOT_FOUND" } } HTTP 404pass_through_endpoint: upstreamline, upstream body absentSELECT status, spend, metadata->'error_information' FROM "LiteLLM_SpendLogs" WHERE request_id='<x-litellm-call-id>':Gemini passthrough success (gemini-3.8-flash)
curl -s -w "\nHTTP %{http_code}\n" http://localhost:4011/gemini/v1beta/models/gemini-3.8-flash:generateContent -H "Authorization: Bearer $KEY" -H "x-goog-api-key: $KEY" -H "content-type: application/json" -d '{"contents":[{"role":"user","parts":[{"text":"hi"}]}]}'HTTP 200, normal candidates body ("Hello! How can I help you today?"),modelVersion: gemini-3.8-flashGemini passthrough 404 with message logging turned off for the key
Same request as the After leg below with a
turn_off_message_loggingkey. At the base the body is never logged in either mode, so the log carries no upstream line and the spend rowerror_messageis404: Upstream passthrough request failed with status 404, the same as the plain 404 caseGemini passthrough streaming 404 (streamGenerateContent?alt=sse)
curl -s -w "\nHTTP %{http_code}\n" "http://localhost:4011/gemini/v1beta/models/gemini-nope-9:streamGenerateContent?alt=sse" -H "Authorization: Bearer $KEY" -H "x-goog-api-key: $KEY" -H "content-type: application/json" -d '{"contents":[{"role":"user","parts":[{"text":"hi"}]}]}'NOT_FOUNDbody as the non-streaming case,HTTP 404, 275 bytesgrep -c "pass_through_endpoint: upstream"-> 0,grep -c NOT_FOUND-> 0After (a947ba7)
Vertex passthrough streamRawPredict, nonexistent model
curl -s -w "\nHTTP %{http_code}\n" http://localhost:4010/vertex_ai/v1/projects/$PROJECT/locations/global/publishers/anthropic/models/claude-nope-9:streamRawPredict -H "Authorization: Bearer $KEY" -H "content-type: application/json" -d '{"anthropic_version":"vertex-2023-10-16","max_tokens":16,"messages":[{"role":"user","content":"hi"}],"stream":true}'{ "error": { "code": 404, "message": "Publisher model `projects/$PROJECT/locations/global/publishers/anthropic/models/claude-nope-9` was not found or your project does not have access to it. Ensure you are using a valid model name and that the model is available in the specified region. For more information, see: https://docs.cloud.google.com/gemini-enterprise-agent-platform/resources/locations.", "status": "NOT_FOUND" } } HTTP 404Gemini passthrough 404
curl -s -w "\nHTTP %{http_code}\n" http://localhost:4010/gemini/v1beta/models/gemini-nope-9:generateContent -H "Authorization: Bearer $KEY" -H "x-goog-api-key: $KEY" -H "content-type: application/json" -d '{"contents":[{"role":"user","parts":[{"text":"hi"}]}]}'x-litellm-call-id: a0171a2b-5ff9-4548-863a-5511d4a946a8):{ "error": { "code": 404, "message": "models/gemini-nope-9 is not found for API version v1beta, or is not supported for generateContent. Call ModelService.ListModels to see the list of available models and their supported methods.", "status": "NOT_FOUND" } } HTTP 404?key=<GEMINI_API_KEY>upstream and the logged URL has no query (grep -c "key=" /tmp/after_tip.log-> 0):SELECT status, spend, metadata->'error_information' FROM "LiteLLM_SpendLogs" WHERE request_id='a0171a2b-5ff9-4548-863a-5511d4a946a8':Gemini passthrough success (gemini-3.8-flash)
curl -s -w "\nHTTP %{http_code}\n" http://localhost:4010/gemini/v1beta/models/gemini-3.8-flash:generateContent -H "Authorization: Bearer $KEY" -H "x-goog-api-key: $KEY" -H "content-type: application/json" -d '{"contents":[{"role":"user","parts":[{"text":"hi"}]}]}'HTTP 200, normal candidates body,modelVersion: gemini-3.8-flash, nopass_through_endpoint: upstreamWARNINGGemini passthrough 404 with message logging turned off for the key
Key generated with
curl -s http://localhost:4010/key/generate -H "Authorization: Bearer $KEY" -H "content-type: application/json" -d '{"metadata":{"logging":[{"callback_name":"prometheus","callback_type":"success_and_failure","callback_vars":{"turn_off_message_logging":"true"}}]}}', used below as$EKEYcurl -s -w "\nHTTP %{http_code}\n" http://localhost:4010/gemini/v1beta/models/gemini-nope-9:generateContent -H "Authorization: Bearer $EKEY" -H "x-goog-api-key: $EKEY" -H "content-type: application/json" -d '{"contents":[{"role":"user","parts":[{"text":"hi"}]}]}'NOT_FOUNDbody as above,HTTP 404, client response unchanged (x-litellm-call-id: e7845dc2-2ede-454a-955b-e65f3fb2f40a)Gemini passthrough streaming 404 (streamGenerateContent?alt=sse, prefix-replay path)
curl -s -w "\nHTTP %{http_code}\n" "http://localhost:4010/gemini/v1beta/models/gemini-nope-9:streamGenerateContent?alt=sse" -H "Authorization: Bearer $KEY" -H "x-goog-api-key: $KEY" -H "content-type: application/json" -d '{"contents":[{"role":"user","parts":[{"text":"hi"}]}]}'HTTP 404, 275 bytes,cmpagainst the Before response body reports no difference (x-litellm-call-id: c6c00e01-1310-45a2-9f75-835819942ece)error_messagecarries the same body,normalized_error: 500_UPSTREAM_PASSTHROUGHAdditional live A/B at base vs head: a scripted upstream serving a 6055-byte chunked 500 body through
streamGenerateContent?alt=ssedelivered all 6055 bytes to the client while the WARNING line ends with... (truncated at 4096 chars). A scripted upstream returning 400"no deployments available for this model"normalized to429_NO_HEALTHY_DEPLOYMENTSon head before the pattern reorder and to500_UPSTREAM_PASSTHROUGHafter itThe same seven live legs were rerun at the tip 23a7284 against the merge base e26a645 after main was merged in (main dropped
tests/integration/contracts.jsonand thecoverscheck, so the merge deletes the manifest and the tip strips the retired markers from the new tests): every client status and body is byte identical between base and tip (the 200 success leg differs only in generated content), and every tip WARNING line and spend row carries the body, theredacted-by-litellmplaceholder for the redacted key, or the 4096 marker for the 6055-byte body, exactly as shown aboveAfter the mid-read guard landed (8618621, 5d57d25 which only closes the upstream response in the new unit test, and c3c5676 which relays the decoded pieces rather than raw compressed bytes on a mid-read failure, so a gzip body cannot reach the client corrupted), the streaming legs C, F and CHUNK were rerun at c3c5676 against e26a645: byte identical to the client, head spend rows carry the body (
/home/ubuntu/scratch/lit6644_live_pr_risk_c3c56762e2.md). The mid-read failure path itself cannot be reached from a live provider, so it is covered by the mutation-checked unit test (removing the catch fails withhttpx.ReadError: peer reset, and the gzip variant fails on 5d57d25 because the client received compressed bytes). The audit head legs were also rerun at c3c5676, 57 nodes, all 20 observability cells green on both runs with no skips (audit_report_c3c56762e2.md); the base leg below still applies since no observability test file changedAdmin UI Logs page (head only, the Before rows carry no body to show)
adminwith the master keygemini-nope-9:generateContentrequest (x-litellm-call-id: 25bd8931-c638-48b0-8da8-518367358f26)Error Code: 404and the message now carries the upstream body:models/gemini-nope-9 is not found for API version v1beta ... "status": "NOT_FOUND"Audit matrix (deterministic tests/integration, tip 23a7284, head legs repeated at c3c5676)
Every row is a checked-in test in
tests/integration/observability/test_passthrough_upstream_error_visibility.pyortest_passthrough_upstream_error_chaos.py, run throughtests/integration/run.py extensionsagainst a real proxy with two workers, real Postgres, real Redis and thewire_serverscripted upstream, no mock inside the proxy, no real provider. Base leg is the merge base e26a645 with the tip's test tree copied in, the head legs are the tip run twice with an identical collected selection of 57 nodes and no skips. Body-visibility cells are red on base and green on both head legs, the cells the fix must not touch (H3, H4, N1) are green on all three legspass_through_endpointsroute 403 withmax budget reachedin the body and a query string on the target: body logged, query stripped, normalized_error stays 500_UPSTREAM_PASSTHROUGHturn_off_message_logging: true: WARNING and row showredacted-by-litellm, never the bodyType
🐛 Bug Fix
Caveats (if any)
Medium
content-encodingandcontent-lengthdropped, since the replayed bytes are already decoded; the client-facing headers already excluded both, so nothing visible changesx-litellm-enable-message-redactionrequest header does not turn redaction on for passthrough requests, because passthrough stores request headers underproxy_server_request.headers, whichshould_redact_message_loggingdoes not read. That is preexisting and applies to the whole passthrough logging path; the global setting and the key/teamturn_off_message_loggingcallback var do applyerror_messageat 2048 chars, so the tail of a long body is kept in the log line onlyLow
store_prompts_in_spend_logsis offFinal Attestation
ran /live-pr-risk and found no regressions/backward incompatible risks
REVIEWER MUST KNOW BEFORE APPROVING
Failure hook exception detail, spend log
error_messageand the Admin UI Logs error text: beforeUpstream passthrough request failed with status 404, after the same sentence followed by:and up to 4096 chars of the upstream body, or: redacted-by-litellmwhen message logging is off. Requested by LIT-6644, approved by yucheng-berri by assigning the ticketProxy WARNING log: before nothing for an upstream 4xx/5xx, after one
pass_through_endpoint: upstream <METHOD> <url without query> returned <status>: <body>line per failure. Requested by LIT-6644, approved by yucheng-berri by assigning the ticketStreaming upstream error whose body stream fails inside the first 4096 bytes (live A/B with a scripted upstream that sends 15 bytes of a 502 then resets, both legs run twice, transcripts in
/home/ubuntu/scratch/lit6644_midread_legs.md): before the client gets the 502 status line and then a reset,curl: (18) transfer closed with outstanding read data remaining, 0 body bytes, spend row with the bare status; after the client gets the 502 with the 15 bytes that arrived and a clean end of body, curl exit 0, spend row carrying those bytes and aread failed after 15 bytes: ReadErrorwarning. Raised by Bugbot on this PR, fixed in 8618621 and c3c5676, not yet approved by a human: yucheng-berri decides on this PRLink to Devin session: https://app.devin.ai/sessions/9943a4e7d8ef47638394734d69c2946e
Open in Devin Desktop: https://app.devin.ai/desktop/session/9943a4e7d8ef47638394734d69c2946e?variant=devin
Requested by: @yucheng-berri