Skip to content

fix(vertex): stop O(n^2) re-parse of accumulated Gemini stream JSON - #31297

Merged
yassin-berriai merged 1 commit into
litellm_internal_stagingfrom
litellm_fix_gemini_stream_quadratic_json
Jun 26, 2026
Merged

yassin-berriai merged 1 commit into
litellm_internal_stagingfrom
litellm_fix_gemini_stream_quadratic_json

Conversation

@yassin-berriai

@yassin-berriai yassin-berriai commented Jun 25, 2026 •

Copy link
Copy Markdown
Contributor

Relevant issues

Fixes #26181

Linear ticket

Resolves LIT-3503 (the large-streaming-payload half; the mid-stream 429 half is handled in a separate PR)

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • My PR passes all unit tests on make test-unit
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have requested a Greptile review by commenting @greptileai and received a Confidence Score of at least 4/5 before requesting a maintainer review (got 5/5)

Type

🐛 Bug Fix

Changes

ModelResponseIterator.handle_accumulated_json_chunk reassembles a Gemini streaming chunk that arrived fragmented across SSE events. It appended each fragment to self.accumulated_json and then called json.loads on the whole buffer after every fragment. When a single value is split across many fragments that is O(n^2) total work, and each json.loads is one CPython C call that holds the GIL, so a large enough response blocks the asyncio event loop. In the default single-process uvicorn proxy that freezes every handler, the k8s liveness probe times out, and the pod is killed and restarted. This is the behavior the py-spy dumps in #26181 captured (MainThread pinned in json.loads called from handle_accumulated_json_chunk)

A complete Gemini stream value is a JSON object or array, so the buffer can only become parseable once its last non-whitespace byte can close one. The fix gates the json.loads attempt on that, so the common fragmented-response case parses roughly once instead of once per fragment. json.loads stays the correctness oracle (a premature attempt simply raises and keeps accumulating); the gate only decides when it is worth attempting, so output is unchanged

Scope is the Vertex/Gemini path that LIT-3503's evidence points at. The same copy-paste pattern exists in the Anthropic and SageMaker iterators and can follow in separate PRs

Screenshots / Proof of Fix

Live proxy, one model pointing at a local mock that streams a 3 MB Gemini response fragmented into 64-byte SSE data lines (the accumulate-growth trigger). A background poller hits /health/liveliness every 100 ms for 30 s while one streaming request runs, so the worst single probe latency is how long the event loop was frozen

Same request, same assembled output (3,145,728 chars), only the proxy code differs:

Before (current litellm_internal_staging):

[unfixed] stream_wall_time=32.49s  probes=14 worst_liveness_probe_latency=12299 ms
[assembled] content_chars=3145728 all_x=True

After (this PR):

[fixed] stream_wall_time=3.49s  probes=281 worst_liveness_probe_latency=8 ms
[assembled] content_chars=3145728 all_x=True

The unfixed event loop is frozen for 12.3 s in a single stretch, so only 14 liveness probes get answered in 30 s and any probe with a sub-12 s timeout fails. With the fix the worst probe is 8 ms and the same content streams through correctly

Driving the exact production method directly (so the quadratic is unambiguous), event-loop-blocked time by payload size, 2 KB fragments:

# before
payload=2MB  event-loop-blocked=  221 ms
payload=4MB  event-loop-blocked= 1754 ms
payload=8MB  event-loop-blocked= 6859 ms     # doubling size ~4x's the freeze -> O(n^2)
# after
payload=2MB  event-loop-blocked=   20 ms
payload=4MB  event-loop-blocked=  105 ms
payload=8MB  event-loop-blocked=  325 ms

Tests (the regression test fails on current code and passes with the fix):

tests/test_litellm/llms/vertex_ai/gemini/test_vertex_and_google_ai_studio_gemini.py
  test_accumulated_json_does_not_reparse_every_fragment
  test_accumulated_json_partial_fragment_returns_none_without_parsing
127 passed

@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

@yassin-berriai

Copy link
Copy Markdown
Contributor Author

@greptileai

@greptile-apps

greptile-apps Bot commented Jun 25, 2026

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR fixes an O(n²) performance regression in ModelResponseIterator.handle_accumulated_json_chunk where json.loads was called on the entire accumulated buffer after every SSE fragment. For large Gemini responses fragmented into small chunks, this held the GIL long enough to freeze the asyncio event loop and trigger k8s liveness-probe timeouts.

  • Core fix: Gates the json.loads attempt on a cheap heuristic — only try when the buffer's last non-whitespace character is } or ] (i.e., a JSON object or array could be closed). For a 3 MB response in 64-byte fragments this reduces parse attempts from ~50 000 to ~1.
  • Tests added: Two new unit tests — a regression test that spies on json.loads call count across many fragments, and a guard test that verifies a partial fragment triggers zero parse calls. No existing tests are modified.

Confidence Score: 5/5

Safe to merge — the change is a six-line, narrowly-scoped guard in a single method, the correctness oracle (json.loads + JSONDecodeError fallback) is preserved, and two new regression tests directly cover the fixed path.

The heuristic gate is correct for the Gemini streaming format (responses are always JSON objects or arrays). The rare edge case where a mid-stream fragment ends with } or ] inside a string value is still handled safely by the existing JSONDecodeError catch. No existing behaviour is changed for callers that receive well-formed fragments.

No files require special attention.

Important Files Changed

Filename Overview
litellm/llms/vertex_ai/gemini/vertex_and_google_ai_studio_gemini.py Six-line change in handle_accumulated_json_chunk: adds a heuristic early-exit before json.loads that checks the buffer's last non-whitespace byte. Logic is correct for all Gemini response shapes (always JSON objects/arrays); edge cases where a string value happens to end with } or ] mid-stream will still call json.loads but will get a JSONDecodeError and continue accumulating — unchanged from before for those rare cases.
tests/test_litellm/llms/vertex_ai/gemini/test_vertex_and_google_ai_studio_gemini.py Two new mock-only tests added: one regression test that uses patch("json.loads", wraps=json.loads) to assert parse calls ≤ 2 across 50+ fragments of a 200k-char payload, and one that asserts zero parse calls for a clearly partial fragment. No existing tests are modified. Tests use only mocks and are appropriate for the unit test folder.

Reviews (1): Last reviewed commit: "fix(vertex): stop O(n^2) re-parse of acc..." | Re-trigger Greptile

@greptile-apps

greptile-apps Bot commented Jun 25, 2026

Copy link
Copy Markdown
Contributor

Greptile Summary

Fixes an O(n²) GIL-holding json.loads loop in ModelResponseIterator.handle_accumulated_json_chunk that caused the asyncio event loop to freeze for seconds (up to 12 s in the benchmark) when Gemini streaming responses were fragmented across many SSE events. The fix adds a cheap last-byte guard — only attempt json.loads when the buffer's trailing non-whitespace character is } or ] — so the parse is attempted roughly once per complete chunk instead of once per fragment.

  • Production change (vertex_and_google_ai_studio_gemini.py): Six-line guard in handle_accumulated_json_chunk reduces event-loop-blocked time from ~12 s to ~8 ms in the provided benchmark for a 3 MB fragmented response; json.loads remains the correctness oracle and all existing semantics are preserved (an early-exit on a false-positive } or ] inside a string value still falls back via JSONDecodeError).
  • Tests (test_vertex_and_google_ai_studio_gemini.py): Two new mock-only regression tests — one asserts json.loads is called ≤ 2 times for a 200 KB multi-fragment payload (would fail against the unfixed code), and one verifies that a partial fragment never triggers a parse attempt.

Confidence Score: 5/5

Safe to merge — the change is a targeted, narrow optimization in a single method with no semantic side-effects and verified benchmark and regression-test evidence.

The guard is strictly additive: it skips json.loads only when the buffer cannot possibly be complete JSON, and any false-positive (a } or ] inside a string value) still falls through to json.loads and the existing JSONDecodeError path. The StopIteration flush path in both __next__ and __anext__ is unaffected because a fully-assembled JSON object always ends with } or ]. New tests exercise both the happy path and the early-exit path with mocks, and the benchmark numbers in the description confirm the fix works end-to-end.

No files require special attention.

Important Files Changed

Filename Overview
litellm/llms/vertex_ai/gemini/vertex_and_google_ai_studio_gemini.py Adds an early-exit guard in handle_accumulated_json_chunk that skips json.loads unless the buffer's last non-whitespace byte is } or ], eliminating the O(n²) re-parse that froze the asyncio event loop on large Gemini streaming responses.
tests/test_litellm/llms/vertex_ai/gemini/test_vertex_and_google_ai_studio_gemini.py Adds two mock-based regression tests: one verifying json.loads is called at most twice for a multi-fragment 200 KB payload, and one confirming an incomplete fragment never triggers a parse attempt; both correctly use MagicMock with no real network calls.

Reviews (2): Last reviewed commit: "fix(vertex): stop O(n^2) re-parse of acc..." | Re-trigger Greptile

@codecov

codecov Bot commented Jun 25, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@yassin-berriai
yassin-berriai enabled auto-merge (squash) June 25, 2026 15:29
handle_accumulated_json_chunk re-ran json.loads on the entire accumulated
buffer after every fragment. For a streaming response fragmented across many
chunks that is O(n^2) total work in a single GIL-holding C call, so a large
enough Gemini response freezes the asyncio event loop for seconds, liveness
probes time out, and the proxy pod gets killed and restarted.

A complete Gemini stream value is a JSON object or array, so the buffer can
only become parseable once its last non-whitespace byte can close one. Gate
the json.loads attempt on that, which makes the common fragmented-response
case parse roughly once instead of once per fragment. An 8MB payload drops
from a 6.9s event-loop freeze to ~0.3s with identical parsed output.

Resolves LIT-3503
Fixes #26181
@yassin-berriai
yassin-berriai force-pushed the litellm_fix_gemini_stream_quadratic_json branch from 0dd04e3 to 9a5127b Compare June 26, 2026 05:50
@yassin-berriai
yassin-berriai merged commit 29c254d into litellm_internal_staging Jun 26, 2026
122 checks passed
@yassin-berriai
yassin-berriai deleted the litellm_fix_gemini_stream_quadratic_json branch June 26, 2026 06:04
# chunk is a JSON object/array, so only attempt the parse once the
# buffer's last non-whitespace byte can close one.
stripped = self.accumulated_json.rstrip()
if not stripped or stripped[-1] not in "}]":

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Medium: Stream parsing denial of service

This gate treats any trailing } or ] as a possible JSON terminator, including those inside a quoted model response or tool argument. An authenticated user can request a large output consisting of repeated closing brackets; when that JSON is fragmented, each prefix passes this check and json.loads rescans the entire growing buffer, allowing the request to stall the shared event loop. Track JSON string/escape and nesting state incrementally, or use a bounded incremental parser rather than relying on the final character.

@veria-ai

veria-ai Bot commented Jul 14, 2026

Copy link
Copy Markdown
Contributor

PR overview

This pull request updates the Vertex AI / Google AI Studio Gemini streaming code to avoid repeatedly reparsing an accumulated JSON buffer while handling streamed Gemini responses. It focuses on improving the stream JSON parsing path for performance during incremental response assembly.

There is still an open availability concern in the updated stream parsing logic: certain fragmented streamed outputs containing repeated closing brackets can cause repeated full-buffer JSON parsing. An authenticated user who can trigger large Gemini streaming responses may be able to stall the shared event loop, so the current security posture still has a meaningful denial-of-service risk despite the intended performance fix.

Open issues (1)

Fixed/addressed: 0 · PR risk: 5/10

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: O(n²) json.loads retry in handle_accumulated_json_chunk blocks event loop, kills liveness probes

3 participants