Skip to content

fix(langsmith): json.dumps with default=str so non-serializable metadata does not crash batch flush - #42424

Merged
yucheng-berri merged 6 commits into
mainfrom
litellm_langsmith_json_safe
Sep 22, 2026
Merged

yucheng-berri merged 6 commits into
mainfrom
litellm_langsmith_json_safe

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Langsmith run metadata can carry datetime, Decimal and similar Python values
  • httpx json= cannot encode them, so the whole batch flush raises TypeError
  • The batch is dropped and nothing reaches Langsmith

How it solves it:

  • Serialize the batch with json.dumps(default=str, allow_nan=False) and send it as content=
  • allow_nan=False still rejects NaN and Infinity instead of emitting invalid JSON
  • The AsyncHTTPHandler retry path now forwards content=, so a retried batch re-sends the same body

Replaces #39133 (same five-file change, re-based on main; the original branch was cut from the retired staging branch and conflicts with main). The change was written by the original contributor and is credited with a Co-authored-by trailer

User Flow

Before: a developer with Langsmith logging enabled gets no runs when request metadata contains non-serializable values

  1. They configure Langsmith as a success callback in their proxy config
  2. They send a chat completion through the litellm SDK (or a custom logger) with metadata such as {"created_at": datetime(...), "spend": Decimal("0.0042")}
  3. The provider answers and the call returns normally
  4. The proxy log shows Langsmith Layer Error ... TypeError: Object of type datetime is not JSON serializable
  5. They open their Langsmith project and the run is missing

After: the same request lands in Langsmith with the values stored as strings

  1. They configure Langsmith as a success callback in their proxy config
  2. They send the same chat completion with the same metadata
  3. The provider answers and the call returns normally
  4. Langsmith accepts the batch with 202
  5. They open their Langsmith project and see the run with created_at: "2026-01-02 03:04:05+00:00" and spend: "0.0042"

Relevant issues

Supersedes #39133

Affected release

Linear ticket

Resolves LIT-8310

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Screenshots / Proof of Fix

Tests: unit regressions in tests/test_litellm/integrations/test_langsmith_init.py, tests/test_litellm/llms/custom_httpx/test_http_handler.py (retry forwards content=, red when the content= forwarding is removed at the post, patch and delete retry sites) and tests/logging_callback_tests/test_langsmith_unit_test.py, plus a live SDK e2e tests/e2e/logging/test_langsmith_batch_serialization_e2e.py against the real OpenAI and LangSmith EU APIs (fails at the poll deadline on 13b37487, passes at this tip). The bug is only reachable through the SDK path since the proxy JSON-decodes request metadata, which is why the e2e drives litellm.acompletion in-process

Rig: two live litellm proxies per leg (aiohttp + pure-httpx twins), each with a forwarding recorder in front of the real LangSmith EU API (https://eu.api.smith.langchain.com). Head 3f038e0a2f246c82b2e612c9fe1ea73f4cd201bb (PR tip) on :4120/:4122, base 13b374873d7b75f400fc1702c759d68195620f62 (origin/main) on :4121/:4123. Real gpt-4o-mini / claude-haiku-4-5 calls. Commits after 4dedd083 touch tests and the CI e2e lane selector only; litellm/ is identical across every head sha

Before (13b3748)

  1. N2 — SDK acompletion with langsmith callback, metadata holding datetime and Decimal:
PYTHONPATH=/home/ubuntu/repos/litellm_ls_mainbase LANGSMITH_BASE_URL=http://127.0.0.1:4221 LANGSMITH_PROJECT=litellm-lit8310-base-r2 python sdk_nonnative.py base N2
litellm.__file__ /home/ubuntu/repos/litellm_ls_mainbase/litellm/__init__.py
response_id chatcmpl-EQnGQSI81yZMKvnMSqq9P3CExhUJG run_id e658efeb-f8ce-4762-9f7b-be12bc066e25
queue_before_flush 1
queue_after_flush 1
LiteLLM:ERROR: langsmith.py:444 - Langsmith Layer Error -
    File ".../httpx/_content.py", line 177, in encode_json
      body = json_dumps(
    TypeError: Object of type datetime is not JSON serializable
  1. Recorder observed: no /api/v1/runs/batch request at all (batch died in json.dumps). GET https://eu.api.smith.langchain.com/api/v1/runs/e658efeb-f8ce-4762-9f7b-be12bc066e25 -> 404

  2. D — retry cell, drop_first on recorder then curl :4123 /v1/chat/completions:

echo drop_first > control-base2.txt
curl -s http://127.0.0.1:4123/v1/chat/completions -H "Authorization: Bearer sk-***" -d '{"model":"gpt-mini","messages":[...],"max_tokens":5}'
response 200, chatcmpl-EQnJDIvAyCBZGlgg4cXTFeujhzrpQ
wire: attempt 1 dropped_first_attempt; attempt 2 forwarded, same sha256 7ba6d994..., Content-Type application/json, 202
run 5775be51-6e8b-47f3-b4f7-3d890e8026cb, GET 200
  1. H1 POST :4121 /v1/chat/completions -> chatcmpl-EQnHZwwZXYcjt2y2gLFJGRdzKlo45, run 66442cf5-4d2a-4642-a17a-658b2695cbbc, GET 200
  2. H5 POST :4121 /v1/messages -> msg_011CfHvMoKxyB6Au3HH4XyZq, run 22460677-d871-4d2f-a2d1-1820f3aa1226, GET 200
  3. H8 POST :4121 /v1/responses -> resp_yMuoGk2jpxnK8ZGsdoxJxEaJDgyD..., run a4369bba-31f6-4fa3-83f6-c5e46db89fc1, GET 200

After (3f038e0, PR tip)

  1. N2, same command against the head tree:
PYTHONPATH=/home/ubuntu/repos/litellm_ls_main LANGSMITH_BASE_URL=http://127.0.0.1:4220 LANGSMITH_PROJECT=litellm-lit8310-head-r2 python sdk_nonnative.py head N2
litellm.__file__ /home/ubuntu/repos/litellm_ls_main/litellm/__init__.py
response_id chatcmpl-EQpNfxkticXm0UFiKyeQC7LB6B2NB run_id d7a7769c-1a4e-453f-aae5-8fa53a510ca4
  1. Recorder observed POST /api/v1/runs/batch -> 202, Content-Type: application/json, batch body extra.requester_metadata = {"cell": "N2", "created_at": "2026-01-02 03:04:05+00:00", "spend": "0.0042"} (non-native values serialized as their str() form). GET https://eu.api.smith.langchain.com/api/v1/runs/d7a7769c-1a4e-453f-aae5-8fa53a510ca4 -> 200

  2. N3, same command with float("nan") in metadata: ValueError: Out of range float values are not JSON compliant: nan, batch dropped, nothing on the wire (allow_nan=False, same as base)

  3. D, retry cell, drop_first on recorder then curl :4122 /v1/chat/completions:

echo drop_first > control-head2.txt
curl -s http://127.0.0.1:4122/v1/chat/completions -H "Authorization: Bearer sk-***" -d '{"model":"gpt-mini","messages":[...],"max_tokens":5,"metadata":{"audit_cell":"D-r4"}}'
response 200, chatcmpl-EQpNxEgAQDxxpLH6Kini6APA2dqvM
wire: attempt 1 dropped_first_attempt; attempt 2 forwarded, same sha256 62ecf2b9282e78c0, Content-Type application/json, 202
run b8c3b7fb-ef44-47a3-ad98-3041130de140, GET 200
  1. H1 POST :4120 /v1/chat/completions -> chatcmpl-EQpQFBST4glkSJrPyamyvuSlcKSm1, run b565c3e8-b1cb-4806-bd39-8d9259886997, GET 200
  2. H5 POST :4120 /v1/messages -> msg_011CfJ6p4sFZ9avXEGGU9PBA, run dada6664-9f12-4527-958e-5004c01fcb78, GET 200
  3. H8 POST :4120 /v1/responses -> resp_9ThGPQ6CAIkSpb0LEUDy_IBDp8ACpcfUJ_B, run 45d1b3cc-f145-45c1-bbb4-e01c54ce7279, GET 200
  4. pytest tests/e2e/logging/test_langsmith_batch_serialization_e2e.py against real OpenAI and LangSmith EU: 1 passed, run 893cadbf-1114-4dd8-8a7d-542c942787fe read back with the str() forms of the datetime and Decimal (the same file copied onto 13b37487 fails at the poll deadline, TypeError: Object of type datetime is not JSON serializable in the log)

The full live audit (happy-path matrix on chat, messages and responses, streaming and non-streaming, OpenAI and Anthropic SDKs and curl, sad paths, key and team level LangSmith settings, idempotency and chaos cells) is in the comment below

Screen recording (before on 13b37487, after on 5a877c55; the only later commit 3f038e0a changes .github/e2e-stack/select_tests.py, nothing under litellm/ or tests/) is posted in the requesting Slack thread, screenshots in the PR comment: base N2 chatcmpl-EQoDbXRryr7KkEpqYJ20t4Fb3n9eV raised the TypeError, run 11b42a8b-9607-4a71-bf48-7768da4e025e GET 404; tip N2 chatcmpl-EQoET2g6714tt3LfOg3hW9gpOVxVG batch 202, run 4503fe51-3e37-4fc3-a1cf-b580846c4274 GET 200; retry D chatcmpl-EQoGmKUG6tiXsfuF7ZihRhVd9VWJu, attempt 2 identical body sha256 62729496e6f91c5e9f0559cd7f3dd134ae97132b9753ab214b8c1f2fa82a01a1, run f78a9633-e0fd-48d7-827a-1d4352df71aa GET 200

Type

🐛 Bug Fix

Caveats (if any)

Low

  • Non-serializable field values are coerced to str; the logged payload type differs from the original Python type
  • Retry with the same body only applies on the pure httpx transport; under the default aiohttp transport a dropped connection surfaces as httpx.ReadError, which the handler does not retry (same on main)
  • Not part of this fix, seen identically on both legs in the previous audit: queued events are lost on destination outage, worker kill or proxy restart

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Link to Devin session: https://app.devin.ai/sessions/63688ea4e4e941ef82072628a246d466
Open in Devin Desktop: https://app.devin.ai/desktop/session/63688ea4e4e941ef82072628a246d466?variant=devin
Requested by: @yucheng-berri

…ata does not crash batch flush

Serialize the runs/batch payload with json.dumps(default=str, allow_nan=False) and send it as content= with an explicit Content-Type, so datetime, Decimal and similar metadata values no longer raise TypeError and drop the batch. Forward content= on the AsyncHTTPHandler retry path so a retried batch re-sends the identical body

Replaces #39133, which was cut from the retired staging branch and conflicts with main

Co-authored-by: Damien Smrt <dsmrt@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@greptile-apps

greptile-apps Bot commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

The PR appears safe to merge; no actionable new issue was introduced since the previous review, and all previous findings are resolved.

Summary

This PR prevents LangSmith batch flushes from failing when metadata contains non-JSON-native Python values and preserves raw request bodies across HTTP retry clients.

  • Serializes LangSmith batches with default=str while continuing to reject NaN and Infinity.
  • Explicitly sends the serialized payload as JSON content with the appropriate content type.
  • Forwards content through POST, PUT, PATCH, and DELETE retry paths.
  • Adds unit and live-E2E coverage for serialization, headers, tenant IDs, errors, and retries.
  • The change since the previous review only excludes the credential-dependent live LangSmith test from the stage-mirror changed-E2E selector.

Reviews (4) · Last reviewed commit: "test(e2e): deselect the LangSmith live e..."

Comment thread tests/test_litellm/integrations/test_langsmith_init.py Outdated
Comment thread tests/test_litellm/llms/custom_httpx/test_http_handler.py Outdated
@codspeed

codspeed Bot commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_langsmith_json_safe (3f038e0) with main (25af172)

Open in CodSpeed

@codecov

codecov Bot commented Sep 22, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 80.00000% with 1 line in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
litellm/llms/custom_httpx/http_handler.py 66.66% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

…client-injecting handler

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

Comment thread tests/test_litellm/llms/custom_httpx/test_http_handler.py
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

Live before/after on tip 5a877c5 vs main 13b3748, real OpenAI and LangSmith EU; recording is in the Slack thread

Before (13b3748) After (5a877c5)
base: TypeError and LangSmith 404 tip: batch 202, run GET 200
pinned revisions LangSmith UI metadata as strings
retry: identical body hash, 202

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 3f038e0. Configure here.

@yucheng-berri
yucheng-berri merged commit 2ef710e into main Sep 22, 2026
94 of 96 checks passed
@yucheng-berri
yucheng-berri deleted the litellm_langsmith_json_safe branch September 22, 2026 19:21

This branch was successfully deployed

1 active deployment
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants