Repository navigation
test(integration): regression tests for July provider translation, routing and streaming bugs - #42693
Conversation
…n the OpenAI Responses wire (Pylon #6619) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…lds are reported and charged (Pylon #6708) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…imeout (Pylon #6505) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ody reaches the provider (Pylon #6645) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ume the role once (Pylon #6681) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ses wire with always_include_stream_usage (Pylon #6466) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
… instead of 400 (Pylon #6539) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…/v1/responses (Pylon #6565) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…patible providers as stop (Pylon #6536) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ned and round-trip through /v1/responses (Pylon #6409) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…v1/messages (Pylon #6727) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…s top_n without sending top_k (Pylon #6401) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ded tokens reach the model tpm (Pylon #6344) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…lock at index zero on /v1/messages streams (Pylon #6337) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ream finishes sending (Pylon #6315) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…eached with bearer auth (Pylon #6249) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…o is null (Pylon #6374) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
… and cached tokens land in spend log metadata (Pylon #6220) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ecutor deployment (Pylon #6212) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ntent with Anthropic ttl (Pylon #6221) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ed before reaching Mantle (Pylon #6262) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ries without thinking blocks (Pylon #6222) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…lon #6221) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…d keepalive timeout (Pylon #6387) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…r and sends hf_model_name as the body model (Pylon #6187) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
… and sends V3 reasoning_effort raw (Pylon #6149) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…fore the provider call (Pylon #6075) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
… a stalled upstream (Pylon #6025) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…antle as top-level tools (Pylon #6012) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…roxy Anthropic key to the caller host (Pylon #6226) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…_pylon_regression_tests_jul_providers # Conflicts: # tests/integration/contracts.json
…_pylon_regression_tests_jul_providers # Conflicts: # tests/integration/contracts.json
…_pylon_regression_tests_jul_providers # Conflicts: # tests/integration/contracts.json
…_pylon_regression_tests_jul_providers # Conflicts: # tests/integration/contracts.json
…_pylon_regression_tests_jul_providers # Conflicts: # tests/integration/contracts.json
…_pylon_regression_tests_jul_providers # Conflicts: # tests/integration/contracts.json
…_pylon_regression_tests_jul_providers # Conflicts: # tests/integration/contracts.json
…_pylon_regression_tests_jul_providers Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> # Conflicts: # tests/integration/contracts.json # tests/integration/providers/test_bedrock_mantle_wire.py
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…under cache and worker sharing Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…eal retention regression check Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".
|
|
|
|
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
…isor executor Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ion_tests_jul_providers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit a5e008a. Configure here.
TLDR
Problem this solves:
How it solves it:
User Flow
Before: an engineer refactors a provider bridge and reintroduces one of the July regressions, for example a messages deployment on a Mantle backend sending max_tokens=1 to the provider
"max_tokens": 1and gets a provider error backAfter: the same refactor turns the providers integration shard red before merge
integration-providersshard fails ontest_messages_over_responses_deployment_with_max_tokens_1_is_clamped_to_16_instead_of_400, showing the unclampedmax_tokensin the outbound bodyRelevant issues
Regression coverage for July customer bug reports, one Pylon ticket per test, listed in the table below. No open issue is fixed by this PR
Affected release
Linear ticket
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
This PR ships tests only, so the proof is the mutation check tests/AGENTS.md asks for: each test was run with its fix commit reverted (red) and with the fix restored (green), through the same real proxy, Postgres, Redis and scripted upstream the CircleCI shard uses. The harness is deterministic by design and does not call real providers. The providers shard passed locally on the tip with two different order seeds, and the extensions and sdk shards passed on both seeds too
Before (fix commit reverted on top of f238aeb)
Per test, the fix sha in the table was reverted (
git revert --no-commit <sha>, or a worktree at the fix parent where the revert conflicted), the proxy restarted, and the new node run. Every node failed on the behavioral assertion. Output is in the fold-outs belowAfter (f238aeb)
Fix restored, same node rerun, passed. The shards:
python tests/integration/run.py providerspasses underINTEGRATION_ORDER_SEED=1and again underINTEGRATION_ORDER_SEED=2, plusextensionsandsdkon both seedsTests and their red/green evidence
37b659e864tests/integration/providers/test_anthropic_legacy_thinking_budget_wire.py::test_claude_4_6_thinking_budget_tokens_on_messages_is_forwarded_instead_of_rewritten_to_adaptive8e524370e1tests/integration/providers/test_bedrock_invoke_cache_usage_wire.py::test_nova_invoke_count_suffixed_cache_usage_fields_are_reported_and_charged14dd98cd5ftests/integration/providers/test_bedrock_role_configuration.py::test_repeat_requests_under_one_session_name_assume_role_once_per_session_nameff71808671tests/integration/providers/test_bedrock_converse_client_metadata_wire.py::test_client_metadata_is_dropped_from_converse_body_while_anthropic_beta_is_kept47eaff19d3tests/integration/providers/test_anthropic_messages_openai_tools_wire.py::test_messages_tool_with_optional_properties_reaches_openai_responses_non_strictf60e99c583tests/integration/providers/test_responses_client_header_forwarding_wire.py::test_client_x_header_is_forwarded_to_the_provider_on_responses0c376d8963tests/integration/providers/test_responses_bridge_incomplete.py::test_messages_over_responses_deployment_with_max_tokens_1_is_clamped_to_16_instead_of_40032a4377acdtests/integration/providers/test_anthropic_messages_fireworks_stop_wire.py::test_messages_stop_sequences_to_fireworks_are_sent_as_stop_not_stop_sequences9121ae3024tests/integration/providers/test_anthropic_wire.py::test_anthropic_messages_slow_upstream_is_cut_off_at_the_deployment_request_timeoutcfb7edb54etests/integration/providers/test_responses_bridge_stream_options.py::test_messages_stream_with_always_include_stream_usage_omits_include_usage_from_responses_request1a8cd8a078tests/integration/providers/test_anthropic_messages_openai_bridge_wire.py::test_midturn_system_correction_is_forwarded_to_openai_responsesf8caaf4d2dtests/integration/providers/test_responses_bridge_namespace_tools.py::test_codex_namespace_tool_is_flattened_for_chat_upstream_and_restored_in_responses_outputf64479e74dtests/integration/providers/test_nvidia_nim_ranking_wire.py::test_nvidia_nim_ranking_keeps_image_passages_and_applies_top_n_without_sending_top_kcaede1c5a0tests/integration/sdk/test_aiohttp_session_rebuild_wire.py::test_rebuilt_shared_session_drops_idle_connection_after_configured_keepalive_timeout61d32c9aactests/integration/providers/test_vertex_batch_output_info_wire.py::test_vertex_batch_create_survives_explicit_null_output_infoe204e629e0tests/integration/routing/test_priority_model_tpm_enforcement.py::test_tpm_only_model_returns_429_to_priority_key_once_recorded_tokens_reach_model_tpm53a206a179, 2f7574d7c1tests/integration/streaming/test_stream_contracts.py::test_messages_stream_opens_thinking_block_at_index_zero_for_reasoning_content_only_chunksc70a3c7093tests/integration/streaming/test_file_content_streaming.py::test_file_content_streams_the_first_megabyte_to_the_client_before_the_upstream_sends_the_rest0c376d8963tests/integration/providers/test_bedrock_mantle_responses_wire.py::test_max_output_tokens_below_mantle_minimum_is_raised_to_16_before_reaching_mantle88c9dd1294tests/integration/compatibility/test_a2a_wire_versions.py::test_agent_serving_its_card_only_at_versioned_path_is_reached_with_bearer_and_answerse59add11cdtests/integration/providers/test_anthropic_thinking_signature_retry_wire.py::test_missing_thinking_signature_400_retries_once_without_thinking_blocks_and_returns_2005fc510a6fdtests/integration/providers/test_gemini_messages_cache_control_wire.py::test_gemini_messages_cache_control_creates_cached_content_and_generates_from_it07b9ea8c3btests/integration/providers/test_anthropic_advisor_wire.py::test_advisor_api_base_without_api_key_is_rejected_before_the_proxy_anthropic_key_reaches_the_caller_hostb3d05bd10btests/integration/providers/test_fireworks_ai_session_affinity_wire.py::test_fireworks_session_id_sends_affinity_header_and_logs_cache_read_tokens96f58fac53tests/integration/routing/test_advisor_failure_cooldown.py::test_advisor_sub_call_401_leaves_the_executor_deployment_serving_the_next_requeste17988f4fetests/integration/providers/test_sagemaker_chat_wire.py::test_sagemaker_chat_signs_the_inference_component_header_and_sends_hf_model_name_as_the_body_model2611f6420atests/integration/providers/test_bedrock_deepseek_reasoning_wire.py::test_deepseek_r1_thinking_and_reasoning_effort_are_dropped_instead_of_leaking_into_converse2611f6420atests/integration/providers/test_bedrock_deepseek_reasoning_wire.py::test_deepseek_v3_reasoning_effort_reaches_converse_raw_instead_of_as_anthropic_thinking950074eea2tests/integration/routing/test_team_model_tpm_limit.py::test_concurrent_team_model_tpm_requests_reserve_tokens_before_reaching_the_provider9121ae3024tests/integration/providers/test_anthropic_messages_timeout_wire.py::test_messages_endpoint_honors_configured_timeout_against_stalled_upstream0c376d8963tests/integration/providers/test_responses_bridge_incomplete.py::test_messages_over_responses_deployment_with_max_tokens_one_reaches_openai_as_sixteen07a355e867tests/integration/providers/test_bedrock_mantle_codex_input_wire.py::test_codex_additional_tools_input_item_reaches_mantle_as_top_level_toolsbb22742025tests/integration/providers/test_rerank_latency_headers_wire.py::test_rerank_response_carries_call_id_latency_and_cost_headers_like_chat_completions2b33201a09tests/integration/providers/test_bedrock_knowledge_base_user_context_wire.py::test_vector_store_search_user_context_reaches_bedrock_retrieve_body568c5713eftests/integration/providers/test_bedrock_marengo_embed_3_wire.py::test_marengo_3_text_embedding_nests_input_text_under_input_typec75fccfd63tests/integration/providers/test_bedrock_mantle_wire.py::test_chat_completions_bridge_signs_mantle_responses_request_with_deployment_aws_keys29c254d3d3tests/integration/providers/test_vertex_gemini_fragmented_stream_wire.py::test_vertex_gemini_stream_split_across_many_fragments_completes_without_stalling2c1d62ce2btests/integration/routing/test_priority_rate_limit_headers.py::test_streaming_chat_completion_success_logs_v3_rate_limit_remaining_values_for_callbacks950074eea2tests/integration/routing/test_key_tpm_reservation.py::test_concurrent_requests_over_key_tpm_are_rejected_before_reaching_provider2967bc9beftests/integration/providers/test_websearch_interception_wire.py::test_database_created_search_tool_backend_receives_the_intercepted_query_over_a_same_named_config_tool90440d75aetests/integration/providers/test_bedrock_mantle_wire.py::test_bedrock_mantle_messages_stream_relays_anthropic_sse_instead_of_failing_on_event_stream_decodePylon 6727: On POST /v1/messages through the Anthropic SDK, a request for a Claude 4.6 model such as claude-sonnet-4-6 with max_token [...]
Red method: worktree_at_fix_parent_then_cherry_pick
Red (fix reverted):
Green (fix present):
Pylon 6708: A proxy client sends POST /v1/chat/completions to a deployment such as bedrock/invoke/us.amazon.nova-pro-v1:0 with a norm [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 6681: POST /v1/chat/completions against a Bedrock model configured with aws_role_name plus aws_session_name (and aws_sts_endpoi [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 6645: PR #35967 fixes POST /v1/responses requests using a Bedrock Converse deployment when the request includes client_metadata [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 6619: The customer sends POST /v1/messages using an Anthropic Messages tool with input_schema properties such as city, unit, an [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 6565: With forward_client_headers_to_llm_api enabled globally or for the model group, a client POST to /v1/responses containing [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 6539: A client sends POST /v1/messages to the proxy for an OpenAI or Azure Responses deployment with model configured for the G [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 6536: A caller sends POST /v1/messages to the proxy for a non-Anthropic OpenAI-compatible deployment such as fireworks_ai with [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 6505: The customer-facing route is Anthropic Messages POST /v1/messages, including Anthropic SDK messages.create calls. Configu [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 6466: POST /v1/messages with an OpenAI Responses-backed model and streaming enabled causes always_include_stream_usage to injec [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 6449: The customer-facing call is POST /v1/messages, typically through the Anthropic SDK, using an OpenAI-backed deployment suc [...]
Red method: worktree_at_fix_parent_then_cherry_pick
Red (fix reverted):
Green (fix present):
Pylon 6409: The proxy route is POST /v1/responses for a Responses API request sent by Codex with tools containing a type=namespace it [...]
Red method: worktree_at_fix_parent_then_cherry_pick
Red (fix reverted):
Green (fix present):
Pylon 6401: POST /v1/rerank through a deployment using nvidia_nim/ranking/ with query text, an image passage such as {"image": [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 6387: The reported SDK call is AsyncAzureOpenAI.embeddings.create(model="text-embedding-ada-002", input="warm-up n"); the equiv [...]
Red method: worktree_at_fix_parent_then_cherry_pick
Red (fix reverted):
Green (fix present):
Pylon 6374: Exercise the OpenAI-compatible proxy routes POST /v1/batches and GET /v1/batches/{batch_id}, or the equivalent SDK calls [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 6344: With dynamic_rate_limiter_v3, configure a model deployment with model-level tpm and priority_reservation, then send POST [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 6337: Send POST /v1/messages with Authorization: Bearer , stream: true, a user message, and a model routed to an O [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 6315: An authenticated client sends GET /v1/files/{file_id}/content (no body; e.g. Authorization: Bearer ). LiteLL [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 6262: For a proxy deployment using the bedrock_mantle provider and Responses API, a request with max_output_tokens below 16, su [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 6249: The regression is observable through POST /a2a/{agent_id} for a registered agent and POST /v1/chat/completions with model [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 6222: The affected flow is the proxy's native POST /v1/messages route, also reachable through an Anthropic SDK client configure [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 6221: POST /v1/messages with a Gemini model and a system or content text block carrying cache_control was translated through th [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 6226: POST /v1/messages on the proxy with a tools[] entry {"type": "advisor_20260301", "model": , "api_base":
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 6220: For proxy POST /v1/chat/completions to a fireworks_ai deployment, and the Anthropic-compatible POST /v1/messages adapter [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 6212: The proxy route is POST /v1/messages using an Anthropic advisor tool. Configure one model group such as claude-sonnet-5 w [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 6187: POST /v1/chat/completions through the proxy using model=sagemaker_chat/, with model_id= an [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 6149: Merged PR #33409 fixes Bedrock Converse DeepSeek reasoning requests. Through the proxy chat completions route, and the an [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 6075: For a proxy chat-completions request using a key attached to a team whose team metadata configures model_tpm_limit for th [...]
Red method: worktree_at_fix_parent_then_cherry_pick
Red (fix reverted):
Green (fix present):
Pylon 6025: Proxy POST /v1/messages (and the SDK anthropic_messages call) ignored the configured timeout: the outbound httpx POST to [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 6008: POST /v1/messages (Anthropic Messages API, used by Claude Code /model) or /v1/chat/completions with max_tokens or max_com [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 6012: The /v1/responses proxy route for a Bedrock Mantle GPT-5.6 model receives a Codex responses-lite request whose input cont [...]
Red method: worktree_at_fix_parent_then_cherry_pick
Red (fix reverted):
Green (fix present):
Pylon 5981: POST /v1/rerank on the proxy (model bedrock/cohere.rerank-v3-5:0 or any scripted rerank upstream) returned 200 with only [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 5991: The customer-facing call is POST /v1/vector_stores/{vector_store_id}/search, including an OpenAI Python SDK call such as [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 5949: Proxy route POST /v1/embeddings for a model configured as bedrock/us.twelvelabs.marengo-embed-3-0-v1:0 (or the eu./base i [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 5870: POST /v1/chat/completions on the proxy against a bedrock_mantle/* deployment (e.g. bedrock_mantle/openai.gpt-5.5) whose l [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 5838: A caller sends POST /v1/chat/completions with a Vertex Gemini deployment and stream=true. LiteLLM sends the translated Ve [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Pylon 5767: Proxy request POST /v1/chat/completions (or /v1/messages) with stream=true and the v3 rate limiter enabled (litellm_setti [...]
Red method: worktree_at_fix_parent_then_cherry_pick
Red (fix reverted):
Green (fix present):
Pylon 5737: For POST /v1/chat/completions through the v3 dynamic rate limiter, a virtual key with tpm_limit=100 and a request contain [...]
Red method: worktree_at_fix_parent_then_cherry_pick
Red (fix reverted):
Green (fix present):
Pylon 5669: For a DB/UI-created search tool, an Anthropic Messages POST /v1/messages containing tools:[{type:"web_search_20250305",na [...]
Red method: worktree_at_fix_parent_then_cherry_pick
Red (fix reverted):
Green (fix present):
Pylon 5596: POST /v1/messages (also /v1/chat/completions) on the proxy with model bedrock_mantle/ and "stream": true. The ou [...]
Red method: revert_fix_at_head
Red (fix reverted):
Green (fix present):
Caveats (if any)
Low
Note
Low Risk
Test-only changes that exercise existing proxy behavior through integration harnesses; no runtime logic is modified in this diff.
Overview
Adds integration regression tests only (no production code in this diff) so ~41 July customer fixes stay pinned through the real proxy, Postgres/Redis, and scripted upstream peers.
New coverage spans provider wire contracts: Anthropic Messages bridges (thinking budget, timeouts, advisor credential rules, Fireworks
stop/reasoning_effort, OpenAI Responses mid-turn system, optional tools), Bedrock (Converse metadata stripping, DeepSeek reasoning fields, Nova invoke cache accounting, STS AssumeRole caching, Mantle clamp/hoist/SigV4/SSE streaming, Marengo embed shape, KBuserContext), Gemini cache_control, Fireworks session affinity, SageMaker inference-component signing, Vertex batch nulloutputInfoand fragmented Gemini streams, rerank/NIM ranking headers and payloads, Responses client-header forwarding and Codex namespace tools, plus A2A agents whose card lives only at/agentCard/v1.0with bearer auth.Routing and quotas gain tests for key/team/model TPM reservation, priority TPM 429s, advisor sub-call failures not cooling down the executor, and v3 rate-limit remainders on streaming callbacks. Streaming/SDK adds file-content early flush, Messages streams that open a thinking block from
reasoning_content-only chunks, and rebuilt aiohttp shared-session keepalive behavior.Existing advisor and A2A tests are lightly refactored (unique prompts,
hosted_vllm/gpt-4o-miniexecutor, formatting).Reviewed by Cursor Bugbot for commit a5e008a. Bugbot is set up for automated code reviews on this repo. Configure here.
Link to Devin session: https://app.devin.ai/sessions/d330510cfc8f463c823bd79c40b7e5e2
Open in Devin Desktop: https://app.devin.ai/desktop/session/d330510cfc8f463c823bd79c40b7e5e2?variant=devin
Requested by: @kerry-berri