Skip to content

test(e2e): add conversational matrix across chat, messages and responses - #42359

Merged
mateo-berri merged 2 commits into
mainfrom
litellm_e2e_conversational_matrix
Sep 22, 2026
Merged

mateo-berri merged 2 commits into
mainfrom
litellm_e2e_conversational_matrix

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Bugs land on one of /v1/chat/completions, /v1/messages, /v1/responses and get fixed one endpoint at a time
  • No single e2e suite runs the same contract across endpoints, providers, models and auth methods
  • Background /v1/models discovery makes strict record/replay of OpenAI deployments non-deterministic

How it solves it:

  • One parameterized matrix suite: 3 endpoints x 3 models x 2 auth methods x 5 behaviors, 90 cases
  • Adding a model, from an existing or a new provider, is one Deployment row in DEPLOYMENTS, no new test code
  • Suite is replayable: live, record, and replay all pass, replay spends nothing
  • New general_settings.disable_model_info_refresh turns off the background poller, set in the record/replay config

User Flow

Before: a maintainer adding a shared helper across the three chat endpoints has no single test run that tells them whether all three still behave the same

  1. They change a helper used by /v1/chat/completions, /v1/messages and /v1/responses
  2. They run the per-endpoint e2e files and each one asserts different things, so a regression in tool-call ids on /v1/messages passes unnoticed because only the chat suite checks tool-call ids
  3. They ship, a customer reports the /v1/messages regression, and the fix plus test gets added to the messages suite only
  4. Separately, a proxy admin who records e2e fixtures sees replay fail with "4 of 5 recorded interactions never consumed, e.g. get /openai/v1/models" because the proxy polled the upstream on its own schedule during the recording

After: the same maintainer runs one suite and sees the same 5 checks for every endpoint, provider, model and auth method

  1. They change the same helper
  2. They run E2E_FIXTURE_MODE=replay pytest tests/e2e/llm_translation/test_conversational_matrix_e2e.py against a local proxy in under 4 minutes with no provider spend
  3. The case test_tool_call_is_returned_named_and_addressable[messages-claude-haiku-4-5-env_ref] fails with "messages tool call has no id, so the caller cannot answer it", naming exactly which endpoint, model and auth path regressed
  4. To cover a new model they add one Deployment(...) row and re-record; all 30 per-model cases exist immediately
  5. The proxy admin adds disable_model_info_refresh: true under general_settings, restarts, and replay consumes every recorded interaction

Relevant issues

Affected release

Linear ticket

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Screenshots / Proof of Fix

Shared setup. Proxy booted from tests/e2e/gateway/record_replay_ci_config.yml with Postgres, OPENAI_API_KEY and LITELLM_MASTER_KEY in the environment. A tiny HTTP sniffer listens on 127.0.0.1:4999 and logs every request it receives (auth header redacted). Two deployments exist in the DB before either boot:

POST /model/new  {"model_name":"proof-sniffed-upstream","litellm_params":{"model":"openai/gpt-4o-mini","api_key":"sk-proof-upstream-bogus","api_base":"http://127.0.0.1:4999/v1"}}
POST /model/new  {"model_name":"proof-gpt-5.4-mini","litellm_params":{"model":"openai/gpt-5.4-mini","api_key":"os.environ/OPENAI_API_KEY"}}

The chat request used in both runs:

curl -s -D - http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $LITELLM_MASTER_KEY" -H "Content-Type: application/json" \
  -d '{"model":"proof-gpt-5.4-mini","messages":[{"role":"user","content":"Reply with exactly: proxy ok"}]}'

Before (1a4e2b1)

Background /v1/models discovery hits the upstream on its own

  1. Start the proxy at the merge base on port 4001 (Uvicorn running at 22:56:37 UTC), send no requests, watch the sniffer
  2. Observed: the sniffer receives a GET two seconds after boot and again exactly 300 s later, with nobody calling the proxy
    2026-09-21T22:56:39.024948+00:00 GET /v1/models {"host": "127.0.0.1:4999", "user-agent": "litellm/1.103.0", "authorization": "<redacted>", ...}
    2026-09-21T23:01:39.023685+00:00 GET /v1/models {"host": "127.0.0.1:4999", "user-agent": "litellm/1.103.0", "authorization": "<redacted>", ...}
    
  3. disable_model_info_refresh: true is not a recognized setting at this commit, so there is nothing an operator can set to stop it

Chat completion through the proxy

  1. Run the shared curl against port 4001
  2. Observed
    HTTP/1.1 200 OK
    x-litellm-call-id: 6f46706d-f2e7-427d-801b-92982e5ad19b
    x-litellm-response-cost: 3.15e-05
    content="proxy ok" prompt_tokens=12 completion_tokens=5
    

After (bb4f670)

Background /v1/models discovery hits the upstream on its own

  1. Start the proxy at the PR tip on port 4000 with disable_model_info_refresh: true in general_settings (boot at 23:05:05 UTC), send no requests, watch the sniffer for 110 s
  2. Observed: zero sniffer lines after the boot timestamp
    $ awk '$1 >= "2026-09-21T23:05:05"' sniff.log
    (no output)
    
  3. GET /v1/models on the proxy itself still lists both deployments, so nothing else about model registration changed
    ['proof-gpt-5.4-mini', 'proof-sniffed-upstream']
    

Chat completion through the proxy

  1. Run the shared curl against port 4000
  2. Observed
    HTTP/1.1 200 OK
    x-litellm-call-id: 38129576-3ef5-4741-a566-f48bb13472ad
    x-litellm-response-cost: 3.15e-05
    content="proxy ok" prompt_tokens=12 completion_tokens=5
    

Matrix runs against the same proxy at the tip, all three lanes, for the record (these are on top of the curl proof, not instead of it):

E2E_FIXTURE_MODE=live    pytest tests/e2e/llm_translation/test_conversational_matrix_e2e.py   -> 90 passed in 304 s at bb4f670a71
E2E_FIXTURE_MODE=record  E2E_FIXTURE_DIR=/tmp/fx pytest ...conversational_matrix_e2e.py       -> 90 passed  (fresh bundle, 0 fixture files mention v1/models)
E2E_FIXTURE_MODE=replay  E2E_FIXTURE_DIR=/tmp/fx OPENAI_API_KEY=bogus ANTHROPIC_API_KEY=bogus pytest ... -> 90 passed in 214 s, every recorded interaction consumed

The tool tests force the call rather than hoping for it: tool_choice names get_weather and parallel calls are off on the first turn (tool_choice={"type":"function","function":{"name":"get_weather"}}, parallel_tool_calls=false on chat, {"type":"tool","name":"get_weather","disable_parallel_tool_use":true} on Messages, {"type":"function","name":"get_weather"}, parallel_tool_calls=false on Responses), so the assertion is exactly one get_weather call and a miss is a translation bug in litellm, not the model's mood. The second turn offers the tool without forcing it, so the model can answer in text. The recorded bundle shows the forced choice reaching both providers: 16 upstream bodies carry "parallel_tool_calls": false and 12 carry "disable_parallel_tool_use": true

Mutation check for the unit test: removing the if model_info_scheduler is not None guard (always scheduling) makes test_proxy_startup_event_honors_disable_model_info_refresh[True-False] fail with disable_model_info_refresh=True but refresh_model_info job is <Job ...>; restoring it goes green

Type

🆕 New Feature
✅ Test

Caveats (if any)

Medium

  • Only OpenAI and Anthropic are in the matrix; Bedrock needs SigV4, which the provider edge cannot re-sign, so it is left for a follow-up
  • Passthrough is not in the matrix: it has no shared "reply with text, usage and tool calls" contract to assert
  • disable_model_info_refresh also hides discovered max_model_len from /model/info for hosted vLLM deployments; it is off by default and only set in the e2e record/replay config

Low

  • test_lifecycle.py has pre-existing ruff findings (unused imports); untouched to keep the diff to the new test
  • The 300 s poller still fires in every other proxy config exactly as before this PR

QA runbook

Prerequisites: a proxy on http://localhost:4000 booted from tests/e2e/gateway/record_replay_ci_config.yml with Postgres, OPENAI_API_KEY, ANTHROPIC_API_KEY and LITELLM_MASTER_KEY ($LITELLM_MASTER_KEY below). Each test runs once per cell, a cell being endpoint x model x auth: endpoints chat_completions (/v1/chat/completions), messages (/v1/messages), responses (/v1/responses); models openai/gpt-4o-mini, openai/gpt-5.4-mini, anthropic/claude-haiku-4-5; auth env_ref ("api_key":"os.environ/OPENAI_API_KEY") or stored_credential (a /credentials entry referenced by litellm_credential_name). The steps below use the chat_completions-gpt-4o-mini-stored_credential cell; swap the route and body for the other cells

  • tests/e2e/llm_translation/test_conversational_matrix_e2e.py::TestConversationalMatrix::test_reply_carries_assistant_text_and_usage - a non-streaming reply on every cell carries an id, assistant text, positive usage, and an x-litellm-call-id

    • POST /credentials with the master key: {"credential_name":"qa-cred","credential_values":{"api_key":"<real OPENAI key>"}}
    • POST /model/new: {"model_name":"qa-chat","litellm_params":{"model":"openai/gpt-4o-mini","litellm_credential_name":"qa-cred"}}; poll GET /v1/models until qa-chat appears
    • POST /key/generate {} and use the returned key for the LLM call
    • POST /v1/chat/completions {"model":"qa-chat","messages":[{"role":"user","content":"Say hello."}]}
    • Expect 200, a non-empty id, non-empty choices[0].message.content, usage.prompt_tokens > 0, usage.completion_tokens > 0, and an x-litellm-call-id response header
    • Sanity check: this test makes sense to add and is not hand-wavey (e.g., assert actual expected spend instead of just spend > 0) or potentially flaky
  • tests/e2e/llm_translation/test_conversational_matrix_e2e.py::TestConversationalMatrix::test_stream_delivers_text_usage_and_a_terminal_event - a streamed reply on every cell arrives as more than one event, carries text, ends with a terminal event, and reports usage

    • Same deployment and key as above
    • POST /v1/chat/completions with "stream":true,"stream_options":{"include_usage":true} (for /v1/messages and /v1/responses plain "stream":true suffices)
    • Expect more than one SSE chunk, at least one chunk with non-empty delta.content, a chunk with finish_reason set, and a final chunk whose usage is non-null (Messages: a message_delta with usage; Responses: response.completed with response.usage)
    • Sanity check: this test makes sense to add and is not hand-wavey (e.g., assert actual expected spend instead of just spend > 0) or potentially flaky
  • tests/e2e/llm_translation/test_conversational_matrix_e2e.py::TestConversationalMatrix::test_cost_header_matches_the_spend_log - on every cell x-litellm-response-cost is positive and equals the one priced /spend/logs row written for a fresh key, which also carries token counts and the deployment's model

    • POST /key/generate {} for a brand new key so it has no prior spend rows
    • POST /v1/chat/completions with that key and a prompt containing a unique marker; read x-litellm-response-cost from the headers and expect it > 0
    • Poll GET /spend/logs?api_key= until a row with spend > 0 appears
    • Expect exactly one such row, prompt_tokens > 0, completion_tokens > 0, model equal to gpt-4o-mini, and spend within 1% of the header value
    • Sanity check: this test makes sense to add and is not hand-wavey (e.g., assert actual expected spend instead of just spend > 0) or potentially flaky
  • tests/e2e/llm_translation/test_conversational_matrix_e2e.py::TestConversationalMatrix::test_tool_call_is_returned_named_and_addressable - on every cell a prompt that needs the get_weather tool produces exactly that tool call, with a non-empty call id and JSON arguments that keep the requested location

    • POST /v1/chat/completions with "tools":[{"type":"function","function":{"name":"get_weather","parameters":{"type":"object","properties":{"location":{"type":"string"}},"required":["location"]}}}], "tool_choice":{"type":"function","function":{"name":"get_weather"}}, "parallel_tool_calls":false and the user message "What is the weather in Paris right now? Use the get_weather tool." (Messages: "tool_choice":{"type":"tool","name":"get_weather","disable_parallel_tool_use":true}; Responses: "tool_choice":{"type":"function","name":"get_weather"}, "parallel_tool_calls":false)
    • Expect choices[0].message.tool_calls with exactly one entry, named get_weather, a non-empty id, and arguments that parse to JSON with location containing "paris" (Messages: a tool_use block with id, name, input; Responses: a function_call item with call_id, name, arguments)
    • Sanity check: this test makes sense to add and is not hand-wavey (e.g., assert actual expected spend instead of just spend > 0) or potentially flaky
  • tests/e2e/llm_translation/test_conversational_matrix_e2e.py::TestConversationalMatrix::test_tool_result_round_trip_reaches_the_model - on every cell, feeding the tool result back under the model's own call id yields a final answer that repeats the returned 22 degrees

    • Run the tool-call step above and keep the returned call id
    • POST the same endpoint again with the original user message, the assistant tool call, and a tool result "Paris: 22 degrees Celsius, clear skies" addressed to that id (chat: role":"tool","tool_call_id"; Messages: tool_result block with tool_use_id; Responses: function_call_output with call_id)
    • Expect 200 and assistant text containing "22"
    • Sanity check: this test makes sense to add and is not hand-wavey (e.g., assert actual expected spend instead of just spend > 0) or potentially flaky

Nuances a manual run will hit: the proxy must have the qa-chat alias propagated before the first LLM call (poll /v1/models) and /v1/messages needs max_tokens (the suite sends 512 everywhere)

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Link to Devin session: https://app.devin.ai/sessions/3990b1604bb34b33aa75d198337d630a
Open in Devin Desktop: https://app.devin.ai/desktop/session/3990b1604bb34b33aa75d198337d630a?variant=devin
Requested by: @mateo-berri


Note

Low Risk
Opt-in startup flag skips background model polling; default unchanged. Most of the diff is new e2e coverage, not production path changes.

Overview
Adds a replayable e2e matrix that runs the same five conversation checks (basic reply/stream, cost vs spend log, forced tool call, tool-result round trip) on every combination of chat completions / messages / responses, OpenAI + Anthropic deployments, and env-ref vs stored credentials, via shared Surface adapters in conversational_matrix.py. Coverage registry entries point at the new suite.

Introduces general_settings.disable_model_info_refresh: when true, the proxy does not schedule the periodic refresh_model_info job (background /v1/models polling). E2e record/replay config turns this on so fixtures stay deterministic; default behavior is unchanged. A lifecycle unit test pins the flag.

Risk: Low for production (opt-in flag, tests dominate). Operators who disable refresh lose automatic discovered model metadata (noted in PR caveats).

Reviewed by Cursor Bugbot for commit bb4f670. Bugbot is set up for automated code reviews on this repo. Configure here.

Parameterizes one behavioral contract (reply, stream, cost log, tool call,
tool round trip) across /v1/chat/completions, /v1/messages and /v1/responses,
OpenAI and Anthropic models, and env-ref vs stored-credential auth, with
record/replay fixtures.

Adds general_settings.disable_model_info_refresh so the proxy fronting a
replay fixture does not poll every OpenAI-compatible deployment's /v1/models
in the background, which otherwise leaves unconsumed interactions in the
recorded bundle.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@codspeed

codspeed Bot commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_e2e_conversational_matrix (bb4f670) with main (c1c1ec4)1

Open in CodSpeed

Footnotes

  1. No successful run was found on main (c21ab96) during the generation of this report, so c1c1ec4 was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩

@greptile-apps

greptile-apps Bot commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

The PR appears safe to merge, with both previous findings resolved and no new actionable issues identified.

Summary

Adds a configurable way to disable background model-info refresh and introduces a replayable conversational E2E matrix across Chat Completions, Messages, and Responses.

  • Covers three models and two credential paths across non-streaming, streaming, cost accounting, tool calls, and tool-result round trips.
  • Forces a single named tool call, addressing the earlier provider-controlled and duplicate-call findings.
  • Disables unsolicited model discovery in the record/replay configuration and verifies scheduler behavior with a lifecycle test.

Reviews (3) · Last reviewed commit: "test(e2e): force the weather tool on the..."

Comment thread tests/e2e/llm_translation/conversational_matrix.py Outdated
Comment thread tests/e2e/llm_translation/test_conversational_matrix_e2e.py Outdated
@codecov

codecov Bot commented Sep 21, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

…er to Deployment

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

1 similar comment
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

Comment on lines +136 to +155
@dataclass(frozen=True, slots=True)
class Cell:
surface: SurfaceName
deployment: Deployment
auth: AuthMethod

@property
def id(self) -> str:
return f"{self.surface}-{self.deployment.label}-{self.auth}"

def registry_id(self, capability: Capability, streaming: Streaming, assertion: Assertion) -> str:
return f"llm.{self.surface}.{self.deployment.route}.{capability}.{streaming}.{assertion}"


CELLS: Final[tuple[Cell, ...]] = tuple(
Cell(surface=surface, deployment=deployment, auth=auth)
for surface in SURFACES
for deployment in DEPLOYMENTS
for auth in AUTH_METHODS
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

beautifully done

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@mateo-berri mateo-berri left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit bb4f670. Configure here.

@mateo-berri
mateo-berri merged commit 30d8b12 into main Sep 22, 2026
94 of 98 checks passed
@mateo-berri
mateo-berri deleted the litellm_e2e_conversational_matrix branch September 22, 2026 05:32

This branch is waiting to be deployed

1 waiting deployment
e2e-changed — bb4f670a Waiting Sep 21, 2026 by devin-ai-integration[bot] via Run changed e2e tests against the stage-mirror stack #9013
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants