Skip to content

test(ci): repair stale tests and move retired OpenAI text-completion fixtures - #43958

Merged
yuneng-berri merged 4 commits into
mainfrom
litellm_/gha-cci-stale-tests-7c96c7
Oct 1, 2026
Merged

yuneng-berri merged 4 commits into
mainfrom
litellm_/gha-cci-stale-tests-7c96c7

Conversation

@yuneng-berri

@yuneng-berri yuneng-berri commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • 13 tests red on main in GHA and CircleCI after intended product changes
  • 10 CircleCI text-completion tests pinned to retired OpenAI models
  • No product behavior is wrong; the tests encode stale contracts, fakes, or fixtures

How it solves it:

User Flow

Before: a contributor's PR inherits red GHA and CircleCI unit jobs that have nothing to do with their change

  1. They open a PR against main and GHA Unit Tests runs proxy-endpoints and misc
  2. Both jobs fail on batches metadata and Google Interactions spec tests
  3. The scheduled CircleCI pipeline also fails unit, local_testing_part1, local_testing_part2, and logging_testing, including 404s for retired OpenAI models

After: those jobs only go red when LiteLLM behavior actually changes

  1. They open a PR against main and GHA Unit Tests runs proxy-endpoints and misc
  2. The repaired tests pass, and still fail when the behavior they guard is broken

Screenshots / Proof of Fix

Tests-only PR, so there is no customer-facing request to curl. Product code is identical in Before and After; the proof is that each repaired test passes on current code and fails when the guarded behavior is mutated

Before (38b0762)

  1. pytest tests/test_litellm/proxy/batches_endpoints/test_endpoints.py -k non_object tests/local_testing/test_http_parsing_utils.py tests/unit/models/test_models.py::TestAutoRouterSession tests/unit/interactions/test_openapi_compliance.py tests/logging_callback_tests/test_gcs_pub_sub.py::test_async_gcs_pub_sub_v1
  2. 13 failed, matching CI: batches metadata 400 x2 (assert 'input_file_id' == 'metadata'), http parsing x3 ('MockRequest' object has no attribute 'scope'), auto-router label x3, Interactions spec x4 (KeyError: 'CreateModelInteractionParams', GET /interactions/{id} endpoint not found), GCS golden x1 (Extra key in actual: billing_agent_id and 5 metadata keys)
  3. test_read_request_body_empty_body and test_read_request_body_unexpected_error were also passing vacuously: both hit the missing-scope fallback, not the path they name

Retired text-completion fixtures

  1. CircleCI local_testing_part1 and local_testing_part2 fail 10 tests with HTTP 404 The model gpt-3.5-turbo-instruct / davinci-002 / gpt-3.5-turbo-1106 has been deprecated
  2. test_completion_openai_with_optional_params was also vacuous: its success callback still asserted model == "gpt-3.5-turbo-1106" after I pointed the call at gpt-6-luna, and the test passed, because callback failures never reach pytest

After (ac9e725)

  1. Same command plus the full mapped files: 357 passed
  2. Product mutations, each restored after the run:
    • Valid JSON body parsed to {}: 3 http parsing tests fail
    • Invalid JSON swallowed instead of raising: 2 http parsing tests fail
    • Non-object batch metadata coerced to {} instead of 400: both batches params fail
    • Partial-comparison guard removed from baseline_model: the multi-baseline partial test fails
    • Guard applied to a single baseline too: the single-baseline partial test fails
    • Tie broken by insertion order: the tie test fails
    • billing_agent_id dropped from the spend log: GCS golden fails with Missing key: billing_agent_id
    • Spec without tools on the model create body, without DELETE on the resource path, or without model required: the matching spec test fails
  3. Live OpenAI, real spend: the 6 OpenAI-backed text-completion tests pass (test_completion_text_openai, _async, test_qwen_text_completion, test_completion_openai_with_optional_params, test_completion_gpt_instruct, test_text_completion_basic)
  4. Mutations for those: GPT-5 config dropping seed or user fails the optional-params test with KeyError; text-completion config dropping logprobs fails test_qwen_text_completion
  5. CircleCI integration-security on main (ed4caeb): 31 tests fail with Owned proxy tried to reach external hosts: [b'CONNECT api.github.com:443 HTTP/1.1'], because the S2 sweep GETs every route and the new ROI route calls https://api.github.com/user/repos under default settings
  6. Fireworks-backed tests (test_completion_openai_prompt, test_completion_text_003_prompt_array, test_text_completion_with_echo[True/False]): with a mocked transport, LiteLLM sends the 2 string prompts, the 2 token-ID prompts, and echo plus logprobs unchanged to https://api.fireworks.ai/inference/v1/completions. Live verification runs in CircleCI local_testing_part2, since no Fireworks key exists outside CI

Why these vehicles

I probed every model on our OpenAI key against /v1/completions. The gpt-5.4 family and gpt-5.1 still accept it; gpt-5.6-terra and gpt-6-luna are chat-only. On every model that accepts it, a single prompt, logprobs up to 10, streaming, and echo alone work, but a prompt list with more than one entry and echo combined with logprobs return HTTP 500. Fireworks documents string and token-ID prompt batches and echo with logprobs, so those cases keep the text-completion-openai code path and only change the upstream

Type

Test

Caveats (if any)

Medium

  • The 4 Fireworks-backed tests are unverified live until CircleCI runs them
    • Fireworks' serverless catalog rotates, so glm-5p3-flash may need a swap later
  • OpenAI's 500 on batches and echo with logprobs is read as unsupported, not as an outage
    • Consistent across 6 models and retries, but OpenAI returned no 4xx naming the limit
  • Interactions tests still read Google's live spec, by design per the file
    • Google renamed {id} to {interactionsId} and folded CreateModelInteractionParams into a ModelInteraction oneOf variant
    • input is no longer marked required upstream, so the test now checks it is accepted rather than required
  • Left out on purpose, still red on main:

…abels, and Interactions spec lookups

Request fakes now carry the scope a real Starlette request has, the GCS pub/sub
spend-log golden gains the agent identity keys from #43722, the auto-router
session tests follow the baseline_models contract from #43348, and the
Interactions spec checks resolve the create body and resource paths from the
live spec instead of hardcoded names
@yuneng-berri
yuneng-berri requested a review from a team October 1, 2026 01:01
@greptile-apps

greptile-apps Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

[Low risk] Test files repaired to match current code and specs.

The PR appears safe to merge; no actionable issue was established in the changed tests.

Summary

This test-only PR updates request fakes, the spend-log golden payload, auto-router baseline assertions, and Google Interactions spec lookups to match current contracts. No actionable regression was established.

Reviews (1) · Last reviewed commit: "test(ci): repair stale request fakes, sp..."

@codecov

codecov Bot commented Oct 1, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

OpenAI still serves native /v1/completions on the gpt-5.4 family, so the
single-prompt cases move to text-completion-openai/gpt-5.4-nano. Multi-prompt
batches and echo with logprobs now 500 on every OpenAI model, so those cases
keep the same text-completion-openai transport pointed at Fireworks, which
documents both. The optional-params test asserts the request body actually
sent instead of a success callback whose assertions were swallowed
@yuneng-berri yuneng-berri changed the title test(ci): repair stale request fakes, spend-log golden, auto-router labels, and Interactions spec lookups test(ci): repair stale tests and move retired OpenAI text-completion fixtures Oct 1, 2026
…tch and echo cases

gpt-oss-20b is on-demand only on Fireworks, so the CI key got 404 model not
deployed; glm-5p3-flash is listed as serverless
@yuneng-berri yuneng-berri added run-ci and removed run-ci labels Oct 1, 2026
…route sweep

GET /roi-calculator/repositories (#43669) lists repositories from the configured
GitHub API, api.github.com by default, so the S2 sweep's GET of every route made
the owned proxy reach an external host and failed the egress check in 31
integration-security tests. It joins /get/latest_release_info in the deny list
@yuneng-berri yuneng-berri added run-ci and removed run-ci labels Oct 1, 2026
@yuneng-berri
yuneng-berri merged commit c168199 into main Oct 1, 2026
121 of 138 checks passed
@yuneng-berri
yuneng-berri deleted the litellm_/gha-cci-stale-tests-7c96c7 branch October 1, 2026 02:20
jan-sauer-reef added a commit to jan-sauer-reef/litellm that referenced this pull request Oct 1, 2026
…ject_key_prefix

* upstream/main: (62 commits)
  fix(guardrails): scan Responses API input in Azure Prompt Shield (BerriAI#43786)
  feat(lens): investigate sampled traces and retain batch results (BerriAI#43942)
  fix(proxy): restore pre-config-wins handling of pass-through endpoints (BerriAI#43962)
  fix(cost-map): raise baseten DeepSeek-V4.1-Flash max output to 262144 (BerriAI#43916)
  chore(cost-map): add deprecation date for anthropic claude-sonnet-4-5 (BerriAI#43898)
  chore(cost-map): add fireworks inkling priority prices from the prices api (BerriAI#43949)
  feat(guardrails): honor litellm_params.timeout in every HTTP guardrail (BerriAI#43134)
  test(e2e): typed per-test metadata for the e2e suite (BerriAI#42044)
  fix(caching): write the response-cache SET to Redis at once instead of on the post-call batch (BerriAI#43973)
  feat(ui): filter tags by name and description on the Tag Management page (BerriAI#42949)
  feat(providers): add Cortecs as an OpenAI-compatible provider (BerriAI#43872)
  feat(e2e): record each e2e test's steps, starting with ProxyClient (BerriAI#42393)
  test(ci): repair stale tests and move retired OpenAI text-completion fixtures (BerriAI#43958)
  feat(proxy): record in spend logs whether a request used a client-forwarded Anthropic OAuth token (BerriAI#43063)
  fix(azure_storage): keep the DataLakeServiceClient alive until its TTL elapses (BerriAI#43082)
  chore(deps): bump gitpython and tornado, extend diskcache osv ignore to Nov 1 (BerriAI#43961)
  fix(guardrails): treat an unknown straiker api_version as unset instead of skipping the guardrail (BerriAI#43956)
  fix(azure_storage): name Data Lake objects without base64 padding or slashes (BerriAI#43914)
  fix(grayswan): send request conversation and tool calls to post-call monitor (BerriAI#43770)
  chore(cost-map): sync openrouter prices from the models API (BerriAI#43950)
  ...
ydidwania added a commit to ydidwania/litellm that referenced this pull request Oct 1, 2026
main fixed the Interactions OpenAPI compliance test in BerriAI#43958; this empty
commit re-runs this PR's checks against it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
RichardoMrMu pushed a commit to RichardoMrMu/litellm that referenced this pull request Oct 3, 2026
Rebased onto latest main to pick up the Interactions spec test repair (BerriAI#43958); no behavior change.
RichardoMrMu pushed a commit to RichardoMrMu/litellm that referenced this pull request Oct 3, 2026
Rebased onto latest main: the Interactions spec test repair (BerriAI#43958) and the
tests/unit move (BerriAI#43186) are picked up; the regression test now lives at its
new path. No behavior change.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants