Skip to content

ci: move proxy, enterprise, mcp, caching, extras and gateway tests into tiers and fix the litellm-tests unit job - #42831

Closed
devin-ai-integration[bot] wants to merge 9 commits into
mainfrom
litellm_ci_migration_proxy_misc
Closed

devin-ai-integration[bot] wants to merge 9 commits into
mainfrom
litellm_ci_migration_proxy_misc

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

What's the problem?

The litellm-tests CircleCI pipeline that PR #42773 introduced is red on main: the default coverage tracer on Python 3.12 pushes the Vertex files streaming peak-memory tests past the 90s timeout and fails the secret-detection linearity test, a failing shard never uploads coverage because the codecov CLI install is skipped, and the relevance check runs after the multi-minute dependency install instead of before it. Separately, the legacy proxy, enterprise, MCP, caching, proxy-extras and gateway test trees still live outside the tests/{unit,integration,e2e} tiers and only run from GitHub Actions. CircleCI also injects every project environment variable into every job, so the unit tests could see provider keys, DATABASE_URL, LITELLM_LICENSE and PROXY_MASTER_KEY that GHA never gave them, and a few moved unit tests read those keys or needed real Azure, OpenAI or Vertex credentials

What's the solution?

Fix the unit job first (COVERAGE_CORE: sysmon, install_codecov_cli with when: always, skip_unless_relevant right after checkout), then move the selected legacy trees into the tiers and run them from .circleci/tests.yml under their existing Codecov flags with their legacy retry settings. Every pytest and run_integration.sh step in .circleci/tests.yml runs under env -i with an allowlist, so unit tests cannot see project secrets. Key-dependent tests move to the tier that matches what they need: a scripted-upstream integration test replaces the provider-bound user_config unit test, and the Vertex Gemini contents token counter test becomes an e2e test. The drained GitHub Actions jobs keep their check names and still run the moved paths for fork pull requests, which CircleCI does not build

How does it fix it?

The unit job now takes a Codecov flag and a reruns count and asks .circleci/scripts/unit_selection.sh for the file list. A selection failure or an empty selection fails the job instead of passing an empty shard. The general unit shards exclude the paths a legacy flag owns and each legacy flag (proxy-db-*, proxy-infra, enterprise-package, enterprise-routing, mcp-integration, caching-local, proxy-extras) gets its own single-shard job with the same workers, dist and Prisma setup it had in GHA, with --reruns 2 --reruns-delay 1 --rerun-except "from pytest-timeout" for every legacy shard and -p no:rerunfailures for mcp-integration (reruns 0), matching the GHA matrix. mcp-integration installs the SDK1 peer (mcp==1.28.1, langchain-mcp-adapters==0.2.1) into a side venv and exports MCP_TEST_PEER_PYTHON. Coverage is uploaded only when a real coverage.xml exists, with the roots unchanged (./litellm, ./enterprise/litellm_enterprise) and the 90s timeout unchanged

The pytest step runs as env -i PATH HOME CI=true COVERAGE_CORE LITELLM_LOCAL_MODEL_COST_MAP [MCP_TEST_PEER_PYTHON] uv run --no-sync pytest ..., and the integration step as env -i PATH HOME CIRCLE_SHA1 CIRCLE_WORKFLOW_ID bash .circleci/scripts/run_integration.sh <suite>, which then creates its own synthetic master key, database URL and Redis URL. pytest-rerunfailures opens a localhost socket for its xdist status server at configure time, so the unit socket policy in tests/unit/conftest.py allows localhost alongside 127.0.0.1 and ::1; outbound hosts stay blocked

Moves, by tree:

tests/proxy_unit_tests -> tests/unit/proxy (with the proxy conftest.py scoped there, test_configs, .env, fixtures and __file__-relative paths intact), except test_skills_db.py -> tests/unit/skills and test_proxy_caching_responses_api.py -> tests/unit/caching
tests/enterprise and tests/test_litellm/enterprise -> tests/unit/enterprise, mirroring litellm_enterprise
tests/mcp_tests -> tests/unit/proxy/_experimental/mcp_server for the in-process manager tests and tests/unit/responses/mcp for the responses helper tests
the four tests/local_testing caching files -> tests/unit/caching
tests/litellm-proxy-extras -> tests/unit/litellm_proxy_extras
tests/test_gateway -> tests/unit/gateway

The proxy-plus-scripted-MCP-server SDK1 suite (test_proxy_mcp_e2e.py, its yaml config, math server and conftest) stays in tests/mcp_tests: it needs the SDK1 peer interpreter that only the mcp-integration unit runner provides, so per tests/integration/AGENTS.md it does not belong under tests/integration. The three provider-bound MCP files (test_mcp_litellm_client.py, test_semantic_tool_filter_e2e.py and the provider half split out of test_aresponses_api_with_mcp.py as test_aresponses_api_with_mcp_providers.py) also stay in tests/mcp_tests because they need real provider keys; they keep running from the unchanged GHA job. .github/workflows/_test-unit-base.yml gains a fork-flag input that appends the flag's selection on fork PRs and skips the Codecov upload when no report was produced. assert_ci_coverage.py, classify_changes.sh, merge-smoke-tests.json, the Makefile, the mutmut selectors and the redis-compat and rust workflows point at the new paths

Key-dependent tests: tests/unit/proxy/test_proxy_pass_user_config.py is deleted and replaced by tests/integration/routing/test_user_config_routing.py (selected by the providers group through routing in run.py GROUPS). It boots an owned proxy with general_settings.allow_client_side_credentials: true against a scripted local upstream, sends a chat completion with a user_config that points at that upstream with a user-supplied key, and asserts exactly one /v1/chat/completions request reached the upstream carrying Authorization: Bearer <user key> and the user_config model. A second case boots the proxy without the opt-in and asserts the same request is rejected with 401. test_vertex_ai_gemini_token_counting_with_contents moves to tests/e2e/llm_translation/test_token_counter_gemini_contents_e2e.py with the skip removed, parametrised over gemini-2.5-flash and gemini-2.5-flash-vertex through /utils/token_counter, and load_vertex_ai_credentials and the credential imports are gone from tests/unit/proxy/test_proxy_token_counter.py. test_proxy_utils.py, enterprise/integrations/test_prometheus_unit_tests.py and test_proxy_custom_auth.py use synthetic literals instead of reading OPENAI_API_KEY, FIREWORKS_API_KEY, AZURE_AI_API_KEY and PROXY_MASTER_KEY

Collection parity against main, by node id, for every legacy flag: caching 44/44, proxy extras 78/78, gateway 10/10, mcp 155/155, proxy 1140/1140 minus the deleted test_proxy_pass_user_config.py::test_chat_completion and the moved Gemini contents test, enterprise 478 and 235 ids on the candidate for the merged enterprise trees. The only other id changes are the two intentional file renames above (test_proxy_server.py::test_gemini_pass_through_endpoint moved to test_proxy_server_gemini_pass_through.py so the gemini test keeps its own module-level fixture, and the three provider tests moved into test_aresponses_api_with_mcp_providers.py)

How does the product experience change?

Nothing user-facing changes. CI contributors see the same check names; the legacy shards now report from CircleCI on same-repo PRs and from GHA on fork PRs

What caveats are there, if any?

The rust-test check on this PR fails in litellm-model-catalog::spec_parity on openrouter/qwen/qwen3-coder-plus *_above_32k_tokens keys. This branch does not touch litellm-rust/ or model_prices_and_context_window.json; those keys landed on main in #42832 after this branch's merge base (eb7eeb54, where the cost map has zero above_32k_tokens occurrences), and the pull_request event tests the merge with current main, so the red is inherited: cargo test -p litellm-model-catalog --test spec_parity on a clean origin/main worktree at 660e6746e4 fails the same three cases with the same key difference (the rust workflow is path-filtered, so it has not run on main since that merge). The mcp-integration shard's test_cancellation_delivers_termination_over_tcp[True-hang-5-scope] timeout seen on an earlier run did not reproduce: the full 48-case parametrised function passes on this branch (and the single case three more times) and passes 48/48 on main, with the test file byte-identical to main. The remaining tests/test_litellm moves are PRs 3 and 4

Linear ticket

How did you test this?

Proof run of litellm-tests on this head (definition a1e47579-d7f3-43f1-8c6c-b20f4c3b38f7, branch as both config and checkout), all 21 jobs green: https://app.circleci.com/pipelines/github/BerriAI/litellm/90245. Codecov uploads for the head sha: https://app.codecov.io/gh/BerriAI/litellm/commit/c17f44074a9b67156080856e763454b1d6730d58. Earlier runs at previous heads: 90241 (first tiered green apart from the now deleted user_config unit test), 90243 (the env -i head, where every legacy shard failed at configure time because pytest-rerunfailures connects to localhost, fixed by the last commit), and the red proof with one assertion flipped in tests/unit/caching/test_cache_preset_key.py at df47af1: https://app.circleci.com/pipelines/github/BerriAI/litellm/90237, reverted in the next commit. Greptile reports 5/5 on this head with both earlier threads resolved

Secret visibility inside pytest: with the same env -i allowlist as the CircleCI step, env | cut -d= -f1 | sort inside a unit test process lists only CI, COVERAGE_CORE, HOME, LC_CTYPE, LITELLM_LOCAL_MODEL_COST_MAP, PATH, UV, UV_RUN_RECURSION_DEPTH and VIRTUAL_ENV. None of OPENAI_API_KEY, AZURE_AI_API_KEY, DATABASE_URL, LITELLM_LICENSE, PROXY_MASTER_KEY or the Redis variables are visible, and the run_integration.sh step sees only PATH, HOME, CIRCLE_SHA1 and CIRCLE_WORKFLOW_ID from the job. No guard test is committed for this

The Gemini contents e2e test runs through the e2e-changed-tests job (.github/workflows/test-e2e-changed.yml): .github/e2e-stack/select_tests.py selects any changed tests/e2e/**/test_*.py on a same-repository PR and runs it three times against the stage-mirror stack once the e2e-changed environment is approved, and the file is not in UNSUPPORTED. Locally it passed for both deployments against real Gemini and Vertex credentials through /utils/token_counter. The routing integration test passed locally with the same environment run_integration.sh builds (real Postgres and Redis, scripted upstream, synthetic master key, no provider keys) on both the opt-in and rejection cases, and run.py already maps routing into the providers group

Locally, every legacy flag was run through unit_selection.sh with the CircleCI pytest flags under the same env -i allowlist with outbound network blocked: caching-local 44 passed, proxy-extras 78 passed, enterprise-routing 188 passed, mcp-integration 110 passed 1 skipped, enterprise-package 286 passed 4 skipped, proxy-infra 10 passed, and the twelve proxy-db-* groups all green. The SDK1 MCP suite passed 30/30 against the peer venv. After the localhost allowlist change, tests/unit/gateway and tests/unit/caching were rerun with --reruns 2 --reruns-delay 1 -n 2 --timeout=90 under env -i (54 passed), the same command shape that failed on run 90243. assert_ci_coverage.py and --shards pass, run_merge_smoke.py reports 11 cases, and make check passes

Link to Devin session: https://app.devin.ai/sessions/d541aba7e5844703bab65085ef2d54b2
Open in Devin Desktop: https://app.devin.ai/desktop/session/d541aba7e5844703bab65085ef2d54b2?variant=devin

Set COVERAGE_CORE=sysmon so the Python 3.12 coverage tracer stops pushing the Vertex streaming peak-memory and secret-detection linearity tests past their budgets, install the Codecov CLI with when: always so failing shards still upload, and run the relevance filter right after checkout so an irrelevant change skips the dependency install
@devin-ai-integration
devin-ai-integration Bot requested a review from a team September 24, 2026 00:36
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@greptile-apps

greptile-apps Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

The PR appears safe to merge; no new actionable failures remain, and both previous review findings are resolved.

Summary

This PR reorganizes legacy test suites into the unit, integration, and e2e tiers while updating CircleCI and GitHub Actions ownership, secret isolation, retries, coverage handling, and path references.

  • Moves proxy, enterprise, MCP, caching, extras, and gateway tests under tiered test paths.
  • Adds deterministic CircleCI selection for legacy Codecov flags and restores their retry and worker settings.
  • Runs unit and integration tests with restricted environments to prevent ambient project credentials from reaching tests.
  • Replaces credential-dependent unit coverage with owned integration and live e2e coverage.
  • Fixes the previously reported empty-selection handling and removes the previously reported nonessential comments.

Reviews (2) · Last reviewed commit: "test: allow localhost loopback in the un..."

Comment thread .circleci/tests.yml Outdated
Comment thread tests/mcp_tests/test_aresponses_api_with_mcp_providers.py Outdated
@codspeed

codspeed Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_ci_migration_proxy_misc (c17f440) with main (8e74bb0)1

Open in CodSpeed

Footnotes

  1. No successful run was found on main (80c39a0) during the generation of this report, so 8e74bb0 was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩

@codecov

codecov Bot commented Sep 24, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

…runfailures can start

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

This branch is waiting to be deployed

1 waiting deployment
e2e-changed — c17f4407 Waiting Sep 24, 2026 by devin-ai-integration[bot] via Run changed e2e tests against the stage-mirror stack #10941
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants