Repository navigation
ci: move proxy, enterprise, mcp, caching, extras and gateway tests into tiers and fix the litellm-tests unit job - #42831
Closed
devin-ai-integration[bot] wants to merge 9 commits into
Closed
devin-ai-integration[bot] wants to merge 9 commits into
devin-ai-integration[bot] wants to merge 9 commits into
Conversation
Set COVERAGE_CORE=sysmon so the Python 3.12 coverage tracer stops pushing the Vertex streaming peak-memory and secret-detection linearity tests past their budgets, install the Codecov CLI with when: always so failing shards still upload, and run the relevance filter right after checkout so an irrelevant change skips the dependency install
…to tiers and run them from tests.yml
This reverts commit df47af1.
Contributor
Author
|
I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".
|
Contributor
|
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
…tests to their tiers
devin-ai-integration
Bot
force-pushed
the
litellm_ci_migration_proxy_misc
branch
from
September 24, 2026 01:47
b37de24 to
e68cf49
Compare
…runfailures can start Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Contributor
This was referenced Sep 24, 2026
Merged
3 tasks done
This branch is waiting to be deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What's the problem?
The
litellm-testsCircleCI pipeline that PR #42773 introduced is red on main: the default coverage tracer on Python 3.12 pushes the Vertex files streaming peak-memory tests past the 90s timeout and fails the secret-detection linearity test, a failing shard never uploads coverage because the codecov CLI install is skipped, and the relevance check runs after the multi-minute dependency install instead of before it. Separately, the legacy proxy, enterprise, MCP, caching, proxy-extras and gateway test trees still live outside thetests/{unit,integration,e2e}tiers and only run from GitHub Actions. CircleCI also injects every project environment variable into every job, so the unit tests could see provider keys,DATABASE_URL,LITELLM_LICENSEandPROXY_MASTER_KEYthat GHA never gave them, and a few moved unit tests read those keys or needed real Azure, OpenAI or Vertex credentialsWhat's the solution?
Fix the unit job first (
COVERAGE_CORE: sysmon,install_codecov_cliwithwhen: always,skip_unless_relevantright after checkout), then move the selected legacy trees into the tiers and run them from.circleci/tests.ymlunder their existing Codecov flags with their legacy retry settings. Every pytest andrun_integration.shstep in.circleci/tests.ymlruns underenv -iwith an allowlist, so unit tests cannot see project secrets. Key-dependent tests move to the tier that matches what they need: a scripted-upstream integration test replaces the provider-bounduser_configunit test, and the Vertex Geminicontentstoken counter test becomes an e2e test. The drained GitHub Actions jobs keep their check names and still run the moved paths for fork pull requests, which CircleCI does not buildHow does it fix it?
The
unitjob now takes a Codecovflagand arerunscount and asks.circleci/scripts/unit_selection.shfor the file list. A selection failure or an empty selection fails the job instead of passing an empty shard. The generalunitshards exclude the paths a legacy flag owns and each legacy flag (proxy-db-*,proxy-infra,enterprise-package,enterprise-routing,mcp-integration,caching-local,proxy-extras) gets its own single-shard job with the same workers, dist and Prisma setup it had in GHA, with--reruns 2 --reruns-delay 1 --rerun-except "from pytest-timeout"for every legacy shard and-p no:rerunfailuresformcp-integration(reruns 0), matching the GHA matrix.mcp-integrationinstalls the SDK1 peer (mcp==1.28.1,langchain-mcp-adapters==0.2.1) into a side venv and exportsMCP_TEST_PEER_PYTHON. Coverage is uploaded only when a realcoverage.xmlexists, with the roots unchanged (./litellm,./enterprise/litellm_enterprise) and the 90s timeout unchangedThe pytest step runs as
env -i PATH HOME CI=true COVERAGE_CORE LITELLM_LOCAL_MODEL_COST_MAP [MCP_TEST_PEER_PYTHON] uv run --no-sync pytest ..., and the integration step asenv -i PATH HOME CIRCLE_SHA1 CIRCLE_WORKFLOW_ID bash .circleci/scripts/run_integration.sh <suite>, which then creates its own synthetic master key, database URL and Redis URL.pytest-rerunfailuresopens alocalhostsocket for its xdist status server at configure time, so the unit socket policy intests/unit/conftest.pyallowslocalhostalongside127.0.0.1and::1; outbound hosts stay blockedMoves, by tree:
tests/proxy_unit_tests->tests/unit/proxy(with the proxyconftest.pyscoped there,test_configs,.env, fixtures and__file__-relative paths intact), excepttest_skills_db.py->tests/unit/skillsandtest_proxy_caching_responses_api.py->tests/unit/cachingtests/enterpriseandtests/test_litellm/enterprise->tests/unit/enterprise, mirroringlitellm_enterprisetests/mcp_tests->tests/unit/proxy/_experimental/mcp_serverfor the in-process manager tests andtests/unit/responses/mcpfor the responses helper teststhe four
tests/local_testingcaching files ->tests/unit/cachingtests/litellm-proxy-extras->tests/unit/litellm_proxy_extrastests/test_gateway->tests/unit/gatewayThe proxy-plus-scripted-MCP-server SDK1 suite (
test_proxy_mcp_e2e.py, its yaml config, math server and conftest) stays intests/mcp_tests: it needs the SDK1 peer interpreter that only themcp-integrationunit runner provides, so pertests/integration/AGENTS.mdit does not belong undertests/integration. The three provider-bound MCP files (test_mcp_litellm_client.py,test_semantic_tool_filter_e2e.pyand the provider half split out oftest_aresponses_api_with_mcp.pyastest_aresponses_api_with_mcp_providers.py) also stay intests/mcp_testsbecause they need real provider keys; they keep running from the unchanged GHA job..github/workflows/_test-unit-base.ymlgains afork-flaginput that appends the flag's selection on fork PRs and skips the Codecov upload when no report was produced.assert_ci_coverage.py,classify_changes.sh,merge-smoke-tests.json, theMakefile, the mutmut selectors and the redis-compat and rust workflows point at the new pathsKey-dependent tests:
tests/unit/proxy/test_proxy_pass_user_config.pyis deleted and replaced bytests/integration/routing/test_user_config_routing.py(selected by theprovidersgroup throughroutinginrun.pyGROUPS). It boots an owned proxy withgeneral_settings.allow_client_side_credentials: trueagainst a scripted local upstream, sends a chat completion with auser_configthat points at that upstream with a user-supplied key, and asserts exactly one/v1/chat/completionsrequest reached the upstream carryingAuthorization: Bearer <user key>and the user_config model. A second case boots the proxy without the opt-in and asserts the same request is rejected with 401.test_vertex_ai_gemini_token_counting_with_contentsmoves totests/e2e/llm_translation/test_token_counter_gemini_contents_e2e.pywith the skip removed, parametrised overgemini-2.5-flashandgemini-2.5-flash-vertexthrough/utils/token_counter, andload_vertex_ai_credentialsand the credential imports are gone fromtests/unit/proxy/test_proxy_token_counter.py.test_proxy_utils.py,enterprise/integrations/test_prometheus_unit_tests.pyandtest_proxy_custom_auth.pyuse synthetic literals instead of readingOPENAI_API_KEY,FIREWORKS_API_KEY,AZURE_AI_API_KEYandPROXY_MASTER_KEYCollection parity against main, by node id, for every legacy flag: caching 44/44, proxy extras 78/78, gateway 10/10, mcp 155/155, proxy 1140/1140 minus the deleted
test_proxy_pass_user_config.py::test_chat_completionand the moved Gemini contents test, enterprise 478 and 235 ids on the candidate for the merged enterprise trees. The only other id changes are the two intentional file renames above (test_proxy_server.py::test_gemini_pass_through_endpointmoved totest_proxy_server_gemini_pass_through.pyso the gemini test keeps its own module-level fixture, and the three provider tests moved intotest_aresponses_api_with_mcp_providers.py)How does the product experience change?
Nothing user-facing changes. CI contributors see the same check names; the legacy shards now report from CircleCI on same-repo PRs and from GHA on fork PRs
What caveats are there, if any?
The
rust-testcheck on this PR fails inlitellm-model-catalog::spec_parityonopenrouter/qwen/qwen3-coder-plus*_above_32k_tokenskeys. This branch does not touchlitellm-rust/ormodel_prices_and_context_window.json; those keys landed on main in #42832 after this branch's merge base (eb7eeb54, where the cost map has zeroabove_32k_tokensoccurrences), and thepull_requestevent tests the merge with current main, so the red is inherited:cargo test -p litellm-model-catalog --test spec_parityon a cleanorigin/mainworktree at660e6746e4fails the same three cases with the same key difference (the rust workflow is path-filtered, so it has not run on main since that merge). Themcp-integrationshard'stest_cancellation_delivers_termination_over_tcp[True-hang-5-scope]timeout seen on an earlier run did not reproduce: the full 48-case parametrised function passes on this branch (and the single case three more times) and passes 48/48 on main, with the test file byte-identical to main. The remainingtests/test_litellmmoves are PRs 3 and 4Linear ticket
How did you test this?
Proof run of
litellm-testson this head (definitiona1e47579-d7f3-43f1-8c6c-b20f4c3b38f7, branch as both config and checkout), all 21 jobs green: https://app.circleci.com/pipelines/github/BerriAI/litellm/90245. Codecov uploads for the head sha: https://app.codecov.io/gh/BerriAI/litellm/commit/c17f44074a9b67156080856e763454b1d6730d58. Earlier runs at previous heads: 90241 (first tiered green apart from the now deleteduser_configunit test), 90243 (theenv -ihead, where every legacy shard failed at configure time becausepytest-rerunfailuresconnects tolocalhost, fixed by the last commit), and the red proof with one assertion flipped intests/unit/caching/test_cache_preset_key.pyatdf47af1: https://app.circleci.com/pipelines/github/BerriAI/litellm/90237, reverted in the next commit. Greptile reports 5/5 on this head with both earlier threads resolvedSecret visibility inside pytest: with the same
env -iallowlist as the CircleCI step,env | cut -d= -f1 | sortinside a unit test process lists onlyCI,COVERAGE_CORE,HOME,LC_CTYPE,LITELLM_LOCAL_MODEL_COST_MAP,PATH,UV,UV_RUN_RECURSION_DEPTHandVIRTUAL_ENV. None ofOPENAI_API_KEY,AZURE_AI_API_KEY,DATABASE_URL,LITELLM_LICENSE,PROXY_MASTER_KEYor the Redis variables are visible, and therun_integration.shstep sees onlyPATH,HOME,CIRCLE_SHA1andCIRCLE_WORKFLOW_IDfrom the job. No guard test is committed for thisThe Gemini contents e2e test runs through the
e2e-changed-testsjob (.github/workflows/test-e2e-changed.yml):.github/e2e-stack/select_tests.pyselects any changedtests/e2e/**/test_*.pyon a same-repository PR and runs it three times against the stage-mirror stack once thee2e-changedenvironment is approved, and the file is not inUNSUPPORTED. Locally it passed for both deployments against real Gemini and Vertex credentials through/utils/token_counter. The routing integration test passed locally with the same environmentrun_integration.shbuilds (real Postgres and Redis, scripted upstream, synthetic master key, no provider keys) on both the opt-in and rejection cases, andrun.pyalready mapsroutinginto theprovidersgroupLocally, every legacy flag was run through
unit_selection.shwith the CircleCI pytest flags under the sameenv -iallowlist with outbound network blocked: caching-local 44 passed, proxy-extras 78 passed, enterprise-routing 188 passed, mcp-integration 110 passed 1 skipped, enterprise-package 286 passed 4 skipped, proxy-infra 10 passed, and the twelveproxy-db-*groups all green. The SDK1 MCP suite passed 30/30 against the peer venv. After thelocalhostallowlist change,tests/unit/gatewayandtests/unit/cachingwere rerun with--reruns 2 --reruns-delay 1 -n 2 --timeout=90underenv -i(54 passed), the same command shape that failed on run 90243.assert_ci_coverage.pyand--shardspass,run_merge_smoke.pyreports 11 cases, andmake checkpassesLink to Devin session: https://app.devin.ai/sessions/d541aba7e5844703bab65085ef2d54b2
Open in Devin Desktop: https://app.devin.ai/desktop/session/d541aba7e5844703bab65085ef2d54b2?variant=devin