Skip to content

test(mcp): verify scoped execution and OAuth credential isolation - #41731

Merged
mateo-berri merged 11 commits into
mainfrom
litellm_mcp_integration_regressions_4506
Sep 19, 2026
Merged

mateo-berri merged 11 commits into
mainfrom
litellm_mcp_integration_regressions_4506

Conversation

@joshua-berri

@joshua-berri joshua-berri commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Security checks lacked durable live gateway regression coverage
  • SDK2 exposed invalid fixture imports and test assumptions

How it solves it:

  • Tests scoped discovery, search, direct calls and virtual calls
  • Observes credentials across two users and same-URL servers
  • Checks expiry, revocation and selected guardrail denials
  • Removes void tests and records remaining acceptance gaps

User Flow

Before: a gateway admin who limits keys to specific MCP servers and stores per-user upstream tokens gets the right answers today, but nothing replays these calls against a live gateway, so a later release can break them unnoticed

  1. The admin sends POST https://litellm-domain/v1/mcp/server twice with the same upstream url and different aliases, then POST https://litellm-domain/key/generate with object_permission.mcp_servers naming only the first server
  2. A developer sends GET https://litellm-domain/mcp-rest/tools/list with that key and sees only the first server's add, multiply and fail tools
  3. They send POST https://litellm-domain/mcp-rest/tools/call with {"server_id": "<first server id>", "name": "add", "arguments": {"a": 3, "b": 5}} and get 200 with 8. The same call naming the second server's id returns 403 "not allowed", and the upstream server receives no request
  4. Two users each send POST https://litellm-domain/v1/mcp/server/{server_id}/oauth-user-credential with their own upstream token, and each user's add call returns 200 with that user's own bearer token reaching the upstream
  5. The first user sends DELETE https://litellm-domain/v1/mcp/server/{server_id}/oauth-user-credential. Their next tools/list and tools/call return 401 {"detail": "Unauthorized"} with a www-authenticate challenge and nothing reaches the upstream, while the second user's calls still return 200
  6. The developer repeats the add call with "guardrails": ["<blocking guardrail>"] in the body and gets 400 with the guardrail's message. The upstream tool never runs, and multiply still returns 200 with 15
  7. A later gateway release changes one of these answers, for example the second server's call starts returning 200 or a deleted credential keeps working. No CI job fails, so the admin finds out in production

After: the same calls return the same answers, and CI now replays them against a live gateway so a release that changes one of them fails before it ships

  1. The admin sends POST https://litellm-domain/v1/mcp/server twice with the same upstream url and different aliases, then POST https://litellm-domain/key/generate with object_permission.mcp_servers naming only the first server
  2. A developer sends GET https://litellm-domain/mcp-rest/tools/list with that key and sees only the first server's add, multiply and fail tools
  3. They send POST https://litellm-domain/mcp-rest/tools/call with {"server_id": "<first server id>", "name": "add", "arguments": {"a": 3, "b": 5}} and get 200 with 8. The same call naming the second server's id returns 403 "not allowed", and the upstream server receives no request
  4. Two users each send POST https://litellm-domain/v1/mcp/server/{server_id}/oauth-user-credential with their own upstream token, and each user's add call returns 200 with that user's own bearer token reaching the upstream
  5. The first user sends DELETE https://litellm-domain/v1/mcp/server/{server_id}/oauth-user-credential. Their next tools/list and tools/call return 401 {"detail": "Unauthorized"} with a www-authenticate challenge and nothing reaches the upstream, while the second user's calls still return 200
  6. The developer repeats the add call with "guardrails": ["<blocking guardrail>"] in the body and gets 400 with the guardrail's message. The upstream tool never runs, and multiply still returns 200 with 15
  7. A later gateway release that changes one of these answers fails the live integration suite in CI, so a key can never quietly reach the other server's tools and a deleted or expired credential can never quietly keep working

Relevant issues

Consumes merged SDK2 compatibility work in #41718. Complements principal coverage in #38680 and retains #41909's ownership of real OAuth consent and restart scenarios. No production code or dependency constraints change

Affected release

Linear ticket

Contributes to LIT-4506. The ten-guard inventory retains explicit UI, JWT and stateful limitations. This PR does not claim full ticket closure

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

This PR only adds and removes tests, so there is no end-user behavior to flip between the two legs. The proof has two parts: the same gateway calls answer the same way at the merge base and at the tip (the author's runs below), and the new tests are non-void, shown by an independent run at the tip plus one disabled production guard per test, each of which makes its test fail (the verification section below)

Author's runs: live gateway, isolated PostgreSQL/Redis, Python 3.12.13 and MCP SDK 2.2.0. The existing integration wrapper grants test entitlement; it does not verify license enforcement. The upstream is an owned SDK arithmetic server, so no model call is needed. Main uses the PR's SDK2-compatible upstream fixtures for this comparison, with unchanged main product source

The matching curl runs below passed on both revisions. Separately, the original PR tip e0b6bae failed five assertions in a live run; the repaired tests use the actual public error envelopes and search-provided qualified tool names

Setup: $KEY is a non-master key granted $ALLOWED_ID only. $VIRTUAL_KEY has the same grant plus tool search enabled. Both registered servers share one URL and use distinct synthetic bearer tokens. $ALLOWED_TOOL is the qualified name returned by search; $FORBIDDEN_TOOL is the other registered server's qualified name. OAuth users have separate scoped keys and saved synthetic upstream tokens

Before (974e4f1)

Scoped discovery and execution

  1. curl -sS http://127.0.0.1:14506/mcp-rest/tools/list -H "Authorization: Bearer $KEY": 200, exactly the granted server's three tools
  2. POST /mcp-rest/tools/call with {"server_id":"$ALLOWED_ID","name":"add","arguments":{"a":3,"b":5}}: 200, isError=false, result 8, one upstream execution. Substitute $FORBIDDEN_ID: 403, zero upstream requests
  3. POST the same endpoint with $VIRTUAL_KEY and {"name":"mcp_tool_search","arguments":{"query":"add","top_k":10}}: only $ALLOWED_TOOL. Call {"name":"mcp_tool_call","arguments":{"tool_name":"$ALLOWED_TOOL","arguments":{"a":3,"b":5}}}: 200, result 8. Substitute $FORBIDDEN_TOOL: 403, zero upstream requests

Tool failure

  1. POST /mcp-rest/tools/call with $KEY and {"server_id":"$ALLOWED_ID","name":"fail","arguments":{}}: 200, isError=true, text Error executing tool fail

Removed static credentials

  1. After a successful authenticated call, clear the server's saved headers through PUT /v1/mcp/server
  2. GET /mcp-rest/tools/list?server_id=$ALLOWED_ID: 500 with detail.error=internal. POST the original add call: 500 with requires a usable upstream credential. Both produce zero upstream requests

OAuth isolation and revocation

  1. Two non-admin users store separate tokens, list tools, then POST the same server's add call: both return 200 and 8; observed upstream bearer tokens differ by user
  2. Delete the first user's saved credential. Its separate discovery and execution requests both return 401 {"detail":"Unauthorized"}, with zero upstream requests
  3. Repeat the second user's add call: 200 and 8, using that user's original upstream bearer token

After (5b9f3d4)

Scoped discovery and execution

  1. curl -sS http://127.0.0.1:14506/mcp-rest/tools/list -H "Authorization: Bearer $KEY": 200, exactly the granted server's three tools
  2. POST /mcp-rest/tools/call with {"server_id":"$ALLOWED_ID","name":"add","arguments":{"a":3,"b":5}}: 200, isError=false, result 8, one upstream execution. Substitute $FORBIDDEN_ID: 403, zero upstream requests
  3. POST the same endpoint with $VIRTUAL_KEY and {"name":"mcp_tool_search","arguments":{"query":"add","top_k":10}}: only $ALLOWED_TOOL. Call {"name":"mcp_tool_call","arguments":{"tool_name":"$ALLOWED_TOOL","arguments":{"a":3,"b":5}}}: 200, result 8. Substitute $FORBIDDEN_TOOL: 403, zero upstream requests

Tool failure

  1. POST /mcp-rest/tools/call with $KEY and {"server_id":"$ALLOWED_ID","name":"fail","arguments":{}}: 200, isError=true, text Error executing tool fail

Removed static credentials

  1. After a successful authenticated call, clear the server's saved headers through PUT /v1/mcp/server
  2. GET /mcp-rest/tools/list?server_id=$ALLOWED_ID: 500 with detail.error=internal. POST the original add call: 500 with requires a usable upstream credential. Both produce zero upstream requests

OAuth isolation and revocation

  1. Two non-admin users store separate tokens, list tools, then POST the same server's add call: both return 200 and 8; observed upstream bearer tokens differ by user
  2. Delete the first user's saved credential. Its separate discovery and execution requests both return 401 {"detail":"Unauthorized"}, with zero upstream requests
  3. Repeat the second user's add call: 200 and 8, using that user's original upstream bearer token

Local validation: BASE_REF=origin/main make lint passed all lint, formatting, strict-rule, type-discipline, test-quality and basedpyright budget gates. The affected guardrail, credential resolver and discovery-outcome backend files passed 134 tests. The full extensions suite passed all 18 tests in 210.05 seconds at the current tip. Combined integration and standalone SDK2 CLI coverage measured 182/182 changed executable lines (100%) and 46/48 branches (95.8%) inside touched functions. The two uncovered branches are pre-existing empty-body and second-receive paths in the upstream recorder, outside the changed SDK setup and security scenarios. No exclusions or budgets were changed. Completed current-tip Codecov reports 82.13% repository line coverage after all relevant uploads, with CI passing. Its production-only patch contains zero measured changed lines; the separately measured 100% figure covers this test-only diff

Independent verification (reviewer's runs)

Each run boots the proxy from the worktree at the named commit with --num_workers 1, a fresh Postgres database, its own Redis and the SDK upstream from tests/integration/_support/upstream.py, all on random ports, mirroring the CircleCI integration-extensions job (Python 3.14.3, pytest 9.0.3, mcp 2.2.0)

The extensions suite at a41b60c: 18 passed, exit 0, all seven new node ids included. At 5b9f3d4, after the main merge, with the venv resynced to litellm-enterprise 0.1.69 and litellm-proxy-extras 0.4.100: 18 passed in 178.76s, exit 0, and test_assert_ci_coverage.py 34 passed. The diff against the new merge base 974e4f1 is the same nine files, and no main commit since the merge base touches tests/integration, tests/mcp_tests or litellm/proxy/_experimental/mcp_server, so the mutation results below carry over

Non-void check, one production guard disabled per run at a41b60c, worktree clean after each revert. Every mutation made its target test fail

Guard disabled Test Result
get_allowed_mcp_servers returns every server scoped grants, both permission variants 3 failed
credential lookup falls back to another user's row OAuth isolation [revoke] 1 failed
pre_mcp_call guardrail hook skipped guardrail effects 1 failed
adapter accepts a missing upstream credential warm credential removal 1 failed
RefreshingTokenStore._is_expired always false OAuth isolation [expire] 1 failed

tests/mcp_tests/mcp_e2e_upstream_server.py has no in-repo runner, so it was booted from the worktree venv on a random MCP_PORT with MCP_HOST=127.0.0.1. Under mcp 2.2.0 the moved run(...) kwargs bind the requested host and port, initialize with a localhost Host header is accepted (DNS rebinding protection off), tools/list returns add and multiply, and tools/call add {"a":3,"b":5} returns 8. Session handling is unchanged: the server was stateful before this PR as well. tests/test_litellm/test_assert_ci_coverage.py (the other reader of contracts.json): 34 passed at the tip

/live-pr-risk. Breaking: none, no production code changes. Backward incompatible: none, the upstream fixture keeps its transport, tool set and defaults. Regression risk: none left unexercised. Dependency graph: tests/integration/_support/mcp.py is imported by compatibility/test_persisted_toolsets.py and observability/test_guardrail_effects.py (both run in the suite above), contracts.json is read by tests/integration/run.py, _support/manifest.py, test_assert_ci_coverage.py and .circleci/scripts/verify_integration_browser.py (the first three ran here, the last only lists browser contracts), mcp_e2e_upstream_server.py is imported only by _support/mcp.py (booted live above), and the two deleted files are referenced nowhere. Not verified: nothing skipped

CircleCI at 5b9f3d4: pipeline 89891. build_and_test 43/44 green, the one red being proxy_e2e_anthropic_messages_tests (job 2191862), whose two failing tests test_bedrock_invoke_messages_with_all_beta_headers[bedrock-claude-opus-4.5-bedrock] and [bedrock-converse-claude-sonnet-4.5-bedrock_converse] get Bedrock's 400 invalid beta flag on main's last four scheduled pipelines too (89818, 89834, 89873, 89877, the same two tests each time, green at 89785 on 2026-09-18), all before this branch's merge base, tracked in LIT-8149. integration 7/8 with integration-extensions (the suite this PR adds to) green and integration-cost red on test_case_bills_expected_cost[fireworks_ai-accounts-fireworks-models-deepseek-v4p1-flash-fallback_cache_read_at_input_rate] (job 2191821), which main's pipeline 89877 (job 2191123, at 2886b8e, an ancestor of the merge base) fails with the same 0.0012456 != 0.0021672; this PR does not touch tests/integration/cost_calculation or the cost map, and the open #41999 repins that case. Neither red is a required check, and both are recorded as Low caveats below

Type

✅ Test

Caveats (if any)

No Severe, High or Medium caveat is left: the PR ships tests only, and the four items the earlier body ranked Medium are scope boundaries or pinned shipped behavior rather than gaps in what this diff ships, so they are recorded as Low with the ticket that owns each

Low

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR
  • 5b9f3d4 passes /live-pr-risk

@joshua-berri
joshua-berri requested a review from a team September 18, 2026 01:37
@joshua-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@joshua-berri

Copy link
Copy Markdown
Contributor Author

@greptileai please review current head daff22a, focusing on historical regression evidence, actual upstream execution assertions, and legacy test removal scope

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@greptile-apps

greptile-apps Bot commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

The scoped MCP regression changes appear safe to merge, with no actionable finding introduced by the latest-main merge

Summary

This PR replaces invalid SDK2 MCP fixtures and test-owned mock coverage with live gateway regressions for scoped discovery, execution authorization, credential isolation, revocation, expiry, and selected guardrail enforcement

  • Uses the SDK2 MCPServer API for controlled upstream peers
  • Verifies exact server visibility and same-URL execution isolation
  • Verifies per-user, per-server OAuth credential isolation and failure after revocation or expiry
  • Verifies selected guardrails block direct and virtual calls before upstream execution
  • Documents tested guarantees and explicitly retained acceptance gaps
  • The latest-main merge does not alter the MCP changes reviewed previously

Reviews (6) · Last reviewed commit: "chore: merge latest main for MCP regress..."

Comment thread tests/integration/mcp/test_mcp_lifecycle.py Outdated
@joshua-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@codecov

codecov Bot commented Sep 18, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@joshua-berri

Copy link
Copy Markdown
Contributor Author

@greptileai please review current head f83992f after the full-result health assertion repair, and refresh the current-tip confidence summary

… rejections

Co-Authored-By: bot_apk <apk@cognition.ai>
@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you all sign our Contributor License Agreement before we can accept your contribution.
1 out of 2 committers have signed the CLA.

✅ joshua-berri
❌ devin-ai-integration[bot]
You have signed the CLA already but the status is still pending? Let us recheck it.

Co-Authored-By: bot_apk <apk@cognition.ai>
@joshua-berri joshua-berri changed the title test(mcp): guard health authorization, credential removal and selected policies test(mcp): verify scoped execution and OAuth credential isolation Sep 19, 2026
@joshua-berri

Copy link
Copy Markdown
Contributor Author

@greptileai @cursor review Please review the current SDK2 fixture changes and new permission, credential-isolation, expiry, and revocation tests.

@joshua-berri

Copy link
Copy Markdown
Contributor Author

@cursor review Please review commit a41b60c for permission isolation, virtual discovery, OAuth failures, and the corrected live regression assertions.

@joshua-berri

Copy link
Copy Markdown
Contributor Author

@greptileai Please review commit a41b60c, especially server-qualified virtual calls, both permission variants, and the revised credential-failure assertions.

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@joshua-berri

Copy link
Copy Markdown
Contributor Author

@greptileai Please review the latest-main merge at 5b9f3d4 and confirm the scoped MCP regressions still have no actionable findings.

@joshua-berri

Copy link
Copy Markdown
Contributor Author

@cursor review Please review commit 5b9f3d4 after the latest-main merge for actionable issues in the scoped MCP security regressions.

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 5b9f3d4. Configure here.

@mateo-berri mateo-berri left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Thanks!

@mateo-berri
mateo-berri merged commit 43c0a26 into main Sep 19, 2026
150 of 155 checks passed
@mateo-berri
mateo-berri deleted the litellm_mcp_integration_regressions_4506 branch September 19, 2026 23:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants