Skip to content

feat(proxy)!: refuse to start with an unset, empty, or publicly known master key - #42019

Merged
mateo-berri merged 10 commits into
mainfrom
litellm_master_key_boot_enforcement
Sep 20, 2026
Merged

mateo-berri merged 10 commits into
mainfrom
litellm_master_key_boot_enforcement

Conversation

@ryan-crabbe-berri

@ryan-crabbe-berri ryan-crabbe-berri commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • No master key means every request is accepted unauthenticated
  • sk-1234 is public, and it boots as the admin key
  • Nothing told the operator either was happening

How it solves it:

  • The proxy refuses to start on an unset, empty, or sk-1234 master key
  • The last thing on screen says where the bad key came from
  • It prints a copy-pastable command that generates a secure key
  • A key that is already exported has to be replaced where it is set, because it wins over .env, and the error says so
  • When the database holds values encrypted with the bad key, the error also asks for LITELLM_MIGRATE_FROM_MASTER_KEY, and the next boot re-encrypts them with the new key
  • dangerously_permit_weak_or_unset_master_key starts anyway, for local development

This is a breaking change on purpose. A deployment that runs on no master key, an empty one, or sk-1234 stops booting after the upgrade until it sets a real key or sets LITELLM_DANGEROUSLY_PERMIT_WEAK_OR_UNSET_MASTER_KEY=true (or general_settings.dangerously_permit_weak_or_unset_master_key: true)

User Flow

Before: an operator who copies the quick start ends up with a proxy that anyone can administer, and nothing tells them

  1. They start the proxy with LITELLM_MASTER_KEY=sk-1234, the value every example used, or with no master key at all
  2. The proxy starts and serves traffic
  3. Anyone who can reach it sends GET http://litellm-domain/v1/models with Authorization: Bearer sk-1234, or with no Authorization header at all when no key is set, and gets 200
  4. With sk-1234 that caller is the proxy admin, so they can also create keys, read spend, and change models

After: the same start command stops with instructions, and the proxy only serves once it has a real key

  1. They start the proxy with LITELLM_MASTER_KEY=sk-1234, or with no master key at all
  2. The proxy exits, and the last thing on screen says why, where the bad key came from, and how to fix it
  3. They run the printed command and start the proxy again
    • With no key set it is echo "LITELLM_MASTER_KEY=sk-$(openssl rand -hex 32)" | tee -a .env
    • With sk-1234 already exported it is echo "sk-$(openssl rand -hex 32)", and they put the output in place of sk-1234 where it is set
  4. It starts. GET http://litellm-domain/v1/models returns 200 with the new key and is rejected with sk-1234
  5. Nobody can administer the proxy with the public key any more, and a proxy with no key can no longer be reached without one
  6. If their database already holds credentials encrypted with sk-1234, step 2 also tells them to set LITELLM_MIGRATE_FROM_MASTER_KEY=sk-1234
    • The next start logs Re-encrypting N stored value(s)... and then Done re-encrypting N stored value(s) with the new master key. You may now delete the LITELLM_MIGRATE_FROM_MASTER_KEY environment variable.
    • Their models, credentials, and callbacks keep working, and leaving the variable set later does nothing except log a reminder

Relevant issues

Companion PRs: #42011 removes sk-1234 from the configs, READMEs, and UI snippets we ship, and BerriAI/litellm-docs#1577 adds the #proxy-refuses-to-start section the error links to

Affected release

Linear ticket

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Every case is a live proxy started from the litellm console script on port 4893 (4894 for the database cases, against a throwaway postgres:16 container on 55893), with LITELLM_MODE=PRODUCTION so no .env is read. Before runs the litellm/ package exported from the merge base, After runs the PR tip. Generated keys are cut to their first 6 characters, and the long temp path of the config is shortened to ./config.yaml

config.yaml:

model_list:
  - model_name: claude-haiku
    litellm_params:
      model: anthropic/claude-haiku-4-5-20251001
      api_key: os.environ/ANTHROPIC_API_KEY

general_settings:
  master_key: os.environ/LITELLM_MASTER_KEY

Before (b946d12)

sk-1234 as the master key

  1. LITELLM_MASTER_KEY=sk-1234 litellm --port 4893 --config config.yaml starts and serves
  2. curl -s -o /dev/null -w '%{http_code}' http://127.0.0.1:4893/v1/models -H 'Authorization: Bearer sk-1234' prints 200

No master key anywhere

  1. litellm --port 4893 --config config.yaml with LITELLM_MASTER_KEY unset starts and serves
  2. curl -s -o /dev/null -w '%{http_code}' http://127.0.0.1:4893/v1/models with no Authorization header prints 200

Run the printed command, then boot

  1. Nothing prints a command on this commit, so this runs the one the PR adds: echo "LITELLM_MASTER_KEY=sk-$(openssl rand -hex 32)" | tee -a .env prints LITELLM_MASTER_KEY=sk-9cd...
  2. With that value exported, litellm --port 4893 --config config.yaml starts and serves
  3. /v1/models prints 200 with the generated key and 400 with Bearer sk-1234

sk-1234 with the override

  1. LITELLM_MASTER_KEY=sk-1234 LITELLM_DANGEROUSLY_PERMIT_WEAK_OR_UNSET_MASTER_KEY=true litellm --port 4893 --config config.yaml starts and serves; the variable means nothing on this commit
  2. The boot log has no warning about the key, and /v1/models prints 200 with Bearer sk-1234

sk-1234 with a database and no salt key

  1. DATABASE_URL=postgresql://postgres:postgres@127.0.0.1:55893/rot_before LITELLM_MASTER_KEY=sk-1234 litellm --port 4894 --config config.yaml starts and serves, with Datasource "client": PostgreSQL database "rot_before" in the log
  2. curl -s -X POST http://127.0.0.1:4894/credentials -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{"credential_name":"before-cred","credential_values":{"api_key":"sk-fake-credential-key"},"credential_info":{"custom_llm_provider":"openai"}}' returns {"success":true,"message":"Credential created successfully"}, so the secret is now stored encrypted with the public key
  3. curl -s http://127.0.0.1:4894/key/list -H 'Authorization: Bearer sk-1234' returns HTTP 200 {"keys":[],"total_count":0,"current_page":1,"total_pages":0}
  4. Lines in the boot log that mention the master key: 0

After (ecf1751)

sk-1234 as the master key

  1. LITELLM_MASTER_KEY=sk-1234 litellm --port 4893 --config config.yaml; echo "exit code $?" prints exit code 3, and the last thing on screen is:
LiteLLM proxy refused to start: the master key is a publicly known default.
It comes from the LITELLM_MASTER_KEY environment variable.

1. Generate a key:
     echo "sk-$(openssl rand -hex 32)"
   Put it in place of the current LITELLM_MASTER_KEY value wherever that is set: a shell export, your
   container or deployment environment, or its line in .env. Do not just add it to .env, because a value
   already exported in the environment wins over .env.

Local development only: set LITELLM_DANGEROUSLY_PERMIT_WEAK_OR_UNSET_MASTER_KEY=true, or
general_settings.dangerously_permit_weak_or_unset_master_key: true, to start anyway.
  1. Nothing listens on 4893, so the /v1/models call from Before cannot be made

No master key anywhere

  1. litellm --port 4893 --config config.yaml; echo "exit code $?" with LITELLM_MASTER_KEY unset prints exit code 3, and the last thing on screen is:
LiteLLM proxy refused to start: no master key is set, so every request would be accepted without authentication.
general_settings.master_key in ./config.yaml is blank, or points at an environment variable that is not set.

1. Make sure ./config.yaml reads the key from the environment:
     general_settings:
       master_key: os.environ/LITELLM_MASTER_KEY
2. Generate a key and save it to .env:
     echo "LITELLM_MASTER_KEY=sk-$(openssl rand -hex 32)" | tee -a .env
   Not using a .env file (docker run, Kubernetes, pip install)? Pass the same value as the
   LITELLM_MASTER_KEY environment variable instead.

Local development only: set LITELLM_DANGEROUSLY_PERMIT_WEAK_OR_UNSET_MASTER_KEY=true, or
general_settings.dangerously_permit_weak_or_unset_master_key: true, to start anyway.
  1. Nothing listens on 4893, so the unauthenticated /v1/models call from Before cannot be made

Run the printed command, then boot

  1. echo "LITELLM_MASTER_KEY=sk-$(openssl rand -hex 32)" | tee -a .env prints LITELLM_MASTER_KEY=sk-36d... (64 hex characters in total)
  2. With that value exported, litellm --port 4893 --config config.yaml starts and serves
  3. /v1/models prints 200 with the generated key and 400 with Bearer sk-1234

sk-1234 with the override

  1. LITELLM_MASTER_KEY=sk-1234 LITELLM_DANGEROUSLY_PERMIT_WEAK_OR_UNSET_MASTER_KEY=true litellm --port 4893 --config config.yaml starts and serves
  2. The boot log now has WARNING: master_key_boot_check.py:122 - dangerously_permit_weak_or_unset_master_key is on, so the proxy is starting with a publicly known master key. Never run this outside local development., and /v1/models prints 200 with Bearer sk-1234

sk-1234 with a database and no salt key

Setup, done once with LITELLM_DANGEROUSLY_PERMIT_WEAK_OR_UNSET_MASTER_KEY=true so the proxy boots on sk-1234 against an empty postgres:16 database mig: POST /model/new with a fake provider api_key, POST /credentials, POST /config/update with one environment variable, POST /team/new plus POST /team/migrate-test-team/callback with fake Langfuse keys, and POST /key/generate for a virtual key sk-bWW.... All returned HTTP 200, so the database now holds secrets encrypted with sk-1234. One plaintext LiteLLM_Config row, {"note": "...", "separator": "-", "allowed_routes": ["*"]}, is inserted with psql as a control, because "-", "*" and "..." base64-decode to no bytes

  1. DATABASE_URL=postgresql://postgres:postgres@127.0.0.1:55893/mig LITELLM_MASTER_KEY=sk-1234 litellm --port 4894 --config config.yaml; echo "exit code $?" prints exit code 3. The proxy counted the stored values that decrypt under sk-1234 (6, the plaintext row is not one of them), and the steps now ask for the key to migrate from:
LiteLLM proxy refused to start: the master key is a publicly known default.
It comes from the LITELLM_MASTER_KEY environment variable.

Your database holds 6 value(s) encrypted with this master key,
which encrypts stored credentials while LITELLM_SALT_KEY is not set. Replacing the key alone makes them
unreadable, so also tell the proxy which key to migrate from:
1. Set the key to migrate from next to LITELLM_MASTER_KEY, wherever that is set (a shell export, your
   container or deployment environment, or .env):
     LITELLM_MIGRATE_FROM_MASTER_KEY=sk-1234
2. Generate a key:
     echo "sk-$(openssl rand -hex 32)"
   Put it in place of the current LITELLM_MASTER_KEY value wherever that is set: a shell export, your
   container or deployment environment, or its line in .env. Do not just add it to .env, because a value
   already exported in the environment wins over .env.
3. Start the proxy again. It re-encrypts the stored values with the new key, then logs that
   LITELLM_MIGRATE_FROM_MASTER_KEY can be removed. Details: https://docs.litellm.ai/docs/proxy/master_key_rotations#proxy-refuses-to-start

Local development only: set LITELLM_DANGEROUSLY_PERMIT_WEAK_OR_UNSET_MASTER_KEY=true, or
general_settings.dangerously_permit_weak_or_unset_master_key: true, to start anyway.
  1. Setting only the new variable is not enough. The same command with LITELLM_MIGRATE_FROM_MASTER_KEY=sk-1234 added and LITELLM_MASTER_KEY still sk-1234 prints exit code 3 again
  2. echo "sk-$(openssl rand -hex 32)" prints sk-c8b.... DATABASE_URL=... LITELLM_MASTER_KEY=sk-c8b... LITELLM_MIGRATE_FROM_MASTER_KEY=sk-1234 litellm --port 4894 --config config.yaml starts and serves. These are the 2nd and 3rd warnings in the boot log, right after the trusted_proxy_ranges one:
18:55:57 - LiteLLM Proxy:WARNING: master_key_migration.py:262 - Re-encrypting 6 stored value(s) from the LITELLM_MIGRATE_FROM_MASTER_KEY key to the new master key.
18:55:57 - LiteLLM Proxy:WARNING: master_key_migration.py:221 - Done re-encrypting 6 stored value(s) with the new master key. You may now delete the LITELLM_MIGRATE_FROM_MASTER_KEY environment variable.
  1. In the database every ciphertext changed: the model's api_key from 1LY6Sw2u... to vMqUp0Eo..., the credential's from MSXseJ5e... to AyEwY01y..., the environment variable from MonDssnS... to NbM_5K1E..., and the team's Langfuse secret from litellm_enc::uGrGNVE0... to litellm_enc::yxNzktDx..., which kept its litellm_enc:: marker. The plaintext control row is byte for byte the same
  2. The secrets still decrypt under the new key, with 0 Error decrypting value lines in the log:
    • curl -s http://127.0.0.1:4894/model/info -H 'Authorization: Bearer sk-c8b...' returns HTTP 200 with the DB model migrate-test-model
    • curl -s http://127.0.0.1:4894/credentials/by_name/migrate-test-cred -H 'Authorization: Bearer sk-c8b...' returns HTTP 200 {"credential_name":"migrate-test-cred","credential_info":{"custom_llm_provider":"openai"},"credential_values":{"api_key":"sk****st"}}
    • curl -s http://127.0.0.1:4894/models -H 'Authorization: Bearer sk-bWW...', the virtual key created under sk-1234, still returns HTTP 200 with migrate-test-model
  3. The proxy is stopped and started again with the variable still set. It starts and serves, the four ciphertexts are unchanged, and the only migration line in the log is:
18:56:35 - LiteLLM Proxy:WARNING: master_key_migration.py:221 - LITELLM_MIGRATE_FROM_MASTER_KEY is still set, but nothing in the database is left to migrate from that key. You may now delete LITELLM_MIGRATE_FROM_MASTER_KEY.
  1. The variable is removed and the proxy is started again with only LITELLM_MASTER_KEY=sk-c8b.... The log has 0 migration lines and 0 decryption errors, and the three calls from step 5 return the same HTTP 200 responses

sk-1234 with a database that holds nothing encrypted

  1. DATABASE_URL=postgresql://postgres:postgres@127.0.0.1:55893/mig_empty LITELLM_MASTER_KEY=sk-1234 litellm --port 4894 --config config.yaml; echo "exit code $?" prints exit code 3. mig_empty has the schema and no stored secrets, so the steps are the same as the first case, with no mention of LITELLM_MIGRATE_FROM_MASTER_KEY:
LiteLLM proxy refused to start: the master key is a publicly known default.
It comes from the LITELLM_MASTER_KEY environment variable.

1. Generate a key:
     echo "sk-$(openssl rand -hex 32)"
   Put it in place of the current LITELLM_MASTER_KEY value wherever that is set: a shell export, your
   container or deployment environment, or its line in .env. Do not just add it to .env, because a value
   already exported in the environment wins over .env.

Local development only: set LITELLM_DANGEROUSLY_PERMIT_WEAK_OR_UNSET_MASTER_KEY=true, or
general_settings.dangerously_permit_weak_or_unset_master_key: true, to start anyway.
  1. The same happens on the seeded mig database when LITELLM_SALT_KEY is set, because the salt key is what encrypts stored values then
  2. With DATABASE_URL pointing at a port nothing listens on, the refusal takes about 10 seconds longer, starts with Your database could not be checked for values encrypted with this master key, and prints the migration steps

A database on an older schema

  1. mig_partial is a copy of the migrated mig with DROP TABLE "LiteLLM_SSOIdentityAssertion" CASCADE, so one of the tables the migration knows about does not exist
  2. DATABASE_URL=.../mig_partial LITELLM_MASTER_KEY=sk-62b... LITELLM_MIGRATE_FROM_MASTER_KEY=sk-c8b... litellm --port 4894 --config config.yaml, with a second generated key, starts and serves. The log has 18:58:36 ... Re-encrypting 6 stored value(s)... followed by the Done re-encrypting 6 stored value(s) line, and curl -s http://127.0.0.1:4894/credentials/by_name/migrate-test-cred -H 'Authorization: Bearer sk-62b...' returns HTTP 200 with "api_key":"sk****st" and 0 decryption errors. The missing table is skipped, and the variable works for any previous key, not only sk-1234

A write that fails during the migration

  1. mig_fail is a copy of the migrated mig with a trigger that makes every UPDATE on "LiteLLM_CredentialsTable" raise simulated write failure
  2. DATABASE_URL=.../mig_fail LITELLM_MASTER_KEY=sk-168... LITELLM_MIGRATE_FROM_MASTER_KEY=sk-c8b... litellm --port 4894 --config config.yaml; echo "exit code $?" prints exit code 3. The proxy does not serve with values it cannot read, and the log says what to do:
18:57:38 - LiteLLM Proxy:WARNING: master_key_migration.py:262 - Re-encrypting 6 stored value(s) from the LITELLM_MIGRATE_FROM_MASTER_KEY key to the new master key.
18:57:38 - LiteLLM Proxy:WARNING: master_key_migration.py:221 - Could not migrate stored values from the LITELLM_MIGRATE_FROM_MASTER_KEY key (RawQueryError: ERROR: simulated write failure). Values still encrypted with the previous key cannot be read until the migration succeeds. Keep LITELLM_MIGRATE_FROM_MASTER_KEY set and restart the proxy once the database is reachable.
  1. The trigger is dropped and the same command is run again. It starts and serves, and finishes what the failed boot left behind:
18:58:03 - LiteLLM Proxy:WARNING: master_key_migration.py:262 - Re-encrypting 4 stored value(s) from the LITELLM_MIGRATE_FROM_MASTER_KEY key to the new master key.
18:58:03 - LiteLLM Proxy:WARNING: master_key_migration.py:221 - Done re-encrypting 4 stored value(s) with the new master key. You may now delete the LITELLM_MIGRATE_FROM_MASTER_KEY environment variable.
  1. curl -s http://127.0.0.1:4894/credentials/by_name/migrate-test-cred -H 'Authorization: Bearer sk-168...' returns HTTP 200 with "api_key":"sk****st", GET /model/info returns HTTP 200 with the DB model, and the log has 0 decryption errors

Type

New Feature

Caveats (if any)

Severe

  • Deployments on no key, an empty key, or sk-1234 stop booting on upgrade
    • This is the point of the PR, and it needs a line in the release notes
    • Way out: set a real key, or set the override to keep the old behavior
  • With a database and no LITELLM_SALT_KEY, swapping the key without LITELLM_MIGRATE_FROM_MASTER_KEY makes stored credentials undecryptable
    • The error detects this case by counting the stored values that decrypt under the bad key, and only then asks for the variable
    • Nothing is destroyed if they swap without it: setting the variable on a later boot still migrates the data
  • The boot-time migration writes to the database before the proxy serves
    • Each row is updated only if it still holds the value that was read, so several workers or replicas booting at once do not overwrite each other or a concurrent edit
    • Only values that decrypt under the previous key are touched. Both ciphers are authenticated, so plaintext and values under another key are left alone
    • Plaintext such as "*", "-" or "" base64-decodes to no bytes, which the legacy cipher reads as an empty plaintext under any key. The migration rejects those before decrypting, and the proof below keeps such a row unchanged
    • It reads and writes through the primary database, never a read replica
    • Tables or columns missing from an older schema are skipped
    • A database error during the migration is logged with the instruction to keep the variable set and restart, and then stops the boot, so a worker never serves with stored values it cannot read
      • The one exception is the rule the Prisma setup already follows at boot: a connection outage while allow_requests_on_db_unavailable is on is tolerated
      • That rule is narrow on purpose. In a live probe, stopping Postgres right after the connect made query_raw raise an engine AttributeError, which the rule does not count as a connection outage, so the boot stopped even with the flag on. Without the flag, a database that is already down at boot fails earlier, in the Prisma setup, with exit code 3
      • The next boot picks up where the failed one stopped, as the proof below shows

Medium

  • Multi-worker uvicorn (--num_workers 4) prints the fix once per worker and the parent exits with code 0
    • Single-worker uvicorn and gunicorn exit 3, hypercorn exits 1, none of them crash-loop
    • Uvicorn's multi-worker parent does this for every startup failure, such as an unreachable database, so it is not changed here
  • A refusal now connects to the database to count encrypted values, where it used to refuse before any connection
    • It only happens when the key is empty or sk-1234, a database is configured, and no salt key is set
    • An unreachable database delays the refusal by about 10 seconds (Prisma's connect retries), and the error then says the database could not be checked and prints the migration steps
  • While LITELLM_MIGRATE_FROM_MASTER_KEY is set, every boot reads the secret-bearing columns once to look for values under the previous key
    • Team, key, and user metadata are filtered in SQL to rows that carry the litellm_enc:: marker, so the big tables are not loaded
    • The reminder to delete the variable is logged at WARNING, because the default log level hides INFO
  • The migration covers more columns than POST /key/regenerate with new_master_key does
    • Extra: callback vars in team, key, and user metadata, CloudZero and Vantage settings, SSO settings, cache settings, and config overrides
    • /key/regenerate is not changed here
  • The fix is written straight to stderr when the process exits, not through the logger
    • The log redactor strips the key-shaped command, and the failed-startup traceback is about 345 lines long
    • A log pipeline that only collects logger output sees the one-line reason and not the full fix
  • POST /key/regenerate with new_master_key while LITELLM_SALT_KEY is set leaves stored secrets unreadable under both keys
    • Already true on main, reproduced live, not changed here. The boot-time migration does nothing when a salt key is set, and says so in the log

Low

  • The .env step only takes effect where the proxy loads .env (repo checkouts, docker compose env_file), so the message also says to pass the env var
  • No strength check yet: sk-12345 or password still boot
  • A key that is already set always gets the replace-in-place wording, even when it came from a line in .env, where appending a second line would also have worked
  • A YAML master_key: sk-1234 plus the same value in the env var is reported as coming from the environment first, then from the config on the next boot
  • The proof makes no LLM call, because the change is at boot and not in the request path
  • CI and test boots that relied on sk-1234 or on no key now set the override; they were not rewritten to use a real key
    • That covers the CircleCI docker runs, tests/e2e/ui/run_e2e.sh, create_proxy_test_client in tests/test_litellm/proxy/conftest.py, and the session fixture in tests/mcp_tests/test_proxy_mcp_e2e.py
  • render.yaml now asks Render to generate LITELLM_MASTER_KEY, so one-click deploys keep booting

QA runbook

No e2e test was added or changed. The only edit under tests/e2e is tests/e2e/ui/run_e2e.sh, which now exports LITELLM_DANGEROUSLY_PERMIT_WEAK_OR_UNSET_MASTER_KEY=true next to its sk-1234 so the UI suite's proxy still boots

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

… master key

The proxy used to boot with no master key (every request accepted without
authentication) and with sk-1234, the key every example used. It now stops at
startup, before it connects to the database, and prints how to fix it: where the
bad key came from, a copy-pastable command that generates a secure key, and,
when the public key is also encrypting a database, a link to the rotation guide

general_settings.dangerously_allow_unsafe_proxy: true or
LITELLM_DANGEROUSLY_ALLOW_UNSAFE_PROXY=true starts the proxy anyway, for local
development. CI and test boots that rely on sk-1234 or on no key set it

BREAKING CHANGE: deployments with no master key, an empty one, or sk-1234 no
longer start until they set a real key or opt in to the override
@ryan-crabbe-berri
ryan-crabbe-berri requested a review from a team September 19, 2026 20:53
@codspeed

codspeed Bot commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_master_key_boot_enforcement (ecf1751) with main (93e39d5)1

Open in CodSpeed

Footnotes

  1. No successful run was found on main (daecea3) during the generation of this report, so 93e39d5 was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩

…0 and keep the lazy OpenAPI snapshot as generated by CI's Python
@greptile-apps

greptile-apps Bot commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

The PR appears safe to merge, with no outstanding findings or newly introduced actionable defects

Summary

This PR adds proxy startup security hardening, an explicit development override, and boot-time migration support for stored encrypted values

  • Rejects unsafe master-key configurations unless the operator explicitly opts into the development override
  • Detects and re-encrypts database values during an intentional master-key rotation
  • Stops startup when a non-tolerated migration failure could leave persisted values unreadable
  • Updates deployment manifests, generated configuration types, CI setup, and focused tests for the new startup contract

Reviews (5) · Last reviewed commit: "refactor(proxy): rename the local develo..."

Comment thread litellm/proxy/auth/master_key_boot_check.py
Comment thread litellm/proxy/proxy_server.py
Comment thread litellm/proxy/auth/master_key_boot_check.py Outdated
@codecov

codecov Bot commented Sep 19, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 96.25850% with 11 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
litellm/proxy/auth/master_key_boot_check.py 95.55% 6 Missing ⚠️
litellm/proxy/proxy_server.py 66.66% 3 Missing ⚠️
litellm/proxy/db/master_key_migration.py 98.46% 2 Missing ⚠️

📢 Thoughts on this report? Let us know!

return source.config_file_path or "your config"


def _source_line(refusal: UnsafeMasterKeyRefused) -> str:
Comment thread litellm/proxy/auth/master_key_boot_check.py Fixed
@ryan-crabbe-berri

Copy link
Copy Markdown
Contributor Author

@greptile re review

Comment thread litellm/proxy/auth/master_key_boot_check.py Fixed

@mateo-berri mateo-berri left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Thanks!

Comment thread .circleci/config.yml Outdated
command: |
docker run -d \
-p 4001:4000 \
-e LITELLM_DANGEROUSLY_ALLOW_UNSAFE_PROXY=true \

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: LITELLM_DANGEROUSLY_PERMIT_WEAK_MASTER_KEY=true <- I feel like this is more descriptive

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Renamed in ecf1751 to LITELLM_DANGEROUSLY_PERMIT_WEAK_OR_UNSET_MASTER_KEY (YAML dangerously_permit_weak_or_unset_master_key), since the override also covers an unset or empty key, not only a weak one

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

late to the party and bit of a nit, but we explicitly are not checking key strength, so calling it "weak" is a misnomer. I think LITELLM_DANGEROUSLY_PERMIT_DEFAULT_OR_UNSET_MASTER_KEY would have been better, but I imagine changing this after merge is a lot more trouble than it's worth

PUBLICLY_KNOWN_MASTER_KEYS: Final = frozenset({"sk-1234"})
ROTATION_DOCS_URL: Final = "https://docs.litellm.ai/docs/proxy/master_key_rotations#proxy-refuses-to-start"
_NEW_MASTER_KEY: Final = "sk-$(openssl rand -hex 32)"
GENERATE_MASTER_KEY_COMMAND: Final = f'echo "{MASTER_KEY_ENV_VAR}={_NEW_MASTER_KEY}" | tee -a .env'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maybe use echo "${MASTER_KEY_ENV_VAR}=${_NEW_MASTER_KEY}" >> .env so it doesn't emit the master key to stdout for the user...?

@ryan-crabbe-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

…ROM_MASTER_KEY so an unsafe key can be replaced while the proxy refuses to start

Rotating through POST /key/regenerate needs a running proxy, which a refused boot does not have. The refusal now counts the stored values that decrypt under the unsafe key. When there are none it only asks for a new key. When there are some it also asks for LITELLM_MIGRATE_FROM_MASTER_KEY, and the next boot with a safe key re-encrypts them and logs that the variable can be deleted. Leaving the variable set afterwards is a no-op with one notice.
return numbered if refusal.migration is None else f"{_migration_lead(refusal.migration)}\n{numbered}"


def _config_steps(source: MasterKeySource) -> tuple[str, ...]:
Comment thread litellm/proxy/db/master_key_migration.py Fixed
return NothingToMigrate.NOTHING_ENCRYPTED_WITH_PREVIOUS_KEY if another_worker_migrated_everything else migrated


def describe_outcome(outcome: MigrationOutcome) -> str:
… ciphertext during the master key migration

A string such as "*" or "..." has no base64 characters, so it decoded to no bytes and read as an empty plaintext under any key. The migration would have counted it and overwritten it with a ciphertext of the empty string. Also read from the writer database instead of a read replica, report a database error during the migration instead of crashing the boot, skip columns the connected schema lacks across every schema on the search path, cap the JSON walk depth for the recursion detector, and move the boot wiring into one tested function.
ReplaceCiphertext = Callable[[str], str | None]


def replace_ciphertexts(value: JsonValue, replacement_for: ReplaceCiphertext, depth: int = 0) -> tuple[JsonValue, int]:
@ryan-crabbe-berri

Copy link
Copy Markdown
Contributor Author

@greptile re review

@ryan-crabbe-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

Comment thread litellm/proxy/db/master_key_migration.py Outdated

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

…ls, unless allow_requests_on_db_unavailable tolerates the outage
@ryan-crabbe-berri

Copy link
Copy Markdown
Contributor Author

@greptile re review

@ryan-crabbe-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread litellm/proxy/db/master_key_migration.py
…_permit_weak_or_unset_master_key so the name says exactly what it permits
ryan-crabbe-berri added a commit to BerriAI/litellm-docs that referenced this pull request Sep 20, 2026
…mit_weak_or_unset_master_key

BerriAI/litellm#42019 renamed the override so the name says exactly what it permits. The environment variable is now LITELLM_DANGEROUSLY_PERMIT_WEAK_OR_UNSET_MASTER_KEY and the YAML key is general_settings.dangerously_permit_weak_or_unset_master_key

Update the general_settings example, the general_settings reference row and the environment variable row. The behaviour is unchanged
@ryan-crabbe-berri

Copy link
Copy Markdown
Contributor Author

@greptile re review

@ryan-crabbe-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit ecf1751. Configure here.

@mateo-berri
mateo-berri merged commit b6dd3d9 into main Sep 20, 2026
96 of 97 checks passed
@mateo-berri
mateo-berri deleted the litellm_master_key_boot_enforcement branch September 20, 2026 02:04
mateo-berri added a commit to BerriAI/litellm-docs that referenced this pull request Sep 20, 2026
…234 (#1577)

* docs(proxy): document the master key boot check and rotating off sk-1234

The proxy is about to refuse to start when the master key is not set, is empty, or is sk-1234, and its error message links to /docs/proxy/master_key_rotations#proxy-refuses-to-start

Add that section to the rotation page with the two ways out: swap and restart when a salt key is set or there is no database, and boot once with the local development override, re-encrypt through POST /key/regenerate, then swap when sk-1234 is also the encryption key

Add dangerously_allow_unsafe_proxy and LITELLM_DANGEROUSLY_ALLOW_UNSAFE_PROXY to the config settings reference, and say on the master_key rows that the proxy will not start without a real key

Fix the existing regenerate example, whose body key had been replaced with a virtual key placeholder. The proxy only rotates the master key when key is the current master key

* docs(proxy): generate the new key first and save it only after the regenerate call

The proxy's error no longer prints the save-to-.env command when sk-1234 is also the encryption key. It prints a command that only generates a key and says to save it once this guide says to, so the rotation steps now follow that order

Back up the database first, generate the key without saving it anywhere the proxy reads, boot once on the old key with the override, call POST /key/regenerate, then stop the proxy right away and only then set LITELLM_MASTER_KEY. The running process still holds the old key after the call and cannot decrypt the re-encrypted rows

Also warn that the regenerate response echoes the new master key in plaintext, repeat that the regenerate call must not be used when LITELLM_SALT_KEY is set, and mirror the error's "make sure the config reads the key from the environment" wording

* docs(proxy): cover the refusal variant for an already exported master key

The startup error now picks its "set a new key" step by whether LITELLM_MASTER_KEY is already set in the proxy's environment. When it is not set, the error still prints the command that generates a key and appends it to .env. When it is set to an unsafe value, the error prints a generate-only command and says to put the new key in place of the current value wherever that is set

Show both commands in the no-rotation case and explain why appending to .env does not work there: the proxy loads .env without overriding, so a value already exported in the environment wins. Say the same in "Where to set the new key", and stop describing the rotation case by the generate-only command, since it is no longer unique to it

* docs(proxy): replace the override and regenerate steps with the boot-time migration

Getting out of a refused boot no longer needs the local development override or a call to POST /key/regenerate. On refusal the proxy counts the stored values that decrypt under the unsafe key, and when there are any it tells the user to set LITELLM_MIGRATE_FROM_MASTER_KEY to the old key next to a new LITELLM_MASTER_KEY. The next boot re-encrypts those values before serving traffic and logs when the variable can be deleted

Rewrite the section around that: the case where nothing is encrypted with the old key, the migration steps with the log lines to expect, what happens when the variable is left set, when the master key is still unsafe, when a value changes during the migration, and why LITELLM_SALT_KEY must not be added halfway. Add a short subsection on what the migration covers and that it also works as an offline alternative to the regenerate call, and point to it from the existing regenerate section

Add LITELLM_MIGRATE_FROM_MASTER_KEY to the environment variable reference. The general POST /key/regenerate docs for a running proxy and the salt key warning stay as they were

* docs(proxy): match the final refusal wording and cover a failed migration

Quote the refusal lead as it is now wrapped, and say the unreachable-database variant only changes the first line

Add the case where a database error interrupts the migration: the proxy keeps starting, logs a WARNING that the values cannot be read until the migration succeeds, and the fix is to keep LITELLM_MIGRATE_FROM_MASTER_KEY set and restart once the database is reachable

Also say that the migration goes through the primary database even with a read replica configured, and that it skips tables or columns an older schema does not have, so it works before and after a schema upgrade

* docs(proxy): mark the error placeholder as code so MDX builds

* docs(proxy): a failed master key migration stops the boot

The proxy no longer keeps starting when the boot-time migration hits a database error. It logs the same WARNING and exits with a non-zero status, so a worker never serves with stored values it cannot read. The exception is a connection outage with general_settings.allow_requests_on_db_unavailable on, where the boot continues as it already does for the database connection

Say that restarting with the same two variables resumes a partial migration, and move the quoted log line into a code block so its <ErrorType: message> placeholder cannot be parsed as JSX

* docs(proxy): rename the local development override to dangerously_permit_weak_or_unset_master_key

BerriAI/litellm#42019 renamed the override so the name says exactly what it permits. The environment variable is now LITELLM_DANGEROUSLY_PERMIT_WEAK_OR_UNSET_MASTER_KEY and the YAML key is general_settings.dangerously_permit_weak_or_unset_master_key

Update the general_settings example, the general_settings reference row and the environment variable row. The behaviour is unchanged

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
pull Bot pushed a commit to SamAcctX/litellm that referenced this pull request Sep 21, 2026
…r key

BerriAI#42019 made the proxy refuse to start on a publicly known master key, and the
session fixture in tests/unified_google_tests started its in-process proxy with
sk-1234, so six tests errored in setup before reaching a provider

Give the fixture, the config it loads, and the SDK client the same non-default
key instead of the override the other harnesses took, so the boot check stays
live in this suite
pull Bot pushed a commit to stnxo2023/litellm that referenced this pull request Sep 21, 2026
The install smoke test starts the proxy on test_config_no_auth.yaml, which has no master key on purpose, and BerriAI#42019's boot check now refuses that, so the three installing_litellm_on_python jobs have been red on main since 2026-09-20. Pass the documented local-dev override to the proxy child so the test keeps its no-auth config and the boot check stays as it is
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants