Repository navigation
feat(proxy)!: refuse to start with an unset, empty, or publicly known master key - #42019
Conversation
… master key The proxy used to boot with no master key (every request accepted without authentication) and with sk-1234, the key every example used. It now stops at startup, before it connects to the database, and prints how to fix it: where the bad key came from, a copy-pastable command that generates a secure key, and, when the public key is also encrypting a database, a link to the rotation guide general_settings.dangerously_allow_unsafe_proxy: true or LITELLM_DANGEROUSLY_ALLOW_UNSAFE_PROXY=true starts the proxy anyway, for local development. CI and test boots that rely on sk-1234 or on no key set it BREAKING CHANGE: deployments with no master key, an empty one, or sk-1234 no longer start until they set a real key or opt in to the override
…y reads the environment
…they save a new key
…0 and keep the lazy OpenAPI snapshot as generated by CI's Python
|
Codecov Report❌ Patch coverage is 📢 Thoughts on this report? Let us know! |
| return source.config_file_path or "your config" | ||
|
|
||
|
|
||
| def _source_line(refusal: UnsafeMasterKeyRefused) -> str: |
… sites and isolate the boot test from a leaked scheduler
… it in place, because it wins over .env
|
@greptile re review |
| command: | | ||
| docker run -d \ | ||
| -p 4001:4000 \ | ||
| -e LITELLM_DANGEROUSLY_ALLOW_UNSAFE_PROXY=true \ |
There was a problem hiding this comment.
nit: LITELLM_DANGEROUSLY_PERMIT_WEAK_MASTER_KEY=true <- I feel like this is more descriptive
There was a problem hiding this comment.
Renamed in ecf1751 to LITELLM_DANGEROUSLY_PERMIT_WEAK_OR_UNSET_MASTER_KEY (YAML dangerously_permit_weak_or_unset_master_key), since the override also covers an unset or empty key, not only a weak one
There was a problem hiding this comment.
late to the party and bit of a nit, but we explicitly are not checking key strength, so calling it "weak" is a misnomer. I think LITELLM_DANGEROUSLY_PERMIT_DEFAULT_OR_UNSET_MASTER_KEY would have been better, but I imagine changing this after merge is a lot more trouble than it's worth
| PUBLICLY_KNOWN_MASTER_KEYS: Final = frozenset({"sk-1234"}) | ||
| ROTATION_DOCS_URL: Final = "https://docs.litellm.ai/docs/proxy/master_key_rotations#proxy-refuses-to-start" | ||
| _NEW_MASTER_KEY: Final = "sk-$(openssl rand -hex 32)" | ||
| GENERATE_MASTER_KEY_COMMAND: Final = f'echo "{MASTER_KEY_ENV_VAR}={_NEW_MASTER_KEY}" | tee -a .env' |
There was a problem hiding this comment.
maybe use echo "${MASTER_KEY_ENV_VAR}=${_NEW_MASTER_KEY}" >> .env so it doesn't emit the master key to stdout for the user...?
|
bugbot run |
…ROM_MASTER_KEY so an unsafe key can be replaced while the proxy refuses to start Rotating through POST /key/regenerate needs a running proxy, which a refused boot does not have. The refusal now counts the stored values that decrypt under the unsafe key. When there are none it only asks for a new key. When there are some it also asks for LITELLM_MIGRATE_FROM_MASTER_KEY, and the next boot with a safe key re-encrypts them and logs that the variable can be deleted. Leaving the variable set afterwards is a no-op with one notice.
| return numbered if refusal.migration is None else f"{_migration_lead(refusal.migration)}\n{numbered}" | ||
|
|
||
|
|
||
| def _config_steps(source: MasterKeySource) -> tuple[str, ...]: |
| return NothingToMigrate.NOTHING_ENCRYPTED_WITH_PREVIOUS_KEY if another_worker_migrated_everything else migrated | ||
|
|
||
|
|
||
| def describe_outcome(outcome: MigrationOutcome) -> str: |
… ciphertext during the master key migration A string such as "*" or "..." has no base64 characters, so it decoded to no bytes and read as an empty plaintext under any key. The migration would have counted it and overwritten it with a ciphertext of the empty string. Also read from the writer database instead of a read replica, report a database error during the migration instead of crashing the boot, skip columns the connected schema lacks across every schema on the search path, cap the JSON walk depth for the recursion detector, and move the boot wiring into one tested function.
| ReplaceCiphertext = Callable[[str], str | None] | ||
|
|
||
|
|
||
| def replace_ciphertexts(value: JsonValue, replacement_for: ReplaceCiphertext, depth: int = 0) -> tuple[JsonValue, int]: |
|
@greptile re review |
|
bugbot run |
…ls, unless allow_requests_on_db_unavailable tolerates the outage
|
@greptile re review |
|
bugbot run |
…_permit_weak_or_unset_master_key so the name says exactly what it permits
…mit_weak_or_unset_master_key BerriAI/litellm#42019 renamed the override so the name says exactly what it permits. The environment variable is now LITELLM_DANGEROUSLY_PERMIT_WEAK_OR_UNSET_MASTER_KEY and the YAML key is general_settings.dangerously_permit_weak_or_unset_master_key Update the general_settings example, the general_settings reference row and the environment variable row. The behaviour is unchanged
|
@greptile re review |
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit ecf1751. Configure here.
…234 (#1577) * docs(proxy): document the master key boot check and rotating off sk-1234 The proxy is about to refuse to start when the master key is not set, is empty, or is sk-1234, and its error message links to /docs/proxy/master_key_rotations#proxy-refuses-to-start Add that section to the rotation page with the two ways out: swap and restart when a salt key is set or there is no database, and boot once with the local development override, re-encrypt through POST /key/regenerate, then swap when sk-1234 is also the encryption key Add dangerously_allow_unsafe_proxy and LITELLM_DANGEROUSLY_ALLOW_UNSAFE_PROXY to the config settings reference, and say on the master_key rows that the proxy will not start without a real key Fix the existing regenerate example, whose body key had been replaced with a virtual key placeholder. The proxy only rotates the master key when key is the current master key * docs(proxy): generate the new key first and save it only after the regenerate call The proxy's error no longer prints the save-to-.env command when sk-1234 is also the encryption key. It prints a command that only generates a key and says to save it once this guide says to, so the rotation steps now follow that order Back up the database first, generate the key without saving it anywhere the proxy reads, boot once on the old key with the override, call POST /key/regenerate, then stop the proxy right away and only then set LITELLM_MASTER_KEY. The running process still holds the old key after the call and cannot decrypt the re-encrypted rows Also warn that the regenerate response echoes the new master key in plaintext, repeat that the regenerate call must not be used when LITELLM_SALT_KEY is set, and mirror the error's "make sure the config reads the key from the environment" wording * docs(proxy): cover the refusal variant for an already exported master key The startup error now picks its "set a new key" step by whether LITELLM_MASTER_KEY is already set in the proxy's environment. When it is not set, the error still prints the command that generates a key and appends it to .env. When it is set to an unsafe value, the error prints a generate-only command and says to put the new key in place of the current value wherever that is set Show both commands in the no-rotation case and explain why appending to .env does not work there: the proxy loads .env without overriding, so a value already exported in the environment wins. Say the same in "Where to set the new key", and stop describing the rotation case by the generate-only command, since it is no longer unique to it * docs(proxy): replace the override and regenerate steps with the boot-time migration Getting out of a refused boot no longer needs the local development override or a call to POST /key/regenerate. On refusal the proxy counts the stored values that decrypt under the unsafe key, and when there are any it tells the user to set LITELLM_MIGRATE_FROM_MASTER_KEY to the old key next to a new LITELLM_MASTER_KEY. The next boot re-encrypts those values before serving traffic and logs when the variable can be deleted Rewrite the section around that: the case where nothing is encrypted with the old key, the migration steps with the log lines to expect, what happens when the variable is left set, when the master key is still unsafe, when a value changes during the migration, and why LITELLM_SALT_KEY must not be added halfway. Add a short subsection on what the migration covers and that it also works as an offline alternative to the regenerate call, and point to it from the existing regenerate section Add LITELLM_MIGRATE_FROM_MASTER_KEY to the environment variable reference. The general POST /key/regenerate docs for a running proxy and the salt key warning stay as they were * docs(proxy): match the final refusal wording and cover a failed migration Quote the refusal lead as it is now wrapped, and say the unreachable-database variant only changes the first line Add the case where a database error interrupts the migration: the proxy keeps starting, logs a WARNING that the values cannot be read until the migration succeeds, and the fix is to keep LITELLM_MIGRATE_FROM_MASTER_KEY set and restart once the database is reachable Also say that the migration goes through the primary database even with a read replica configured, and that it skips tables or columns an older schema does not have, so it works before and after a schema upgrade * docs(proxy): mark the error placeholder as code so MDX builds * docs(proxy): a failed master key migration stops the boot The proxy no longer keeps starting when the boot-time migration hits a database error. It logs the same WARNING and exits with a non-zero status, so a worker never serves with stored values it cannot read. The exception is a connection outage with general_settings.allow_requests_on_db_unavailable on, where the boot continues as it already does for the database connection Say that restarting with the same two variables resumes a partial migration, and move the quoted log line into a code block so its <ErrorType: message> placeholder cannot be parsed as JSX * docs(proxy): rename the local development override to dangerously_permit_weak_or_unset_master_key BerriAI/litellm#42019 renamed the override so the name says exactly what it permits. The environment variable is now LITELLM_DANGEROUSLY_PERMIT_WEAK_OR_UNSET_MASTER_KEY and the YAML key is general_settings.dangerously_permit_weak_or_unset_master_key Update the general_settings example, the general_settings reference row and the environment variable row. The behaviour is unchanged --------- Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
…r key BerriAI#42019 made the proxy refuse to start on a publicly known master key, and the session fixture in tests/unified_google_tests started its in-process proxy with sk-1234, so six tests errored in setup before reaching a provider Give the fixture, the config it loads, and the SDK client the same non-default key instead of the override the other harnesses took, so the boot check stays live in this suite
The install smoke test starts the proxy on test_config_no_auth.yaml, which has no master key on purpose, and BerriAI#42019's boot check now refuses that, so the three installing_litellm_on_python jobs have been red on main since 2026-09-20. Pass the documented local-dev override to the proxy child so the test keeps its no-auth config and the boot check stays as it is
TLDR
Problem this solves:
sk-1234is public, and it boots as the admin keyHow it solves it:
sk-1234master key.env, and the error says soLITELLM_MIGRATE_FROM_MASTER_KEY, and the next boot re-encrypts them with the new keydangerously_permit_weak_or_unset_master_keystarts anyway, for local developmentThis is a breaking change on purpose. A deployment that runs on no master key, an empty one, or
sk-1234stops booting after the upgrade until it sets a real key or setsLITELLM_DANGEROUSLY_PERMIT_WEAK_OR_UNSET_MASTER_KEY=true(orgeneral_settings.dangerously_permit_weak_or_unset_master_key: true)User Flow
Before: an operator who copies the quick start ends up with a proxy that anyone can administer, and nothing tells them
LITELLM_MASTER_KEY=sk-1234, the value every example used, or with no master key at allAuthorization: Bearer sk-1234, or with noAuthorizationheader at all when no key is set, and gets 200sk-1234that caller is the proxy admin, so they can also create keys, read spend, and change modelsAfter: the same start command stops with instructions, and the proxy only serves once it has a real key
LITELLM_MASTER_KEY=sk-1234, or with no master key at allecho "LITELLM_MASTER_KEY=sk-$(openssl rand -hex 32)" | tee -a .envsk-1234already exported it isecho "sk-$(openssl rand -hex 32)", and they put the output in place ofsk-1234where it is setsk-1234sk-1234, step 2 also tells them to setLITELLM_MIGRATE_FROM_MASTER_KEY=sk-1234Re-encrypting N stored value(s)...and thenDone re-encrypting N stored value(s) with the new master key. You may now delete the LITELLM_MIGRATE_FROM_MASTER_KEY environment variable.Relevant issues
Companion PRs: #42011 removes
sk-1234from the configs, READMEs, and UI snippets we ship, and BerriAI/litellm-docs#1577 adds the#proxy-refuses-to-startsection the error links toAffected release
Linear ticket
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Every case is a live proxy started from the
litellmconsole script on port 4893 (4894 for the database cases, against a throwawaypostgres:16container on 55893), withLITELLM_MODE=PRODUCTIONso no.envis read. Before runs thelitellm/package exported from the merge base, After runs the PR tip. Generated keys are cut to their first 6 characters, and the long temp path of the config is shortened to./config.yamlconfig.yaml:Before (b946d12)
sk-1234 as the master key
LITELLM_MASTER_KEY=sk-1234 litellm --port 4893 --config config.yamlstarts and servescurl -s -o /dev/null -w '%{http_code}' http://127.0.0.1:4893/v1/models -H 'Authorization: Bearer sk-1234'prints200No master key anywhere
litellm --port 4893 --config config.yamlwithLITELLM_MASTER_KEYunset starts and servescurl -s -o /dev/null -w '%{http_code}' http://127.0.0.1:4893/v1/modelswith noAuthorizationheader prints200Run the printed command, then boot
echo "LITELLM_MASTER_KEY=sk-$(openssl rand -hex 32)" | tee -a .envprintsLITELLM_MASTER_KEY=sk-9cd...litellm --port 4893 --config config.yamlstarts and serves/v1/modelsprints200with the generated key and400withBearer sk-1234sk-1234 with the override
LITELLM_MASTER_KEY=sk-1234 LITELLM_DANGEROUSLY_PERMIT_WEAK_OR_UNSET_MASTER_KEY=true litellm --port 4893 --config config.yamlstarts and serves; the variable means nothing on this commit/v1/modelsprints200withBearer sk-1234sk-1234 with a database and no salt key
DATABASE_URL=postgresql://postgres:postgres@127.0.0.1:55893/rot_before LITELLM_MASTER_KEY=sk-1234 litellm --port 4894 --config config.yamlstarts and serves, withDatasource "client": PostgreSQL database "rot_before"in the logcurl -s -X POST http://127.0.0.1:4894/credentials -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{"credential_name":"before-cred","credential_values":{"api_key":"sk-fake-credential-key"},"credential_info":{"custom_llm_provider":"openai"}}'returns{"success":true,"message":"Credential created successfully"}, so the secret is now stored encrypted with the public keycurl -s http://127.0.0.1:4894/key/list -H 'Authorization: Bearer sk-1234'returns HTTP 200{"keys":[],"total_count":0,"current_page":1,"total_pages":0}After (ecf1751)
sk-1234 as the master key
LITELLM_MASTER_KEY=sk-1234 litellm --port 4893 --config config.yaml; echo "exit code $?"printsexit code 3, and the last thing on screen is:/v1/modelscall from Before cannot be madeNo master key anywhere
litellm --port 4893 --config config.yaml; echo "exit code $?"withLITELLM_MASTER_KEYunset printsexit code 3, and the last thing on screen is:/v1/modelscall from Before cannot be madeRun the printed command, then boot
echo "LITELLM_MASTER_KEY=sk-$(openssl rand -hex 32)" | tee -a .envprintsLITELLM_MASTER_KEY=sk-36d...(64 hex characters in total)litellm --port 4893 --config config.yamlstarts and serves/v1/modelsprints200with the generated key and400withBearer sk-1234sk-1234 with the override
LITELLM_MASTER_KEY=sk-1234 LITELLM_DANGEROUSLY_PERMIT_WEAK_OR_UNSET_MASTER_KEY=true litellm --port 4893 --config config.yamlstarts and servesWARNING: master_key_boot_check.py:122 - dangerously_permit_weak_or_unset_master_key is on, so the proxy is starting with a publicly known master key. Never run this outside local development., and/v1/modelsprints200withBearer sk-1234sk-1234 with a database and no salt key
Setup, done once with
LITELLM_DANGEROUSLY_PERMIT_WEAK_OR_UNSET_MASTER_KEY=trueso the proxy boots onsk-1234against an emptypostgres:16databasemig:POST /model/newwith a fake providerapi_key,POST /credentials,POST /config/updatewith one environment variable,POST /team/newplusPOST /team/migrate-test-team/callbackwith fake Langfuse keys, andPOST /key/generatefor a virtual keysk-bWW.... All returned HTTP 200, so the database now holds secrets encrypted withsk-1234. One plaintextLiteLLM_Configrow,{"note": "...", "separator": "-", "allowed_routes": ["*"]}, is inserted withpsqlas a control, because"-","*"and"..."base64-decode to no bytesDATABASE_URL=postgresql://postgres:postgres@127.0.0.1:55893/mig LITELLM_MASTER_KEY=sk-1234 litellm --port 4894 --config config.yaml; echo "exit code $?"printsexit code 3. The proxy counted the stored values that decrypt undersk-1234(6, the plaintext row is not one of them), and the steps now ask for the key to migrate from:LITELLM_MIGRATE_FROM_MASTER_KEY=sk-1234added andLITELLM_MASTER_KEYstillsk-1234printsexit code 3againecho "sk-$(openssl rand -hex 32)"printssk-c8b....DATABASE_URL=... LITELLM_MASTER_KEY=sk-c8b... LITELLM_MIGRATE_FROM_MASTER_KEY=sk-1234 litellm --port 4894 --config config.yamlstarts and serves. These are the 2nd and 3rd warnings in the boot log, right after thetrusted_proxy_rangesone:api_keyfrom1LY6Sw2u...tovMqUp0Eo..., the credential's fromMSXseJ5e...toAyEwY01y..., the environment variable fromMonDssnS...toNbM_5K1E..., and the team's Langfuse secret fromlitellm_enc::uGrGNVE0...tolitellm_enc::yxNzktDx..., which kept itslitellm_enc::marker. The plaintext control row is byte for byte the sameError decrypting valuelines in the log:curl -s http://127.0.0.1:4894/model/info -H 'Authorization: Bearer sk-c8b...'returns HTTP 200 with the DB modelmigrate-test-modelcurl -s http://127.0.0.1:4894/credentials/by_name/migrate-test-cred -H 'Authorization: Bearer sk-c8b...'returns HTTP 200{"credential_name":"migrate-test-cred","credential_info":{"custom_llm_provider":"openai"},"credential_values":{"api_key":"sk****st"}}curl -s http://127.0.0.1:4894/models -H 'Authorization: Bearer sk-bWW...', the virtual key created undersk-1234, still returns HTTP 200 withmigrate-test-modelLITELLM_MASTER_KEY=sk-c8b.... The log has 0 migration lines and 0 decryption errors, and the three calls from step 5 return the same HTTP 200 responsessk-1234 with a database that holds nothing encrypted
DATABASE_URL=postgresql://postgres:postgres@127.0.0.1:55893/mig_empty LITELLM_MASTER_KEY=sk-1234 litellm --port 4894 --config config.yaml; echo "exit code $?"printsexit code 3.mig_emptyhas the schema and no stored secrets, so the steps are the same as the first case, with no mention ofLITELLM_MIGRATE_FROM_MASTER_KEY:migdatabase whenLITELLM_SALT_KEYis set, because the salt key is what encrypts stored values thenDATABASE_URLpointing at a port nothing listens on, the refusal takes about 10 seconds longer, starts withYour database could not be checked for values encrypted with this master key, and prints the migration stepsA database on an older schema
mig_partialis a copy of the migratedmigwithDROP TABLE "LiteLLM_SSOIdentityAssertion" CASCADE, so one of the tables the migration knows about does not existDATABASE_URL=.../mig_partial LITELLM_MASTER_KEY=sk-62b... LITELLM_MIGRATE_FROM_MASTER_KEY=sk-c8b... litellm --port 4894 --config config.yaml, with a second generated key, starts and serves. The log has18:58:36 ... Re-encrypting 6 stored value(s)...followed by theDone re-encrypting 6 stored value(s)line, andcurl -s http://127.0.0.1:4894/credentials/by_name/migrate-test-cred -H 'Authorization: Bearer sk-62b...'returns HTTP 200 with"api_key":"sk****st"and 0 decryption errors. The missing table is skipped, and the variable works for any previous key, not onlysk-1234A write that fails during the migration
mig_failis a copy of the migratedmigwith a trigger that makes everyUPDATEon"LiteLLM_CredentialsTable"raisesimulated write failureDATABASE_URL=.../mig_fail LITELLM_MASTER_KEY=sk-168... LITELLM_MIGRATE_FROM_MASTER_KEY=sk-c8b... litellm --port 4894 --config config.yaml; echo "exit code $?"printsexit code 3. The proxy does not serve with values it cannot read, and the log says what to do:curl -s http://127.0.0.1:4894/credentials/by_name/migrate-test-cred -H 'Authorization: Bearer sk-168...'returns HTTP 200 with"api_key":"sk****st",GET /model/inforeturns HTTP 200 with the DB model, and the log has 0 decryption errorsType
New Feature
Caveats (if any)
Severe
sk-1234stop booting on upgradeLITELLM_SALT_KEY, swapping the key withoutLITELLM_MIGRATE_FROM_MASTER_KEYmakes stored credentials undecryptable"*","-"or""base64-decodes to no bytes, which the legacy cipher reads as an empty plaintext under any key. The migration rejects those before decrypting, and the proof below keeps such a row unchangedallow_requests_on_db_unavailableis on is toleratedquery_rawraise an engineAttributeError, which the rule does not count as a connection outage, so the boot stopped even with the flag on. Without the flag, a database that is already down at boot fails earlier, in the Prisma setup, with exit code 3Medium
--num_workers 4) prints the fix once per worker and the parent exits with code 0sk-1234, a database is configured, and no salt key is setLITELLM_MIGRATE_FROM_MASTER_KEYis set, every boot reads the secret-bearing columns once to look for values under the previous keylitellm_enc::marker, so the big tables are not loadedPOST /key/regeneratewithnew_master_keydoes/key/regenerateis not changed herePOST /key/regeneratewithnew_master_keywhileLITELLM_SALT_KEYis set leaves stored secrets unreadable under both keysLow
.envstep only takes effect where the proxy loads.env(repo checkouts, docker composeenv_file), so the message also says to pass the env varsk-12345orpasswordstill boot.env, where appending a second line would also have workedmaster_key: sk-1234plus the same value in the env var is reported as coming from the environment first, then from the config on the next bootsk-1234or on no key now set the override; they were not rewritten to use a real keytests/e2e/ui/run_e2e.sh,create_proxy_test_clientintests/test_litellm/proxy/conftest.py, and the session fixture intests/mcp_tests/test_proxy_mcp_e2e.pyrender.yamlnow asks Render to generateLITELLM_MASTER_KEY, so one-click deploys keep bootingQA runbook
No e2e test was added or changed. The only edit under
tests/e2eistests/e2e/ui/run_e2e.sh, which now exportsLITELLM_DANGEROUSLY_PERMIT_WEAK_OR_UNSET_MASTER_KEY=truenext to itssk-1234so the UI suite's proxy still bootsFinal Attestation