Conversation
- Replace the upstream migrations job with one that runs 'litellm --skip_server_startup --use_v2_migration_resolver'. Upstream's entrypoint is pinned to the v1 resolver, whose diff-and-force recovery can drop and recreate tables during rolling deploys — this wiped LiteLLM_ProxyModelTable on amazeeai-us1. Job pods are kept for a day (upstream's 120s TTL destroyed the logs before diagnosis was possible). - Default proxy_config.model_list to [] so upstream's demo models (gpt-3.5-turbo, fake-openai-endpoint) never leak into clusters that get their models from the amazee.ai model catalog. Clusters that define their own model_list are unaffected.
dan2k3k4
marked this pull request as draft
August 11, 2026 11:11
…rade-only The purge of LiteLLM_ProxyModelTable happens at proxy pod startup, not in the migration job: the chart's deployment never disables schema updates, so every pod start runs litellm's DB setup with the default v1 resolver, whose diff-and-force recovery can drop and recreate tables when pods contend during rolling deploys or restarts. Upstream's migration job never re-ran after install under plain Helm (ArgoCD-only hook annotations + immutable completed Jobs), so it was not the actor. - add --use_v2_migration_resolver to the deployment args so every pod migrates with the safe resolver (this is the actual fix) - run the replacement job as a pre-upgrade-only hook: migrations complete before pods roll; pre-install dropped because the DB secret is created by this chart's own ExternalSecret and cannot exist while a pre-install hook blocks the first install - correct the comments to reflect the real failure path
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
Stops LiteLLM's database migrations from wiping
LiteLLM_ProxyModelTable(all DB-managed models) during rollouts and restarts, and removes upstream's demo models from the default config. Chart0.7.2 → 0.7.3.Root cause. The destructive actor is LiteLLM's default v1 migration resolver, whose diff-and-force recovery can drop and recreate tables. LiteLLM's own CLI warns:
Where it actually runs shaped this PR:
DISABLE_SCHEMA_UPDATEand we don't setdisable_prisma_schema_update, so each pod start runs the DB setup with the v1 resolver — old and new pods contend during rolling deploys, and plain restarts run it too. This matches models vanishing on restarts where no helm upgrade happens at all.migrationJob.hooks.helm.enableddefaults tofalsein the pinned chart) and a completedJobis immutable, so on our Sveltos/Helm-driven clusters it ran once at first install and never again. Its entrypoint (python litellm/proxy/prisma_migration.py) is also hard-wired to v1 with no env-var/values hookup for the flag, so it can't be fixed via values.Changes:
args: [--config, /etc/litellm/config.yaml, --use_v2_migration_resolver]— the chart's default args plus the flag, so every proxy pod migrates with the v2 resolver. This is the actual fix for the purge. Flag verified present inv1.96.2(the tag all clusters pin): [Feature] Proxy: opt-in v2 migration resolver BerriAI/litellm#26194.templates/migrations-job.yaml— replacement migrations job runninglitellm --skip_server_startup --use_v2_migration_resolveras a pre-upgrade-only helm hook, so schema migrations complete before pods roll.pre-installis deliberately omitted: the DB secret comes from this chart's own ExternalSecret, which can't exist while a pre-install hook blocks the first install — on first install the pods create the schema themselves (v2, via args).ttlSecondsAfterFinished: 86400keeps job logs for a day (upstream's 120s TTL is why the us1 logs were already gone). Renders only whilelitellm-helm.migrationJob.enabled=false, so re-enabling upstream never yields two jobs.migrationJob.enabled: false— disables the upstream v1 job.proxy_config.model_list: []— upstream's default model_list ships two demo models (gpt-3.5-turbo,fake-openai-endpoint) that leak into any cluster without its own model_list (currently visible on us1 and de1). Catalog-managed clusters get models from api.amazee.ai; clusters that definemodel_listin their cluster file override this and are unaffected.Related issue(s)
Related changes
amazeeai-k0rdent-clusters(to be opened once this merges and chart0.7.3is published): ServiceTemplatelitellm-0.5.5→ chart0.7.3+ chain, bumping us1 and de1 only at first.Type of changes
How has this been tested?
helm template(with dependency build) on the first revision: exactly one migrations job rendered (the v2 one),model_list: []in the proxy config, zero demo-model references. Re-render pending for the latest commits (deployment args + pre-upgrade-only hook).--use_v2_migration_resolverand--skip_server_startupverified present in thev1.96.2image ([Feature] Proxy: opt-in v2 migration resolver BerriAI/litellm#26194, add skip server startup flag to cli BerriAI/litellm#10665).argsornumWorkers, so the explicitargsoverride changes nothing besides adding the flag.1.83.14-stable.patch.3inspected from ghcr to confirm hook defaults (argocd: true,helm: false) and that the deployment sets noDISABLE_SCHEMA_UPDATE.How and when this is going to be rolled out?
Staged via the companion clusters PR: us1 and de1 first; prod clusters stay on
litellm-0.5.4until this proves out. The upgrade that delivers the fix is the last one where old pods still run v1 — take a DB snapshot before rolling it. If anything destructive happens again, the amazee.ai reconcile cron restores models within 5 minutes, and this time the migration-job logs will still exist. After one or two clean upgrade cycles, the 5-minute model-sync workaround can be retired.Checklist
Status