Skip to content

litellm 0.7.3: v2 migration resolver + drop upstream demo models - #43

Closed
dan2k3k4 wants to merge 3 commits into
mainfrom
litellm-v2-migrations-and-no-demo-models
Closed

dan2k3k4 wants to merge 3 commits into
mainfrom
litellm-v2-migrations-and-no-demo-models

Conversation

@dan2k3k4

@dan2k3k4 dan2k3k4 commented Aug 11, 2026 •

Copy link
Copy Markdown
Member

What does this PR do?

Stops LiteLLM's database migrations from wiping LiteLLM_ProxyModelTable (all DB-managed models) during rollouts and restarts, and removes upstream's demo models from the default config. Chart 0.7.2 → 0.7.3.

Root cause. The destructive actor is LiteLLM's default v1 migration resolver, whose diff-and-force recovery can drop and recreate tables. LiteLLM's own CLI warns:

Using default (v1) migration resolver. If your deployment has seen schema thrashing during rolling deploys, try --use_v2_migration_resolver (safer: avoids the diff-and-force recovery that caused the thrash).

Where it actually runs shaped this PR:

  • Proxy pods run it on every startup. The chart's deployment never sets DISABLE_SCHEMA_UPDATE and we don't set disable_prisma_schema_update, so each pod start runs the DB setup with the v1 resolver — old and new pods contend during rolling deploys, and plain restarts run it too. This matches models vanishing on restarts where no helm upgrade happens at all.
  • The upstream migration job is nearly inert under plain Helm. Its hook annotations are ArgoCD-only (migrationJob.hooks.helm.enabled defaults to false in the pinned chart) and a completed Job is immutable, so on our Sveltos/Helm-driven clusters it ran once at first install and never again. Its entrypoint (python litellm/proxy/prisma_migration.py) is also hard-wired to v1 with no env-var/values hookup for the flag, so it can't be fixed via values.

Changes:

  1. args: [--config, /etc/litellm/config.yaml, --use_v2_migration_resolver] — the chart's default args plus the flag, so every proxy pod migrates with the v2 resolver. This is the actual fix for the purge. Flag verified present in v1.96.2 (the tag all clusters pin): [Feature] Proxy: opt-in v2 migration resolver BerriAI/litellm#26194.
  2. templates/migrations-job.yaml — replacement migrations job running litellm --skip_server_startup --use_v2_migration_resolver as a pre-upgrade-only helm hook, so schema migrations complete before pods roll. pre-install is deliberately omitted: the DB secret comes from this chart's own ExternalSecret, which can't exist while a pre-install hook blocks the first install — on first install the pods create the schema themselves (v2, via args). ttlSecondsAfterFinished: 86400 keeps job logs for a day (upstream's 120s TTL is why the us1 logs were already gone). Renders only while litellm-helm.migrationJob.enabled=false, so re-enabling upstream never yields two jobs.
  3. migrationJob.enabled: false — disables the upstream v1 job.
  4. proxy_config.model_list: [] — upstream's default model_list ships two demo models (gpt-3.5-turbo, fake-openai-endpoint) that leak into any cluster without its own model_list (currently visible on us1 and de1). Catalog-managed clusters get models from api.amazee.ai; clusters that define model_list in their cluster file override this and are unaffected.

Related issue(s)

  • Relates to the 2026-08-10 us1 incident (DB-managed models wiped during rollout; keys and access groups survived).

Related changes

  • Companion PR to amazeeai-k0rdent-clusters (to be opened once this merges and chart 0.7.3 is published): ServiceTemplate litellm-0.5.5 → chart 0.7.3 + chain, bumping us1 and de1 only at first.

Type of changes

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Refactoring (non-breaking change which adds no functionality)
  • Breaking change (fix or feature that would cause existing functionality to change)

How has this been tested?

  • helm template (with dependency build) on the first revision: exactly one migrations job rendered (the v2 one), model_list: [] in the proxy config, zero demo-model references. Re-render pending for the latest commits (deployment args + pre-upgrade-only hook).
  • --use_v2_migration_resolver and --skip_server_startup verified present in the v1.96.2 image ([Feature] Proxy: opt-in v2 migration resolver BerriAI/litellm#26194, add skip server startup flag to cli BerriAI/litellm#10665).
  • Verified no cluster sets args or numWorkers, so the explicit args override changes nothing besides adding the flag.
  • Pinned chart 1.83.14-stable.patch.3 inspected from ghcr to confirm hook defaults (argocd: true, helm: false) and that the deployment sets no DISABLE_SCHEMA_UPDATE.

How and when this is going to be rolled out?

Staged via the companion clusters PR: us1 and de1 first; prod clusters stay on litellm-0.5.4 until this proves out. The upgrade that delivers the fix is the last one where old pods still run v1 — take a DB snapshot before rolling it. If anything destructive happens again, the amazee.ai reconcile cron restores models within 5 minutes, and this time the migration-job logs will still exist. After one or two clean upgrade cycles, the 5-minute model-sync workaround can be retired.

Checklist

  • Assign yourself as the Assignee of this MR.
  • Review your changes before adding any reviewers.
  • (If applicable) Documentation is updated (README, Notion etc.)
  • If you have multiple commits, please combine them into a few logically organized commits by squashing them.
  • This PR contains substantial use of AI-assisted coding tools.

Status

  • Ready for Review

- Replace the upstream migrations job with one that runs
  'litellm --skip_server_startup --use_v2_migration_resolver'. Upstream's
  entrypoint is pinned to the v1 resolver, whose diff-and-force recovery
  can drop and recreate tables during rolling deploys — this wiped
  LiteLLM_ProxyModelTable on amazeeai-us1. Job pods are kept for a day
  (upstream's 120s TTL destroyed the logs before diagnosis was possible).
- Default proxy_config.model_list to [] so upstream's demo models
  (gpt-3.5-turbo, fake-openai-endpoint) never leak into clusters that get
  their models from the amazee.ai model catalog. Clusters that define
  their own model_list are unaffected.
@dan2k3k4
dan2k3k4 requested a review from a team as a code owner August 11, 2026 11:02
@dan2k3k4
dan2k3k4 marked this pull request as draft August 11, 2026 11:11
…rade-only

The purge of LiteLLM_ProxyModelTable happens at proxy pod startup, not in
the migration job: the chart's deployment never disables schema updates, so
every pod start runs litellm's DB setup with the default v1 resolver, whose
diff-and-force recovery can drop and recreate tables when pods contend
during rolling deploys or restarts. Upstream's migration job never re-ran
after install under plain Helm (ArgoCD-only hook annotations + immutable
completed Jobs), so it was not the actor.

- add --use_v2_migration_resolver to the deployment args so every pod
  migrates with the safe resolver (this is the actual fix)
- run the replacement job as a pre-upgrade-only hook: migrations complete
  before pods roll; pre-install dropped because the DB secret is created by
  this chart's own ExternalSecret and cannot exist while a pre-install hook
  blocks the first install
- correct the comments to reflect the real failure path
@dan2k3k4 dan2k3k4 self-assigned this Aug 14, 2026
@dan2k3k4 dan2k3k4 closed this Sep 8, 2026
@dan2k3k4
dan2k3k4 deleted the litellm-v2-migrations-and-no-demo-models branch September 8, 2026 09:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant