Skip to content

fix(proxy-extras): only spend a migrate-deploy attempt when a pass made no progress - #39506

Merged
mateo-berri merged 2 commits into
litellm_internal_stagingfrom
litellm_fix_v2_migration_resolver_attempt_accounting
Sep 3, 2026
Merged

mateo-berri merged 2 commits into
litellm_internal_stagingfrom
litellm_fix_v2_migration_resolver_attempt_accounting

Conversation

@mateo-berri

@mateo-berri mateo-berri commented Sep 3, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • A --use_prisma_db_push database can never boot under --use_v2_migration_resolver
  • Every recovery burned one of only 4 migrate attempts
  • The proxy exits before binding its port, so nothing serves

How it solves it:

  • Attempts are only spent when a pass made no progress
  • Baselining, and each migration newly marked applied, cost nothing
  • Timeouts, deadlock retries, and repeat recoveries still spend
  • A repeat of the same recovery still ends the run

User Flow

Before: an operator who first brought their database up with --use_prisma_db_push cannot switch that same database to --use_v2_migration_resolver, and the proxy never starts

  1. They start the proxy once with --use_prisma_db_push against an empty Postgres, and it serves POST https://litellm-domain/v1/chat/completions normally
  2. They restart the same proxy against the same database with --use_v2_migration_resolver instead
  3. Boot logs Schema exists but no migrations ledger — creating baseline, then three SQL hit idempotent error — marking applied and retrying lines
  4. Boot stops with Database migration cannot proceed. Database migration failed after 4 attempts and the process exits
  5. GET https://litellm-domain/health/readiness never answers because the port was never bound, and every request their apps send fails to connect

After: the same restart works through the objects the push already created and the proxy comes up

  1. They start the proxy once with --use_prisma_db_push against an empty Postgres, and it serves POST https://litellm-domain/v1/chat/completions normally
  2. They restart the same proxy against the same database with --use_v2_migration_resolver instead
  3. Boot logs the same baseline line, then works through each SQL hit idempotent error — marking applied and retrying in turn instead of stopping at the fourth
  4. Boot finishes and the proxy binds its port
  5. GET https://litellm-domain/health/readiness returns 200 with {"status":"healthy","db":"connected"}, and POST https://litellm-domain/v1/chat/completions serves normally again

Relevant issues

Linear ticket

Resolves LIT-6802

Type

🐛 Bug Fix

Caveats (if any)

Medium

  • Boot takes one migrate pass per pre-existing object
    • Roughly 25s each on the database this was measured against
    • Bounded by the migrations shipped, and it is a boot that now finishes

Low

  • attempt N in the logs now counts no-progress attempts, not passes
  • A baseline that fails to land now backs off before the retry, like every other no-progress path

Pre-Submission checklist

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Screenshots / Proof of Fix

Both legs run against the same Postgres 16 container, and each leg starts from a database that was just brought up by --use_prisma_db_push and then dropped and recreated so the second boot sees exactly the push-created state.

Shared setup:

docker run -d --name lit6802-pg-push -p 127.0.0.1:45871:5432 \
  -e POSTGRES_USER=lit6802 -e POSTGRES_PASSWORD=lit6802 -e POSTGRES_DB=lit6802push postgres:16

cat > config.yaml <<'YAML'
model_list:
  - model_name: gpt-5-mini
    litellm_params:
      model: openai/gpt-5-mini
      api_key: os.environ/OPENAI_API_KEY

general_settings:
  master_key: sk-lit6802
YAML

export DATABASE_URL='postgresql://lit6802:lit6802@127.0.0.1:45871/lit6802push'

Both legs are preceded by the same push boot, which is what creates the schema without a ledger:

python litellm/proxy/proxy_cli.py --config config.yaml --port 45881 --num_workers 2 --use_prisma_db_push
# serves normally, then is stopped
docker exec lit6802-pg-push psql -U lit6802 -d lit6802push -tAc \
  "SELECT count(*) FROM information_schema.tables WHERE table_schema='public';"
# 85
docker exec lit6802-pg-push psql -U lit6802 -d lit6802push -tAc \
  "SELECT coalesce(to_regclass('public._prisma_migrations')::text,'NO_LEDGER');"
# NO_LEDGER

Before (ff17e8b)

  1. Boot that same database with the v2 resolver:
python litellm/proxy/proxy_cli.py --config config.yaml --port 45881 --num_workers 2 --use_v2_migration_resolver
INFO - Using v2 migration resolver (--use_v2_migration_resolver)
INFO - Found 163 migrations at .../litellm-proxy-extras/litellm_proxy_extras
INFO - Schema exists but no migrations ledger — creating baseline
INFO - Generating baseline migration...
INFO - Marking baseline migration as applied...
Migration 0_init marked as applied.
INFO - Migration 20250329084805_new_cron_job_table SQL hit idempotent error — marking applied and retrying
INFO - Migration 20250806095134_rename_alias_to_server_name_mcp_table SQL hit idempotent error — marking applied and retrying
INFO - Migration 20260224203854_add_agent_object_permissions_table SQL hit idempotent error — marking applied and retrying
LiteLLM Proxy: Database migration cannot proceed. Database migration failed after 4 attempts ...
EXIT_CODE=2
  1. The port was never bound:
lsof -nP -iTCP:45881 -sTCP:LISTEN
(no output)
  1. Readiness never answers, so there is nothing to send a completion to:
curl -s -m 20 -w '\nHTTP %{http_code}\n' http://127.0.0.1:45881/health/readiness
HTTP 000
  1. The ledger is stuck three migrations in:
docker exec lit6802-pg-push psql -U lit6802 -d lit6802push -tAc \
  "SELECT count(*), count(finished_at) FROM _prisma_migrations;"
95|92

After (8572544)

  1. Same command against the same freshly push-created database:
python litellm/proxy/proxy_cli.py --config config.yaml --port 45881 --num_workers 2 --use_v2_migration_resolver
INFO - Using v2 migration resolver (--use_v2_migration_resolver)
INFO - Found 163 migrations at .../litellm-proxy-extras/litellm_proxy_extras
INFO - Schema exists but no migrations ledger — creating baseline
INFO - Marking baseline migration as applied...
Migration 0_init marked as applied.
INFO - Migration 20250329084805_new_cron_job_table SQL hit idempotent error — marking applied and retrying
INFO - Migration 20250806095134_rename_alias_to_server_name_mcp_table SQL hit idempotent error — marking applied and retrying
INFO - Migration 20260224203854_add_agent_object_permissions_table SQL hit idempotent error — marking applied and retrying
INFO - Migration 20260331000000_add_prompt_environment_and_created_by SQL hit idempotent error — marking applied and retrying
... 13 more, 17 recoveries in total, where the old code stopped after 3 ...
INFO:     Uvicorn running on http://0.0.0.0:45881 (Press CTRL+C to quit)
  1. The port is bound:
lsof -nP -iTCP:45881 -sTCP:LISTEN
python3.1 17999 mateo    7u  IPv4 0xf3d2a60da5fdae2f      0t0  TCP *:45881 (LISTEN)
python3.1 54519 mateo    7u  IPv4 0xf3d2a60da5fdae2f      0t0  TCP *:45881 (LISTEN)
python3.1 54520 mateo    7u  IPv4 0xf3d2a60da5fdae2f      0t0  TCP *:45881 (LISTEN)
  1. Readiness answers:
curl -s -m 20 -w '\nHTTP %{http_code}\n' http://127.0.0.1:45881/health/readiness
{"status":"healthy","db":"connected"}
HTTP 200
  1. A real gpt-5-mini completion through the recovered database:
curl -s -m 90 -w '\nHTTP %{http_code}\n' http://127.0.0.1:45881/v1/chat/completions \
  -H 'Authorization: Bearer sk-lit6802' -H 'Content-Type: application/json' \
  -d '{"model":"gpt-5-mini","messages":[{"role":"user","content":"Reply with exactly: ok LIT6802 refactor"}]}'
{"id":"chatcmpl-EJw0bXKB7b5ycwJs1NYkDix8gim8o","created":1788419925,"model":"gpt-5-mini","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"ok LIT6802 refactor","role":"assistant",...}}],"usage":{"completion_tokens":80,"prompt_tokens":17,"total_tokens":97,...}}
HTTP 200
  1. The ledger reached every migration (163 shipped plus the 0_init baseline), where the Before leg stopped at 92:
docker exec lit6802-pg-push psql -U lit6802 -d lit6802push -tAc \
  "SELECT count(*), count(finished_at) FROM _prisma_migrations;"
181|164

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

…de no progress

The v2 migration resolver gave `prisma migrate deploy` four attempts, and
every recovery path ended in a bare `continue`, so each one burned an attempt.
A database first brought up with `--use_prisma_db_push` has a full schema and
no migrations ledger, so the baseline spent attempt one and the first three
migrations whose objects already existed spent the rest. The proxy then exited
before binding its port, and that database could never be moved onto the
resolver.

The retry budget now counts only attempts that got nowhere. Creating the
baseline, and each migration newly marked applied, leaves the budget alone, so
a push-created database works through its pre-existing objects one pass at a
time. Timeouts, deadlock rollbacks, advisory-lock waits, and a repeat of a
recovery that already ran still spend an attempt, so a run that stops making
progress gives up exactly as before.
@greptile-apps

greptile-apps Bot commented Sep 3, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR changes migration retry accounting so successful recovery progress does not consume the no-progress attempt budget

  • Adds immutable retry-budget state and extracts deploy-failure recovery
  • Adds regression coverage for push-created databases, repeated recoveries, timeouts, failed baselines, and unrecoverable errors

Confidence Score: 5/5

The PR appears safe to merge

No blocking failure remains

Important Files Changed

Filename Overview
litellm-proxy-extras/litellm_proxy_extras/utils.py Refactors migration failure recovery and charges retry attempts only when a deploy pass makes no progress
tests/litellm-proxy-extras/test_litellm_proxy_extras_utils.py Adds focused retry-accounting tests covering progressive and non-progressing migration recovery paths

Reviews (2): Last reviewed commit: "refactor(proxy-extras): pull the migrate..." | Re-trigger Greptile

Comment thread litellm-proxy-extras/litellm_proxy_extras/utils.py Outdated
@codecov

codecov Bot commented Sep 3, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

…o a budget helper

_setup_database_v2 decided the next attempt budget inline in eight
branches, each rebinding budget before continuing. The branches now live
in _budget_after_deploy_failure, which returns the budget the next pass
runs under, and the two identical idempotent-recovery blocks share
_mark_migration_applied. The loop backs off whenever a pass spent an
attempt, which is the same set of paths that slept before.
@mateo-berri mateo-berri added run-ci and removed run-ci labels Sep 3, 2026
@mateo-berri

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 8572544. Configure here.

@mateo-berri
mateo-berri merged commit 45495e1 into litellm_internal_staging Sep 3, 2026
117 of 122 checks passed
@mateo-berri
mateo-berri deleted the litellm_fix_v2_migration_resolver_attempt_accounting branch September 3, 2026 15:36
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants