Conversation
createManagedDbBackup() (core.ts) -- the writer behind health-check-repair backups, taken every time the periodic health check runs autoRepair -- never called into the shared retention policy at all. backup.ts's manual/auto path already resolved settings and pruned after writing; this was the one backup producer with no retention call whatsoever. Observed live: db_backups/ grew to 570 GB / 60 files against a small live database, all health-check-repair snapshots, none ever pruned. Extracted the duplicated "read persisted dbBackup setting" logic (previously copy-pasted between backup.ts's getStoredDbBackupInteger and migrationRunner.ts's now-removed readStoredBackupSetting) into backupRetention.ts as readStoredDbBackupSetting() / resolveDbBackupRetentionSettings(), and added pruneManagedDbBackups() as the shared resolve+prune+log sequence. core.ts's health-check-repair path now calls it right after every successful VACUUM INTO, the same way backup.ts's backupDbFile() already does. Does NOT touch the pre-migration backup path (migrationRunner/preMigrationBackup.ts): db-pre-migration-backup-retention-10421.test.ts explicitly asserts retention stays outside the migration window, by design -- confirmed by running the full existing suite before finalizing this change. Added tests/unit/db-backup-retention-shared.test.ts proving pruneManagedDbBackups actually caps a seeded 30-file backup directory at the configured maxFiles, never throws when pruning itself fails, and that resolveDbBackupRetentionSettings' env-override precedence holds. createManagedDbBackup() itself is gated behind isAutomatedTestProcess() (an intentional, pre-existing production-safety check) and cannot be exercised end-to-end from this test runner -- which is also exactly why the missing prune call here went unnoticed by the existing suite. 60/60 backup + health-check tests pass. eslint clean on all touched files.
…check-repair snapshots) into dev/omniroute-dev-combined
…check-repair snapshots) into dev/omniroute-dev-combined
…check-repair snapshots) into dev/omniroute-dev-combined
…check-repair snapshots) into dev/omniroute-dev-combined
…check-repair snapshots) into dev/omniroute-dev-combined
…check-repair snapshots) into dev/omniroute-dev-combined
…check-repair snapshots) into dev/omniroute-dev-combined
…check-repair snapshots) into dev/omniroute-dev-combined
…check-repair snapshots) into dev/omniroute-dev-combined
…check-repair snapshots) into dev/omniroute-dev-combined
…check-repair snapshots) into dev/omniroute-dev-combined
|
Thanks for the thorough writeup and the 570 GB / 60-file evidence — this was a Triage note: this is the review recommendation — the close itself happens only after the maintainer's per-PR sign-off (and, where a superseding PR is named, after it has landed). Nothing is being closed by this comment. |
|
Confirmed #13404 (closing #13308) merged, so per your own sequencing that's the go — closing this in its favor. Thanks for the flag on the gap: the merged fix reads maxFiles/retentionDays only from env vars, not the persisted Storage-page setting this branch also handled. Might come back and open that as a narrower follow-up off this branch. |
What this fixes
createManagedDbBackup()(src/lib/db/core.ts) — the writer behindhealth-check-repairbackups, taken every time the periodic health check runsautoRepair— never called into the shared retention policy at all.backup.ts's manual/auto path already resolves the operator's settings and prunes right after writing; this was the one backup producer with no retention call whatsoever.Observed live:
db_backups/grew to 570 GB / 60 files against a small live database, allhealth-check-repairsnapshots, none ever pruned.What changed
dbBackupsetting" logic (previously copy-pasted betweenbackup.ts'sgetStoredDbBackupIntegerandmigrationRunner.ts's now-removedreadStoredBackupSetting) intobackupRetention.tsasreadStoredDbBackupSetting()/resolveDbBackupRetentionSettings().pruneManagedDbBackups()tobackupRetention.tsas the shared resolve+prune+log sequence.core.ts's health-check-repair path now callspruneManagedDbBackups()right after every successfulVACUUM INTO, the same waybackup.ts'sbackupDbFile()already does.What this deliberately does NOT touch
The pre-migration backup path (
migrationRunner/preMigrationBackup.ts).db-pre-migration-backup-retention-10421.test.tsexplicitly asserts retention stays outside the concurrent migration window — that's an intentional design decision (content-addressed snapshots, reused for an identical DB state), not a gap. I initially assumed the missing prune call there (also dropped when this file was extracted frommigrationRunner.ts) was an unrelated regression and tried to fix it too; running the full existing suite caught that immediately (successful migrations retain existing backups and do not prune inside the migration windowfailed), so I reverted that part. Only the health-check-repair path — which has no such constraint — is touched here.Evidence
tests/unit/db-backup-retention-shared.test.ts— provespruneManagedDbBackupscaps a seeded 30-file backup directory at the configuredmaxFiles(Pruned 25 old backup file(s) (5 kept...)), never throws when pruning itself fails, and thatresolveDbBackupRetentionSettings's env-override precedence holds.createManagedDbBackup()itself is gated behindisAutomatedTestProcess()(an intentional, pre-existing production-safety check) and cannot be exercised end-to-end from this test runner — which is also exactly why the missing prune call went unnoticed by the existing suite for as long as it did.db-pre-migration-backup-retention-10421,db-backup-extended,db-backup-autobackup-setting-5871,db-backups-skills-3500,cli-backup-command,db-health-check,db-health-driver, plus the new file).eslintclean on all touched files (3 pre-existing unused-var errors incore.tsare unrelated to this change and present on unmodifiedrelease/v3.8.51too).Note for maintainer
Also relevant to my own deployment: I have
DB_BACKUP_MAX_FILES=5set via env, which this fix now actually enforces forhealth-check-repairbackups too.