fix(db): stagger the cleanup/model-sync 6h schedulers and index conversation_turn_nodes.last_seen_at (#13973) - #14567
Merged
diegosouzapw merged 4 commits intoSep 24, 2026
Conversation
…rsation_turn_nodes.last_seen_at (#13973)
…enumber migration to 186 (#13973) The stagger added 45 min to the model-sync PERIOD (6h45m), so the two schedulers only drifted and re-collided, and MODEL_SYNC_INTERVAL_HOURS was silently overridden. Arm the recurring interval after a one-shot 45 min phase timer instead, keeping effectiveIntervalMs unchanged, and cancel the pending arm on stop. Migration 185 collided with 185_usage_history_cpa_auth_index.sql on the release tip (#14544), which aborts boot; renumber to 186.
…phase test Without the flush, the fire-and-forget import armed a real loopback poll after the stubs were restored, leaking a fetch into the concurrency-cap test.
This was referenced Sep 24, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #13973
Root cause
PR #14005 already removed the periodic
wal_checkpoint(TRUNCATE)that caused the original SIGBUS. Two items were explicitly left out of that fix and stayed the remaining scope of this issue:src/lib/db/cleanup.ts'sCLEANUP_INTERVAL_MSandsrc/shared/services/modelSyncScheduler.ts'sDEFAULT_INTERVAL_MSare both exactly 6h with zero relative offset, and both schedulers are started back-to-back in the same boot sequence (src/instrumentation-node.ts→startCleanupScheduler()alongsideensureCloudSyncInitialized()→startModelSyncScheduler()), so their periodic ticks land in the same wall-clock second every 6 hours for the life of the process.conversation_turn_nodeshas no index onlast_seen_at(migration 156 only indexesconversation_id/parent_id/content_hash), whichcleanup.ts's own doc comment already flagged as making the retentionDELETEa full table scan.Fix
src/shared/services/modelSyncScheduler.ts: addedMODEL_SYNC_STAGGER_OFFSET_MS(45 minutes) and apply it to the periodicsetIntervaldelay (effectiveIntervalMs + MODEL_SYNC_STAGGER_OFFSET_MS), so the model-sync scheduler's tick is phase-shifted relative tocleanup.ts's own un-offset 6h interval.cleanup.tsitself is unchanged — it keeps its plain 6h cadence, which is now the reference phase.src/lib/db/migrations/185_conversation_turn_nodes_last_seen_index.sql: additiveCREATE INDEX IF NOT EXISTS idx_turn_nodes_last_seen ON conversation_turn_nodes(last_seen_at).src/lib/db/cleanup.ts: updated thecleanupConversationTurnNodesdoc comment to stop claiminglast_seen_athas no index.Remaining
A third scheduler with the same default 6h period exists (
src/lib/db/core.ts's DB health-check timer,getDbHealthCheckIntervalMs()/startDbHealthCheckScheduler()), gated off entirely during automated test runs (isAutomatedTestProcess()), so no test currently proves or disproves a collision with it, andsrc/lib/db/core.tsis at 1799/1800 lines on the frozen file-size baseline (essentially no headroom to add a stagger there safely). Left out of this PR's surgical scope; flagging as a candidate follow-up rather than guessing at an untested change to a near-frozen file.Regression test
tests/unit/scheduler-6h-stagger-13973.test.ts— spies on the globalsetInterval/setTimeoutconstructors around the real, unmockedstartCleanupScheduler()/startModelSyncScheduler()calls and captures the delay each one registers.AssertionError [ERR_ASSERTION]: cleanup scheduler and model-sync scheduler must register a staggered 6h setInterval with a non-zero relative offset ... { actual: 21600000, expected: 21600000, operator: 'notStrictEqual' }— both schedulers registered the identical 21600000ms delay.pass 1, fail 0— cleanup keeps21600000ms, model-sync registers21600000 + 2700000 = 24300000ms (45-minute stagger), and the assertions on both the delay-length and the "cleanup unstaggered / model-sync staggered" relationship hold.Existing tests
tests/unit/model-sync-scheduler.test.ts— aligned the one pre-existing assertion that hardcoded the periodic interval delay as exactly6 * 60 * 60 * 1000(that assertion encoded the old, now-fixed same-second-collision contract) to6 * 60 * 60 * 1000 + scheduler.MODEL_SYNC_STAGGER_OFFSET_MS. No other assertion in that file touches the interval delay value. Ran green (EXIT:0).tests/unit/db-cleanup-conversation-nodes-12453.test.ts— unrelated to the index/stagger change, ran to confirm no regression incleanupConversationTurnNodes. Ran green (EXIT:0).Gates run
npm run typecheck:core→ exit 0npx eslint --suppressions-location config/quality/eslint-suppressions.json <every changed file>→ exit 0node scripts/check/check-file-size.mjs→ OK (154 frozen files, 4773 checked; neither touched file is on the frozen list)node scripts/check/check-complexity-ratchets.mjs --base-ref origin/release/v3.8.51→ OK — new-code mode, 2 changed files in scope, 0 cyclomatic violations, cognitive complexity unchanged at its pre-existing base of 1node scripts/check/check-changelog-integrity.mjs→ OKnode scripts/check/check-mutation-test-coverage.mjs --strict→ fails, but on 4 modules this PR does not touch (open-sse/services/accountFallback.ts,src/sse/services/auth.ts,open-sse/services/combo/comboStructure.ts,open-sse/services/combo/quotaScoring.ts) and 8 pre-existing test files not registered instryker.conf.json'stap.testFiles; none of this PR's files or new test appear in that gate's output. Pre-existing gap unrelated to this diff.node --import tsx/esm --test tests/unit/scheduler-6h-stagger-13973.test.ts→ RED before fix, GREEN after (see above)node --import tsx/esm --test tests/unit/model-sync-scheduler.test.ts→ GREEN after aligning the superseded assertionnode --import tsx/esm --test tests/unit/db-cleanup-conversation-nodes-12453.test.ts→ GREEN, unaffectedPlan-file:
_tasks/pipeline/bugs/2-implementing/13973-fix-backend-three-6h-schedulers-fire-in-the-same-second-trunca.plan.mdRework (merge-batch 2026-09-23)
Two defects found in review, fixed in this branch:
185_conversation_turn_nodes_last_seen_index.sqlcollided with185_usage_history_cpa_auth_index.sql, which landed on the release tip in feat(providers): correlate X-CPA-TRACE-ID auth_index with usage history #14544. The migration runner throwsMigration version collision detected, so boot fails. The migration is now186_conversation_turn_nodes_last_seen_index.sql, and the comment incleanup.tswas updated to match. The index itself is unchanged. Heads-up: open PRs fix(call-logs): record encrypted reasoning presence, duration and effort #14680, fix(dashboard): TPS over generation time (duration − TTFT) + reasoning-aware numerator (#13130) #13373 and feat(api): authenticate Claude Code to /v1/* with Entra ID SSO #14665 also add a186_*migration, so whichever of them lands after this one will need to renumber.setIntervalperiod (6h45m). With a different period the two schedulers drift apart and then collide again, and the operator'sMODEL_SYNC_INTERVAL_HOURSwas silently overridden. The fix arms the recurring interval from a one-shotMODEL_SYNC_STAGGER_OFFSET_MS(45 min, unref'd) phase timer. The period stays exactlyeffectiveIntervalMs, andstopModelSyncScheduler()also cancels a pending phase timer.Tests:
tests/unit/scheduler-6h-stagger-13973.test.tswas rewritten. It asserts that cleanup arms 6h at boot, that model-sync arms no interval at boot and only a phase timer ofMODEL_SYNC_STAGGER_OFFSET_MS, and that the interval it arms is exactly 6h. The capture timers are inert, so no real DB/HTTP work runs.tests/unit/model-sync-scheduler.test.tswas aligned to the new behavior. A new test coversMODEL_SYNC_INTERVAL_HOURS=4: the period stays at exactly 4h, the first tick is delayed by the offset, and stop cancels the pending arm.e2a731af) the new or updated tests fail (3 tests). With the fix,model-sync-schedulerpasses 16/16 (2 runs) andscheduler-6h-stagger-13973passes 1/1.origin/release/v3.8.51: migration-135-numbering-collision, db-migration-version-uniqueness, check-migration-numbering, db-migration-runner, db-core-migration and db-migration-missing-physical-schema.check-migration-numberingis OK.Gates:
typecheck:coreshows only the inheritedcliproxyAccountHealth.ts(157,5)error.check:open-sse-typecheckhas 2 inheritedauggie.tsregressions from the tip (#14215, tracked in #14547); this PR has noopen-sse/diff. eslint is clean on the changed files, and file-size is OK when measured after the commit.