Conversation
Long-running cron jobs were blocking the entire scheduler tick loop because run_job() executed while the tick lock was still held. Scope the lock to due-job selection only, then release it before executing jobs. This prevents slow jobs from delaying unrelated due jobs and blocking the next 60-second gateway tick. Fixes NousResearch#3752
|
Thanks for identifying the lock starvation issue @Sertug17! After reviewing, we think the current approach is safe enough — Moving the job loop outside the lock would introduce a race window between Appreciate you thinking about scheduler robustness though — this is a real concern for heavy cron users. |
|
Thanks for the detailed explanation @teknium1! Makes sense — if I'll close this PR. Appreciate you taking the time to review! |
What does this PR do?
Long-running cron jobs were blocking the entire scheduler tick loop because
run_job()executed while the tick lock was still held. One slow job could delay all unrelated due jobs and block the next 60-second gateway tick — effectively causing scheduler starvation.This PR scopes the file lock to due-job selection only, then releases it before executing jobs. Each job now runs outside the global tick lock.
Related Issue
Fixes #3752
Type of Change
Changes Made
cron/scheduler.py: Movedrun_job()loop outside the tick lock scope. Lock is released immediately afterget_due_jobs()completes.How to Test
Checklist