fix(cron): reduce polling interval and recover advance_next_run crash window - #51061
Closed
JoaoMarcos44 wants to merge 3 commits into
Closed
JoaoMarcos44 wants to merge 3 commits into
JoaoMarcos44 wants to merge 3 commits into
Conversation
… window (NousResearch#51038) Two complementary fixes for cron jobs that miss their trigger window when the gateway crashes between advance_next_run() and mark_job_run(): 1. Reduce InProcessCronScheduler default polling interval from 60 s to 15 s via _resolve_tick_interval(). Overridable with HERMES_CRON_INTERVAL env var or cron.tick_interval_seconds in config.yaml. Shorter interval shrinks the crash window proportionally. 2. Add catch-up elif in _get_due_jobs_locked(): when a job has last_run_at=None and next_run_at is far in the future (> grace), check whether now is past the job's expected first run (created_at + interval, or croniter for cron kind). If so, the job is a crash-window victim and is returned as due. The _next_run_was_recovered flag prevents the same path from double-firing a legitimately new job that was just assigned its first next_run_at. Closes NousResearch#51038 Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…up function More conservative approach: instead of injecting the advance_next_run crash catch-up logic into _get_due_jobs_locked() (called every tick), isolate it in a new recover_crash_window_jobs() function called once at startup from InProcessCronScheduler.start(). - Removes elif from _get_due_jobs_locked(): hot-path is unchanged from pre-fix - Removes _next_run_was_recovered local variable (no longer needed) - Adds recover_crash_window_jobs(): scans jobs.json once at startup, resets next_run_at=now for any crash-window victims so the first tick fires them - Reverts default interval to 60 s (kept _resolve_tick_interval for opt-in) - Updates 4 crash-window tests to exercise recover_crash_window_jobs() directly; adds multi-victim test covering the 11-job scenario from NousResearch#51038 Zero regressions: 14 failed, 525 passed (all pre-existing, same baseline) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…tion Add a CRASH WINDOW paragraph to advance_next_run() docstring explaining that a gateway crash after this call but before mark_job_run() leaves next_run_at pointing to the next period while last_run_at stays null. Points future readers to recover_crash_window_jobs() as the mitigation. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
13 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #51038
O que aconteceu
Em 22/06/2026, 11 jobs de cron falharam em disparar nos horários agendados. Dois confirmados:
Sintoma:
next_run_atpassou normalmente, maslast_run_atpermaneceunull. O clock estava sincronizado via NTP, então não era fuso nem deriva.Investigação
Tracei o fluxo do
tick()emcron/scheduler.py:Se o gateway crashar entre (1) e (2):
Esse é o crash window do design at-most-once: o
advance_next_runprotege contra re-disparo em loop, mas cria um ponto cego quando o processo cai entre a escrita e a execução. O intervalo de 60s de polling amplifica a janela — qualquer restart dentro desse minuto bate no problema.Solução
Criei
recover_crash_window_jobs()emcron/jobs.py: uma função de recuperação one-shot chamada uma única vez no startup do ticker, antes do primeiro tick. Ela escaneiajobs.jsone reseta onext_run_atde qualquer vítima paranow, fazendo o primeiro tick normal disparar o job.Optei por startup-only em vez de inline a cada tick para não tocar no hot-path de
get_due_jobs().Critérios de detecção de vítima — todos precisam ser verdadeiros:
last_run_at = nullnext_run_at > now(next_run_at - now) > gracenow > first_run_expectedcreated_at + intervalpara interval;croniter.get_next(created_at)para cron)O guard
first_run_expectedé o principal antídoto contra falso positivo: um job novo cujo primeiro disparo ainda não chegou nunca é acionado prematuramente.Também adicionei
_resolve_tick_interval()emscheduler_provider.pypara tuning viaHERMES_CRON_INTERVAL(env var) oucron.tick_interval_seconds(config.yaml), sem alterar o default de 60s.Testes
Comparei as duas abordagens (inline elif a cada tick vs. startup recovery) com 500 repetições em 100 jobs — diferença de performance dentro do ruído de medição (~1.8ms vs ~2.0ms por call). Escolhi startup recovery por ser mais cirúrgica e não tocar no hot-path.
8 novos testes em
TestResolveTickInterval— env var tem prioridade, config fallback, clamping mínimo 5s, valor inválido, default 60s.4 novos testes em
TestGetDueJobs— job resetado quandofirst_run_expectedpassou; job intocado quando ainda no futuro; job comlast_run_atignorado; múltiplas vítimas resetadas em um único pass (cobre o cenário de 11 jobs da issue).Bateria final:
14 falhas são todas pré-existentes: testes de permissão Unix no Windows,
croniter/httpx/dotenvausentes no CI. Zero regressões.