Skip to content

fix(cron): keep server-local jobs on their wall clock across DST changes - #119970

Open
jonpol01 wants to merge 2 commits into
NousResearch:mainfrom
jonpol01:sweep/cron-server-local-tz-dst-fixed-offset
Open

jonpol01 wants to merge 2 commits into
NousResearch:mainfrom
jonpol01:sweep/cron-server-local-tz-dst-fixed-offset

Conversation

@jonpol01

Copy link
Copy Markdown

What does this PR do?

With timezone unset (the default), hermes_time.now() falls back to datetime.now().astimezone(), which has a fixed UTC offset. Cron treated that offset as the zone in three places:

  • compute_next_run attached the last run's offset to croniter's next wall clock.
  • parse_schedule attached the creation-time offset to a naive one-shot.
  • _ensure_aware read legacy naive timestamps with today's offset.

As a result, every wall clock on the far side of a DST change landed an hour off. On a New York host the first 0 9 * * * run after spring-forward was stored at 09:00-05:00 and fired at 10:00 EDT, with a false cron.timezone_migration.catch_up WARNING. A naive one-shot typed in February for 1 April also fired at 10:00 EDT.

All three sites now go through one helper, _wall_clock_readings(wall, zone):

  • Configured zone: it returns the same fold=0/fold=1 readings as before.
  • No zone configured: it reads the wall clock with the host offset in force a day either side of it, and keeps the readings the host agrees with at that instant.
  • Spring-forward gap: it agrees with neither, so both readings stay, earlier offset first. That is how zoneinfo orders them.

Server-local mode and the same zone configured now resolve every wall clock to the same instant. I checked this every 15 minutes across 2026 in seven zones, including Lord Howe's 30-minute shift and southern-hemisphere rules, with 0 mismatches.

Why not plain naive.astimezone()? On CPython 3.11 its fold=0 reading of a skipped wall clock is the earlier instant, the reverse of zoneinfo. A 30 2 * * * job would have moved from 03:30 EDT to 01:30 EST on spring-forward day.

The fix stays in cron/jobs.py and does not change hermes_time.now(), so no other clock consumer sees a different tzinfo.

Related: #66436 / #66456 are the configured-zone half of this bug (croniter reusing the base offset), which main fixed in d6d29b0 by anchoring croniter to the configured IANA zone. That fix cannot reach the default: with no timezone configured, the zone is a fixed UTC offset, so re-attaching it keeps the wrong offset. This PR covers that default case.

Related Issue

Fixes #119969

Type of Change

  • 🐛 Bug fix (non-breaking change that fixes an issue)
  • ✨ New feature (non-breaking change that adds functionality)
  • 🔒 Security fix
  • 📝 Documentation update
  • ✅ Tests (adding or improving test coverage)
  • ♻️ Refactor (no behavior change)
  • 🎯 New skill (bundled or hub)

Changes Made

  • cron/jobs.py:

    • New _wall_clock_readings() helper.
    • compute_next_run uses zone = get_timezone(), so an unset zone means the host's zone rather than base_time's fixed offset, and builds its strictly-after candidates from the helper.
    • parse_schedule reads naive one-shots through the helper.
    • _ensure_aware reads legacy naive values through the helper.
    • parse_schedule's timestamp guard also catches OverflowError, so a naive timestamp at datetime.min/max stays an "Invalid timestamp" ValueError. Base rejects those inputs later in create_job with a ValueError, and this keeps that contract.
  • tests/cron/test_server_local_dst.py (new): sets up a host in America/New_York the way a real host does (TZ + time.tzset()), with no Hermes zone configured, and pins the scheduler clock to that host clock. It covers:

    • the first run after each DST change
    • a 365-day walk
    • parity with the configured-zone path across both transition nights for 30 2 * * *, 30 1 * * * and */20 * * * *
    • the real store loop on spring-forward day (due at 09:00, no migration catch-up)
    • naive one-shots across both transitions
    • legacy naive timestamps
    • the datetime-bounds rejection

    It is skipped where time.tzset does not exist (native Windows).

  • tests/cron/test_cron_timezone_migration_catchup.py and tests/cron/test_jobs.py: six tests stood in for a configured zone by patching the clock with a fixed offset while leaving the zone unset. Unset now means the host's zone (UTC in the suite), so they configure the zone they describe: Europe/Brussels and Asia/Kolkata. Their assertions are unchanged, and they pass on main and with this change.

  • tests/cron/test_timezone.py: the legacy-naive oracle built its expected value from today's host offset, the reading under fix. It now uses the host's reading of that date.

How to Test

  1. TZ=UTC pytest tests/cron/test_server_local_dst.py -q: 11 pass. With cron/jobs.py from main, 9 fail. The 2 bounds-guard tests pass on both by design.
  2. Run the repro script from the issue under TZ=America/New_York with a scratch HERMES_HOME. Main prints 2026-03-08T09:00:00-05:00 and 2026-04-01T09:00:00-05:00 (10:00 EDT). This branch prints -04:00 for both (09:00 EDT), the same as main with HERMES_TIMEZONE=America/New_York.
  3. Sabotage: I reverted each part in turn and only its own tests failed.
    • compute_next_run hunk → the 6 recurring tests
    • parse_schedule hunk → the 2 one-shot tests
    • _ensure_aware hunk → the legacy test
    • dropping the host-agreement filter → the 6 recurring tests
    • flipping the gap order → [30 2 * * *] parity
    • dropping OverflowError from the guard → the 2 bounds tests

Checklist

Code

  • I've read the Contributing Guide
  • My commit messages follow Conventional Commits (fix(scope):, feat(scope):, etc.)
  • I searched for existing PRs to make sure this isn't a duplicate
  • My PR contains only changes related to this fix/feature (no unrelated commits)
  • I've run pytest tests/ -q and all tests pass. Not the full suite. I ran every test file that exercises the cron scheduling functions, one file at a time under the hermetic runner env: all 85 files in tests/cron, the cron-related files in tests/hermes_cli, tests/gateway, tests/tools, tests/plugins and tests/agent, and the new file. 1562 passed, 0 failed.
  • I've added tests for my changes (required for bug fixes, strongly encouraged for features)
  • I've tested on my platform: macOS 27.0

Documentation & Housekeeping

  • I've updated relevant documentation (README, docs/, docstrings) — or N/A. The configuration docs already say empty means server-local time; this makes that hold across DST.
  • I've updated cli-config.yaml.example if I added/changed config keys — or N/A
  • I've updated CONTRIBUTING.md or AGENTS.md if I changed architecture or workflows — or N/A
  • I've considered cross-platform impact (Windows, macOS) per the compatibility guide — or N/A. The helper uses only datetime.astimezone(), which works on every platform. The new tests skip where time.tzset is missing. scripts/check-windows-footguns.py --diff origin/main reports only a pre-existing read_text() at tests/cron/test_jobs.py:1327, which this PR does not touch.
  • I've updated tool descriptions/schemas if I changed tool behavior — or N/A

Screenshots / Logs

The real store loop on a New York host with no timezone configured: the Saturday 09:00 run is recorded, then the clock is stepped to Sunday.

main:
stored next_run_at: 2026-03-08T09:00:00-05:00
Sun 09:00:30 EDT due=False catchups=0
  WARNING cron.timezone_migration.catch_up job='morning' ... stored=2026-03-08T09:00:00-05:00 normalized=2026-03-08T10:00:00-04:00
Sun 10:00:30 EDT due=True catchups=1

this branch:
stored next_run_at: 2026-03-08T09:00:00-04:00
Sun 09:00:30 EDT due=True catchups=0

@alt-glitch alt-glitch added type/bug Something isn't working P2 Medium — degraded but workaround exists comp/cron Cron scheduler and job management sweeper:risk-compatibility Sweeper risk: may break existing users, config, migrations, defaults, or upgrades labels Sep 23, 2026
teknium1 added a commit that referenced this pull request Sep 23, 2026
…y cells

A strict xfail on a gap whose fix is an open PR turns main red the moment
that fix merges (XPASS), and the non-strict ones guarded nothing. Each gap
now has a probe (tests/e2e/core/delivery/_pending_fixes.py) that reproduces
the defect's mechanism on the tree under test in a throwaway interpreter;
expect_gap() applies the strict xfail only while the probe still reproduces
it, so the cell becomes a plain test once the fix is in the tree, whatever
the merge order. Covered: #120314, #119970 (soak), #120315, #120377,
#120444 (C12) and #120450 (cell 5). Each probe was checked against every
fix head: it flips on its own PR and on no other.

C12 cells made deterministic (identical outcome on every run):
- long_split streams the whole reply as one chunk; long_streamed and
  stream_timeout_first_send pace chunks so each lands in its own consumer
  tick. The five former coin-flip xfails are now two plain cells and three
  strict #120315 gap cells.
- sent_ack_lost waits until the answer is persisted before the kill, so it
  pins the #120377 recovery; the streamed-before-persisted order is its own
  cell (stream_accepted_unpersisted, a strict live gap with a stalled
  provider stream).
- zzz_unclean_restart compares director.resumes against a snapshot taken
  before its kill instead of requiring it empty: crash cells on their own
  homes may legitimately resume.
- the whole-run audit skips the reconnect replay only while #120444's gap
  is open.
teknium1 added a commit that referenced this pull request Sep 23, 2026
The soak's second ticker process never contends for a fire: the tick lock
admits one ticker per instant and the winner advances next_run_at before it
releases, so neutralising claim_job_for_fire left the soak green.

test_two_replicas_contend_for_every_fire drives the path the claim actually
guards, CronScheduler.fire_due ("exactly one of N replicas runs a job"): this
process and a second OS process receive the same fire for every due
occurrence of two jobs. A file barrier parks each replica inside
claim_job_for_fire until both are there (it only delays), and the winner's
run is held open, so the loser always meets a live claim. Per round exactly
one replica claims, one run starts for that slot and is delivered once, the
loser's attempt is recorded as not acquired, and the store re-arms past now.
Skipping the live-claim check or the cross-process jobs lock turns it red
on the first round.

The soak's DST scenarios gate their xfails on the #120314 / #119970 probes.
teknium1 added a commit that referenced this pull request Sep 23, 2026
…y cells

A strict xfail on a gap whose fix is an open PR turns main red the moment
that fix merges (XPASS), and the non-strict ones guarded nothing. Each gap
now has a probe (tests/e2e/core/delivery/_pending_fixes.py) that reproduces
the defect's mechanism on the tree under test in a throwaway interpreter;
expect_gap() applies the strict xfail only while the probe still reproduces
it, so the cell becomes a plain test once the fix is in the tree, whatever
the merge order. Covered: #120314, #119970 (soak), #120315, #120377,
#120444 (C12) and #120450 (cell 5). Each probe was checked against every
fix head: it flips on its own PR and on no other.

C12 cells made deterministic (identical outcome on every run):
- long_split streams the whole reply as one chunk; long_streamed and
  stream_timeout_first_send pace chunks so each lands in its own consumer
  tick. The five former coin-flip xfails are now two plain cells and three
  strict #120315 gap cells.
- sent_ack_lost waits until the answer is persisted before the kill, so it
  pins the #120377 recovery; the streamed-before-persisted order is its own
  cell (stream_accepted_unpersisted, a strict live gap with a stalled
  provider stream).
- zzz_unclean_restart compares director.resumes against a snapshot taken
  before its kill instead of requiring it empty: crash cells on their own
  homes may legitimately resume.
- the whole-run audit skips the reconnect replay only while #120444's gap
  is open.
teknium1 added a commit that referenced this pull request Sep 23, 2026
The soak's second ticker process never contends for a fire: the tick lock
admits one ticker per instant and the winner advances next_run_at before it
releases, so neutralising claim_job_for_fire left the soak green.

test_two_replicas_contend_for_every_fire drives the path the claim actually
guards, CronScheduler.fire_due ("exactly one of N replicas runs a job"): this
process and a second OS process receive the same fire for every due
occurrence of two jobs. A file barrier parks each replica inside
claim_job_for_fire until both are there (it only delays), and the winner's
run is held open, so the loser always meets a live claim. Per round exactly
one replica claims, one run starts for that slot and is delivered once, the
loser's attempt is recorded as not acquired, and the store re-arms past now.
Skipping the live-claim check or the cross-process jobs lock turns it red
on the first round.

The soak's DST scenarios gate their xfails on the #120314 / #119970 probes.
teknium1 added a commit that referenced this pull request Sep 23, 2026
…y cells

A strict xfail on a gap whose fix is an open PR turns main red the moment
that fix merges (XPASS), and the non-strict ones guarded nothing. Each gap
now has a probe (tests/e2e/core/delivery/_pending_fixes.py) that reproduces
the defect's mechanism on the tree under test in a throwaway interpreter;
expect_gap() applies the strict xfail only while the probe still reproduces
it, so the cell becomes a plain test once the fix is in the tree, whatever
the merge order. Covered: #120314, #119970 (soak), #120315, #120377,
#120444 (C12) and #120450 (cell 5). Each probe was checked against every
fix head: it flips on its own PR and on no other.

C12 cells made deterministic (identical outcome on every run):
- long_split streams the whole reply as one chunk; long_streamed and
  stream_timeout_first_send pace chunks so each lands in its own consumer
  tick. The five former coin-flip xfails are now two plain cells and three
  strict #120315 gap cells.
- sent_ack_lost waits until the answer is persisted before the kill, so it
  pins the #120377 recovery; the streamed-before-persisted order is its own
  cell (stream_accepted_unpersisted, a strict live gap with a stalled
  provider stream).
- zzz_unclean_restart compares director.resumes against a snapshot taken
  before its kill instead of requiring it empty: crash cells on their own
  homes may legitimately resume.
- the whole-run audit skips the reconnect replay only while #120444's gap
  is open.
teknium1 added a commit that referenced this pull request Sep 23, 2026
The soak's second ticker process never contends for a fire: the tick lock
admits one ticker per instant and the winner advances next_run_at before it
releases, so neutralising claim_job_for_fire left the soak green.

test_two_replicas_contend_for_every_fire drives the path the claim actually
guards, CronScheduler.fire_due ("exactly one of N replicas runs a job"): this
process and a second OS process receive the same fire for every due
occurrence of two jobs. A file barrier parks each replica inside
claim_job_for_fire until both are there (it only delays), and the winner's
run is held open, so the loser always meets a live claim. Per round exactly
one replica claims, one run starts for that slot and is delivered once, the
loser's attempt is recorded as not acquired, and the store re-arms past now.
Skipping the live-claim check or the cross-process jobs lock turns it red
on the first round.

The soak's DST scenarios gate their xfails on the #120314 / #119970 probes.
teknium1 added a commit that referenced this pull request Sep 23, 2026
…y cells

A strict xfail on a gap whose fix is an open PR turns main red the moment
that fix merges (XPASS), and the non-strict ones guarded nothing. Each gap
now has a probe (tests/e2e/core/delivery/_pending_fixes.py) that reproduces
the defect's mechanism on the tree under test in a throwaway interpreter;
expect_gap() applies the strict xfail only while the probe still reproduces
it, so the cell becomes a plain test once the fix is in the tree, whatever
the merge order. Covered: #120314, #119970 (soak), #120315, #120377,
#120444 (C12) and #120450 (cell 5). Each probe was checked against every
fix head: it flips on its own PR and on no other.

C12 cells made deterministic (identical outcome on every run):
- long_split streams the whole reply as one chunk; long_streamed and
  stream_timeout_first_send pace chunks so each lands in its own consumer
  tick. The five former coin-flip xfails are now two plain cells and three
  strict #120315 gap cells.
- sent_ack_lost waits until the answer is persisted before the kill, so it
  pins the #120377 recovery; the streamed-before-persisted order is its own
  cell (stream_accepted_unpersisted, a strict live gap with a stalled
  provider stream).
- zzz_unclean_restart compares director.resumes against a snapshot taken
  before its kill instead of requiring it empty: crash cells on their own
  homes may legitimately resume.
- the whole-run audit skips the reconnect replay only while #120444's gap
  is open.
teknium1 added a commit that referenced this pull request Sep 23, 2026
The soak's second ticker process never contends for a fire: the tick lock
admits one ticker per instant and the winner advances next_run_at before it
releases, so neutralising claim_job_for_fire left the soak green.

test_two_replicas_contend_for_every_fire drives the path the claim actually
guards, CronScheduler.fire_due ("exactly one of N replicas runs a job"): this
process and a second OS process receive the same fire for every due
occurrence of two jobs. A file barrier parks each replica inside
claim_job_for_fire until both are there (it only delays), and the winner's
run is held open, so the loser always meets a live claim. Per round exactly
one replica claims, one run starts for that slot and is delivered once, the
loser's attempt is recorded as not acquired, and the store re-arms past now.
Skipping the live-claim check or the cross-process jobs lock turns it red
on the first round.

The soak's DST scenarios gate their xfails on the #120314 / #119970 probes.
teknium1 added a commit that referenced this pull request Sep 23, 2026
…y cells

A strict xfail on a gap whose fix is an open PR turns main red the moment
that fix merges (XPASS), and the non-strict ones guarded nothing. Each gap
now has a probe (tests/e2e/core/delivery/_pending_fixes.py) that reproduces
the defect's mechanism on the tree under test in a throwaway interpreter;
expect_gap() applies the strict xfail only while the probe still reproduces
it, so the cell becomes a plain test once the fix is in the tree, whatever
the merge order. Covered: #120314, #119970 (soak), #120315, #120377,
#120444 (C12) and #120450 (cell 5). Each probe was checked against every
fix head: it flips on its own PR and on no other.

C12 cells made deterministic (identical outcome on every run):
- long_split streams the whole reply as one chunk; long_streamed and
  stream_timeout_first_send pace chunks so each lands in its own consumer
  tick. The five former coin-flip xfails are now two plain cells and three
  strict #120315 gap cells.
- sent_ack_lost waits until the answer is persisted before the kill, so it
  pins the #120377 recovery; the streamed-before-persisted order is its own
  cell (stream_accepted_unpersisted, a strict live gap with a stalled
  provider stream).
- zzz_unclean_restart compares director.resumes against a snapshot taken
  before its kill instead of requiring it empty: crash cells on their own
  homes may legitimately resume.
- the whole-run audit skips the reconnect replay only while #120444's gap
  is open.
teknium1 added a commit that referenced this pull request Sep 23, 2026
The soak's second ticker process never contends for a fire: the tick lock
admits one ticker per instant and the winner advances next_run_at before it
releases, so neutralising claim_job_for_fire left the soak green.

test_two_replicas_contend_for_every_fire drives the path the claim actually
guards, CronScheduler.fire_due ("exactly one of N replicas runs a job"): this
process and a second OS process receive the same fire for every due
occurrence of two jobs. A file barrier parks each replica inside
claim_job_for_fire until both are there (it only delays), and the winner's
run is held open, so the loser always meets a live claim. Per round exactly
one replica claims, one run starts for that slot and is delivered once, the
loser's attempt is recorded as not acquired, and the store re-arms past now.
Skipping the live-claim check or the cross-process jobs lock turns it red
on the first round.

The soak's DST scenarios gate their xfails on the #120314 / #119970 probes.
teknium1 added a commit that referenced this pull request Sep 23, 2026
…xfail scoped to its deviation

ev_long_run used to flag a missing claim refresh as stalled_heartbeat
and carry on, so neutralising heartbeat_fire_claim in _heartbeat_loop
left all five scenarios green. Every virtual 30 s of the 15-minute hold
now (a) waits for the run's heartbeat to restamp the claim and fails if
it never does, and (b) has a contender call the real claim_job_for_fire
and asserts it loses, also once the first stamp is older than the TTL.
With the heartbeat replaced by True all 5 scenarios are red on the
refresh assertion; on main they are green.

The #119970 cell no longer carries a whole-scenario strict xfail. While
its behavioural probe reproduces, only the oracle's fired-set and
next_run_at checks may deviate: each deviation is recorded, the model
follows the stored slot, and the soak runs every virtual day with all
other checks live, XFAILing at the end only if a deviation was seen
(and failing if the probe says open but none was). The #120314 entry
is gone (merged). A scenario that fails mid-hold now releases its held
run before teardown so it cannot bleed into the next scenario.
teknium1 added a commit that referenced this pull request Sep 23, 2026
…y cells

A strict xfail on a gap whose fix is an open PR turns main red the moment
that fix merges (XPASS), and the non-strict ones guarded nothing. Each gap
now has a probe (tests/e2e/core/delivery/_pending_fixes.py) that reproduces
the defect's mechanism on the tree under test in a throwaway interpreter;
expect_gap() applies the strict xfail only while the probe still reproduces
it, so the cell becomes a plain test once the fix is in the tree, whatever
the merge order. Covered: #120314, #119970 (soak), #120315, #120377,
#120444 (C12) and #120450 (cell 5). Each probe was checked against every
fix head: it flips on its own PR and on no other.

C12 cells made deterministic (identical outcome on every run):
- long_split streams the whole reply as one chunk; long_streamed and
  stream_timeout_first_send pace chunks so each lands in its own consumer
  tick. The five former coin-flip xfails are now two plain cells and three
  strict #120315 gap cells.
- sent_ack_lost waits until the answer is persisted before the kill, so it
  pins the #120377 recovery; the streamed-before-persisted order is its own
  cell (stream_accepted_unpersisted, a strict live gap with a stalled
  provider stream).
- zzz_unclean_restart compares director.resumes against a snapshot taken
  before its kill instead of requiring it empty: crash cells on their own
  homes may legitimately resume.
- the whole-run audit skips the reconnect replay only while #120444's gap
  is open.
teknium1 added a commit that referenced this pull request Sep 23, 2026
The soak's second ticker process never contends for a fire: the tick lock
admits one ticker per instant and the winner advances next_run_at before it
releases, so neutralising claim_job_for_fire left the soak green.

test_two_replicas_contend_for_every_fire drives the path the claim actually
guards, CronScheduler.fire_due ("exactly one of N replicas runs a job"): this
process and a second OS process receive the same fire for every due
occurrence of two jobs. A file barrier parks each replica inside
claim_job_for_fire until both are there (it only delays), and the winner's
run is held open, so the loser always meets a live claim. Per round exactly
one replica claims, one run starts for that slot and is delivered once, the
loser's attempt is recorded as not acquired, and the store re-arms past now.
Skipping the live-claim check or the cross-process jobs lock turns it red
on the first round.

The soak's DST scenarios gate their xfails on the #120314 / #119970 probes.
teknium1 added a commit that referenced this pull request Sep 23, 2026
…xfail scoped to its deviation

ev_long_run used to flag a missing claim refresh as stalled_heartbeat
and carry on, so neutralising heartbeat_fire_claim in _heartbeat_loop
left all five scenarios green. Every virtual 30 s of the 15-minute hold
now (a) waits for the run's heartbeat to restamp the claim and fails if
it never does, and (b) has a contender call the real claim_job_for_fire
and asserts it loses, also once the first stamp is older than the TTL.
With the heartbeat replaced by True all 5 scenarios are red on the
refresh assertion; on main they are green.

The #119970 cell no longer carries a whole-scenario strict xfail. While
its behavioural probe reproduces, only the oracle's fired-set and
next_run_at checks may deviate: each deviation is recorded, the model
follows the stored slot, and the soak runs every virtual day with all
other checks live, XFAILing at the end only if a deviation was seen
(and failing if the probe says open but none was). The #120314 entry
is gone (merged). A scenario that fails mid-hold now releases its held
run before teardown so it cannot bleed into the next scenario.
@tsposato

tsposato commented Oct 4, 2026

Copy link
Copy Markdown

Checked this against current main (158fd638da). I cherry-picked bc18813 onto it and ran tests/e2e/core/delivery/test_cron_virtual_clock_soak.py. unset_on_newyork_spring passes, which includes the 30 2 * * * job on the 2026-03-08 spring-forward day.

Two notes in case they help review:

  • That scenario is currently an xfail on main, tied to this PR through GAPS and the 119970 probe in _pending_fixes.py. Once this merges, the probe reports "fixed" and the scenario runs as a plain test, so the GAPS entry and the PROBES[119970] entry can be deleted in the same change.
  • A narrower alternative that only fixes compute_next_run does not pass that scenario. The due scan's _cron_next_run_matches_expr can't see the DST gap through a fixed offset, so it re-anchors the 02:30 job without firing it. This PR's single _wall_clock_readings helper avoids that.

With no `timezone` configured (the shipped default), hermes_time.now() falls
back to datetime.now().astimezone(), a fixed UTC offset, and cron treated that
offset as the zone. compute_next_run attached the last run's offset to
croniter's next wall clock, parse_schedule attached the creation-time offset
to a naive one-shot, and _ensure_aware read legacy naive timestamps with
today's offset. Every wall clock on the far side of a DST change landed an
hour off. On a New York host the first `0 9 * * *` run after spring-forward
was stored at 09:00-05:00 and fired at 10:00 EDT, logging a false
timezone_migration.catch_up warning, and a naive one-shot typed in February
for 1 April fired at 10:00 EDT.

All three sites now resolve naive wall clocks through one helper. A
configured zone keeps the existing fold=0/fold=1 readings. With none, the
helper takes the host offset a day either side of the wall clock and keeps
the readings the host agrees with at that instant. A spring-forward gap
matches neither and keeps both, earlier offset first, the way zoneinfo
orders them. Server-local and configured modes now resolve every wall clock
to the same instant: checked every 15 minutes across 2026 in seven zones,
including Lord Howe's 30-minute shift and southern-hemisphere rules.
parse_schedule reports a timestamp at the datetime bounds as an invalid
timestamp, so the helper's day-either-side probe cannot surface an
OverflowError.

Six existing tests stood in for a configured zone by patching the clock
while leaving the zone unset. They now configure the zone they describe,
and their assertions are unchanged. One oracle built its expected value from
today's host offset; it now uses the host's reading of that date.
@jonpol01
jonpol01 force-pushed the sweep/cron-server-local-tz-dst-fixed-offset branch from bc18813 to 848a4a9 Compare October 4, 2026 03:44
@jonpol01

jonpol01 commented Oct 4, 2026

Copy link
Copy Markdown
Author

Thanks @tsposato, good catch. Removed the GAPS["unset_on_newyork_spring"] entry and PROBES[119970] in 848a4a9 (branch rebased onto main), so the scenario runs as a plain test once this lands. The removed probe prints fixed on this head and open on main.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/cron Cron scheduler and job management P2 Medium — degraded but workaround exists sweeper:risk-compatibility Sweeper risk: may break existing users, config, migrations, defaults, or upgrades type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: With no timezone configured, cron jobs and naive one-shots fire an hour off after a DST change

3 participants